Principal Site Reliability Engineer (m/f/x)

Commerce Tools
Berlin, London, Munich
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureExperience: 7+ yearsSkills: ["Problem-solving","Leadership","Mentoring","Communication","Customer focus"]

Own resiliency end-to-end for mission-critical commerce infrastructure, driving incident management maturity, real-time operational visibility, and data-driven process improvements. Scale and lead organization-wide readiness for peak events like Black Friday, standardizing detection, response, communication, and postmortems. Partner with product engineering and domain Principal Engineers across infrastructure, cloud, security, API, architecture, and performance, while fostering documentation, training, and cross-team adoption.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Commerce Tools
Commerce Tools
2 days ago

Principal Site Reliability Engineer (m/f/x)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 1 hour agoStatus: Live

Job Summary

Own resiliency end-to-end for mission-critical commerce infrastructure, driving incident management maturity, real-time operational visibility, and data-driven process improvements. Scale and lead organization-wide readiness for peak events like Black Friday, standardizing detection, response, communication, and postmortems. Partner with product engineering and domain Principal Engineers across infrastructure, cloud, security, API, architecture, and performance, while fostering documentation, training, and cross-team adoption.
Location: Berlin, London, Munich
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Standardize incident management across detection, response, communication, and postmortems.
  • •Build real-time metrics, dashboards, and signals to track system health and incident trends.
  • •Use operational data to identify process gaps and partner with product engineering teams to fix them.
  • •Lead organization-wide readiness for peak-traffic events like Black Friday.
  • •Drive cross-team resiliency initiatives and enable knowledge sharing through documentation and training.

Pay and Benefits

Perks:Health InsuranceLearning BudgetLanguage TrainingParental LeaveEquity

Key Requirements

  • •7+ years driving incident management and operational excellence, plus 5+ years leading org-wide resiliency and reliability initiatives.
  • •Proven experience scaling systems for high-stakes, multi-team operational events such as Black Friday and major launches.
  • •Strong data literacy to analyze metrics, diagnose issues, and measure process improvements.
  • •Ability to evaluate both technical issues and organizational/human dynamics behind incidents.
  • •Demonstrated success managing large-scale, multi-team initiatives in an Agile environment, with strong written and verbal communication.
Experience:7+ years
Skills:Problem-solvingLeadershipMentoringCommunicationCustomer focus
Languages:English

Company Brief

Commerce Tools
Provides a frontend-as-a-service platform for headless commerce, enabling teams to build, deploy, and manage composable storefronts. Integrates with commercetools and other backend services to deliver omnichannel customer experiences.
Industry: SaaS
Company Size: Medium (51 to 250 employees)
Growth: Growth Stage Startup
Headquarters: Berlin, Germany
WebsiteLinkedIn