Principal Site Reliability Engineer

Commerce Tools
Berlin, London
Workplace: HybridFull timeFunction: Communications, PR & CommunitySkills: ["Problem-solving","Leadership","Communication","Mentoring","Data analysis"]

Drive commercetools’ resiliency practice end-to-end by standardizing incident management, improving operational visibility with real-time metrics and dashboards, and using incident data to drive process improvements with product and engineering teams. Lead organization-wide readiness for peak traffic events like Black Friday, collaborate across teams to close resilience gaps, and foster knowledge sharing through communication, documentation, and training.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Commerce Tools
Commerce Tools
1 month ago

Principal Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

Drive commercetools’ resiliency practice end-to-end by standardizing incident management, improving operational visibility with real-time metrics and dashboards, and using incident data to drive process improvements with product and engineering teams. Lead organization-wide readiness for peak traffic events like Black Friday, collaborate across teams to close resilience gaps, and foster knowledge sharing through communication, documentation, and training.
Location: Berlin, London
Workplace: Hybrid
Employment Type: Full time
Job Function: Communications, PR & Community
Seniority: Mid level

Key Responsibilities

  • •Standardize end-to-end incident management processes for detection, response, communication, and postmortems across the company.
  • •Enhance system visibility with real-time metrics, dashboards, and signals to track health and incident trends.
  • •Use operational data to identify process gaps and partner with product engineering teams to fix them.
  • •Own and scale organization-wide readiness programs for massive traffic spikes such as Black Friday.
  • •Lead cross-team and cross-functional initiatives by collaborating with engineering leadership and principal engineers (e.g., Cloud, Security, API, Architecture, Performance) and driving knowledge sharing through documentation and training.

Pay and Benefits

Perks:Health InsuranceLearning BudgetParental LeaveEquity

Key Requirements

  • •7+ years driving incident management and operational excellence, plus 5+ years leading organization-wide resiliency and reliability initiatives.
  • •Demonstrated experience managing and scaling systems for high-stakes, multi-team operational events (e.g., Black Friday, major launches).
  • •Strong data literacy to analyze metrics, diagnose technical issues, and measure process improvements.
  • •Ability to evaluate technical issues alongside organizational and human dynamics, and lead through influence.
  • •Experience managing large-scale initiatives across multiple engineering teams in an Agile environment, with fluent English communication.
Experience:ResiliencyReliabilityIncident managementOperational excellence
Skills:Problem-solvingLeadershipCommunicationMentoringData analysis
Languages:English

Company Brief

Commerce Tools
Provides a frontend-as-a-service platform for headless commerce, enabling teams to build, deploy, and manage composable storefronts. Integrates with commercetools and other backend services to deliver omnichannel customer experiences.
Industry: SaaS
Company Size: Medium (51 to 250 employees)
Growth: Growth Stage Startup
Headquarters: Berlin, Germany
WebsiteLinkedIn