Senior Site Reliability Engineer

Runware
United Kingdom, France, Germany, Spain, Italy, Sweden
Workplace: RemoteFull timeFunction: DevOps, Cloud & InfrastructureEducation: bachelorsSkills: ["Ownership","Troubleshooting","Incident management","Collaboration","Automation-minded mindset"]

Own and improve the reliability, performance, and availability of critical production services for a serverless AI inference platform. Define SRE practices (SLIs, SLOs, alerting, observability, and production-readiness), investigate complex distributed-system incidents, and lead incident reviews/RCAs. Reduce operational toil via automation and safer deployments while partnering with Engineering and DevOps on capacity planning, scaling, and architectural improvements.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Runware
Runware
1 month ago

Senior Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 7 hours agoStatus: Live

Job Summary

Own and improve the reliability, performance, and availability of critical production services for a serverless AI inference platform. Define SRE practices (SLIs, SLOs, alerting, observability, and production-readiness), investigate complex distributed-system incidents, and lead incident reviews/RCAs. Reduce operational toil via automation and safer deployments while partnering with Engineering and DevOps on capacity planning, scaling, and architectural improvements.
Location: United Kingdom, France, Germany, Spain, Italy, Sweden
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Own and improve reliability, availability, and performance of critical production services across the platform.
  • •Define and evolve reliability practices (SLIs, SLOs, alerting, observability, and production-readiness standards).
  • •Investigate complex production issues across distributed systems, APIs, networking, queues, databases, and GPU-backed workloads.
  • •Lead and contribute to incident reviews and RCAs, converting recurring failure modes into lasting improvements.
  • •Reduce operational toil via automation, automated remediation, and improvements to deployment safety, recovery, and resilience.

Pay and Benefits

Perks:Paid LeaveEquityRemote WorkFlexible HoursParental LeaveCompany Retreats

Key Requirements

  • •Strong experience operating and troubleshooting production systems at scale in an SRE, Production Engineering, Platform Engineering, or similar role.
  • •Comfort debugging distributed systems across applications, databases, queues, containers, networking, and infrastructure.
  • •Experience designing and operating observability systems using metrics, logs, and distributed tracing.
  • •Knowledge of SRE principles including SLIs, SLOs, error budgets, capacity planning, incident management, and reducing operational toil.
  • •Experience with Kubernetes, containers, IaC, automated deployment practices, and the ability to write automation/software (e.g., Python, Go, PHP).
Experience:SREPlatform engineeringDistributed systemsObservabilityProduction operations
Education:Bachelor's
Skills:OwnershipTroubleshootingIncident managementCollaborationAutomation-minded mindset
Tech Stack:KubernetesContainersIaCPythonGoPHPMetricsLogsDistributed tracingSLIsSLOsAlertingObservabilityMySQLRedisClickHouseRabbitMQCDNLoad balancingGPU

Company Brief

Runware
Runware.ai offers a platform for building, deploying, and managing LLM-powered applications, focusing on orchestration, observability, and operational tooling to ensure reliable, scalable, and safe production AI workflows.
Industry: Developer Tools
Website