Senior Site Reliability Engineer
Runware
United Kingdom, France, Germany, Spain, Italy, Sweden
Workplace: RemoteFull timeFunction: DevOps, Cloud & InfrastructureEducation: bachelorsSkills: ["Ownership","Troubleshooting","Incident management","Collaboration","Automation-minded mindset"]Own and improve the reliability, performance, and availability of critical production services for a serverless AI inference platform. Define SRE practices (SLIs, SLOs, alerting, observability, and production-readiness), investigate complex distributed-system incidents, and lead incident reviews/RCAs. Reduce operational toil via automation and safer deployments while partnering with Engineering and DevOps on capacity planning, scaling, and architectural improvements.

