Site Reliability Engineer

Adaptyv
Switzerland
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Systems thinking","Calm under pressure","Root-cause analysis","Incident response","Automation mindset"]

Own reliability for Adaptyv’s software-driven automated lab, ensuring APIs, data pipelines, job queues, and integrations keep running for customer experiments. Build end-to-end observability (metrics, logs, traces, dashboards), define SLOs and alerting, harden pipelines, and lead incident triage and blameless postmortems. Improve deploy safety, automate runbooks to eliminate toil, and partner with software and lab-automation teams so reliability is built into the system.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Adaptyv
Adaptyv
1 month ago

Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 19 hours agoStatus: Live

Job Summary

Own reliability for Adaptyv’s software-driven automated lab, ensuring APIs, data pipelines, job queues, and integrations keep running for customer experiments. Build end-to-end observability (metrics, logs, traces, dashboards), define SLOs and alerting, harden pipelines, and lead incident triage and blameless postmortems. Improve deploy safety, automate runbooks to eliminate toil, and partner with software and lab-automation teams so reliability is built into the system.
Location: Switzerland
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Own reliability across the full stack supporting LabOS and the customer-facing platform.
  • •Build observability across services and instrument SLOs with effective alerting to page only on real issues.
  • •Harden data and processing pipelines to prevent silent corruption or stalled experiments.
  • •Own incident response: triage outages, mitigate quickly, and run blameless postmortems for permanent fixes.
  • •Improve deploy safety and rollback and automate runbooks to reduce toil and firefighting.
  • •Partner with software and lab-automation teams to make reliability a system property.

Key Requirements

  • •Write production code in Python and/or TypeScript while owning infrastructure.
  • •Demonstrate real SRE/production ownership experience including running services and carrying a pager.
  • •Build observability (metrics, logging, tracing, dashboards, alerting) using tools like Grafana/Prometheus/Loki.
  • •Lead incident response with strong incident instinct: triage, root-cause quickly, prevent repeat outages.
  • •Automate recurring operational tasks and improve reliability practices rather than relying on manual runbooks.
Experience:SaaSProduction ownershipObservabilityIncident responseAutomation
Skills:Systems thinkingCalm under pressureRoot-cause analysisIncident responseAutomation mindset
Tech Stack:PythonTypeScriptGrafanaPrometheusLokiVercelSupabaseModalEdge functionsAPIsDatabasesJob queuesClaude Code

Company Brief

Adaptyv
Biotechnology company developing adaptive immunotherapy platforms and precision biologic treatments aimed at enhancing immune responses against disease. Focuses on engineering biologic solutions and translational research to accelerate therapeutic candidates from discovery to clinical development.
Industry: Biotech
Website