Site Reliability Engineer

Baseten
San Francisco, New York, United States, Toronto, Canada, Montreal
Workplace: HybridFull time135,000 - 285,000 annuallyFunction: DevOps, Cloud & InfrastructureSkills: ["Kubernetes","Terraform","Helm","GitOps","Incident management","Grafana","Prometheus","Loki","ELK","VictoriaMetrics","Observability","Runbooks","Automation","Alarms","Post-mortems"]

Site Reliability Engineer to own Baseten’s multi-cloud Kubernetes infrastructure, build observability tooling, runbooks, and automated mitigations; collaborate with engineering, forward-deployed and product teams to improve SRE practices, incident response, and platform reliability at scale.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Baseten
Baseten
10 months ago

Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 10 hours agoStatus: Live

Job Summary

Site Reliability Engineer to own Baseten’s multi-cloud Kubernetes infrastructure, build observability tooling, runbooks, and automated mitigations; collaborate with engineering, forward-deployed and product teams to improve SRE practices, incident response, and platform reliability at scale.
Location: San Francisco, New York, United States, Toronto, Canada, Montreal
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Own the reliability of Baseten's multi-cloud Kubernetes infrastructure, including incident response, post-mortems, and remediation tracking
  • •Build and maintain observability infrastructure — metrics, logging, dashboards, and alerting — as code
  • •Author, validate, and improve runbooks for recurring failure patterns, ensuring they're structured for low-context, safe execution
  • •Identify high-frequency failure patterns and convert them into automated mitigations or self-healing automations
  • •Diagnose and resolve runtime issues related to latency, memory behavior, GPU utilization, concurrency, and model lifecycle management

Pay and Benefits

Salary: 135,000 - 285,000 annually
Equity and Bonus:Equity
Perks:Health InsuranceEquity401kRemote WorkPaid Parental

Key Requirements

  • •Extensive hands-on experience with Kubernetes (multi-cloud experience across EKS, GKE, or similar is a strong plus)
  • •Experience in building and maintaining scalable infrastructure
  • •Strong foundation in observability tooling: metrics, logging, dashboards, and alerting pipelines (observability-as-code is a plus)
  • •Experience with infrastructure-as-code (Terraform, Helm) and GitOps workflows (Flux CD, ArgoCD)
  • •Experience writing and improving runbooks, leading incident response, and doing post-mortem analysis
Experience:CloudKubernetesSREInfrastructureMulti-cloud
Skills:KubernetesTerraformHelmGitOpsIncident managementGrafanaPrometheusLokiELKVictoriaMetricsObservabilityRunbooksAutomationAlarmsPost-mortems
Languages:English
Tech Stack:KubernetesEKSGKEPrometheusGrafanaLokiELKVictoriaMetricsTerraformHelmFlux CDArgoCDIncident.io

Company Brief

Baseten
Baseten provides an inference-first ML infrastructure platform that lets engineering and ML teams deploy, serve, and scale machine-learning models with optimized performance, autoscaling, and GPU-backed hosting for production AI applications. ([crunchbase.com](https://www.crunchbase.com/organization/baseten?utm_source=openai))
Industry: AI & Machine Learning
Company Size: Medium (51 to 250 employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series E+
Headquarters: San Francisco, United States
Founded: 2019
Glassdoor
Glassdoor: 5.0
WebsiteLinkedInGlassdoor