Staff Site Reliability Engineer

Garner Health
New York
Workplace: RemoteFull timeUSD 241,000 - 270,000 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 7+ yearsSkills: ["Communication","Mentorship","Technical direction","Incident analysis","Accountability"]

Own the reliability strategy for Garner’s cloud infrastructure and AI/ML workloads on the Platform Engineering team. Define and architect SLOs, incident response, observability, and automation standards, using AWS, Kubernetes, and Terraform to drive resilience, performance, and cost efficiency. Lead complex incident escalations, mentor engineers, and ensure security and HIPAA compliance, translating ambiguous scaling needs into reliable, composable infrastructure.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Garner Health
Garner Health
2 days ago

Staff Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 14 hours agoStatus: Live

Job Summary

Own the reliability strategy for Garner’s cloud infrastructure and AI/ML workloads on the Platform Engineering team. Define and architect SLOs, incident response, observability, and automation standards, using AWS, Kubernetes, and Terraform to drive resilience, performance, and cost efficiency. Lead complex incident escalations, mentor engineers, and ensure security and HIPAA compliance, translating ambiguous scaling needs into reliable, composable infrastructure.
Location: New York
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Architect and own end-to-end reliability, performance, and resilience for Garner’s cloud environments, including AI/ML workloads.
  • •Set the standard for incident response: on-call leadership, complex escalation handling, root-cause deep dives, and blameless review culture.
  • •Architect monitoring, alerting, and observability so issues are detected and resolved before users are impacted, and product health is quickly visible.
  • •Translate ambiguous scaling and reliability requirements into automated, composable infrastructure-as-code deliverables (Terraform), improving cost efficiency and performance.
  • •Build and own deployment and observability standards to help engineers ship AI features faster and mentor engineers across the organization.
Travel: Low travel

Pay and Benefits

Salary: USD 241,000 - 270,000 annually
Equity and Bonus:Equity
Perks:Paid LeaveHealth InsuranceDentalVision401k

Key Requirements

  • •7+ years operating production cloud infrastructure at scale in an SRE, DevOps, or platform engineering role.
  • •Deep expertise with Kubernetes and Terraform in a cloud-first environment (AWS preferred), including reliability architecture at scale.
  • •Experience designing reliability practices such as SLO frameworks, observability platforms, incident response programs, and blameless post-incident reviews.
  • •Strong Python or Go skills applied to infrastructure automation (Kubernetes API experience a plus).
  • •Track record driving cloud cost-efficiency and performance optimization across compute, storage, and networking.
Experience:7+ yearsSaaSAI/MLHealthcareCloud infrastructure
Skills:CommunicationMentorshipTechnical directionIncident analysisAccountability
Languages:English
Tech Stack:AWSKubernetesTerraformIstioPythonGoTypeScriptPostgresNATSDatadogGitLabInfrastructure as codeSLOsObservabilityIncident responseHIPAA complianceAWS cloud

Company Brief

Garner Health
Provides employer-sponsored healthcare navigation and care coordination services, combining personalized digital coaching, benefits navigation, and care management to help employees access appropriate care and reduce healthcare costs.
Industry: HealthTech
Website