Senior Site Reliability Engineer

Garner Health
Anywhere
Workplace: RemoteFull timeUSD 191,000 - 226,000 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 4+ yearsSkills: ["Incident response","Root cause analysis","Automation-first mindset","Clear communication","Accountability"]

Own the reliability, performance, and resilience of cloud infrastructure powering Garner’s products and AI/ML workloads. As part of Platform Engineering, define and uphold SLOs, lead incident response and root-cause analysis, and build observability, monitoring, and alerting. Drive automation and standards with Terraform-based infrastructure-as-code, optimize cost and performance, and ensure security and HIPAA compliance while enabling teams to ship faster.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Garner Health
Garner Health
2 days ago

Senior Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 11 hours agoStatus: Live

Job Summary

Own the reliability, performance, and resilience of cloud infrastructure powering Garner’s products and AI/ML workloads. As part of Platform Engineering, define and uphold SLOs, lead incident response and root-cause analysis, and build observability, monitoring, and alerting. Drive automation and standards with Terraform-based infrastructure-as-code, optimize cost and performance, and ensure security and HIPAA compliance while enabling teams to ship faster.
Location: Anywhere
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Own end-to-end reliability, performance, and resilience of Garner’s cloud environments (AWS, Kubernetes), including AI/ML workloads; define, measure, and uphold SLOs.
  • •Lead incident response via on-call rotation, drive root-cause analysis, ensure corrective actions are completed, and review infrastructure changes with rigor.
  • •Build and maintain monitoring, alerting, and observability systems to detect and resolve issues before users are impacted.
  • •Create automated, composable infrastructure-as-code deliverables with Terraform, and identify cost-efficiency and performance improvements across the stack.
  • •Automate away operational toil using AI tools, enable engineering with deployment and observability standards, and uphold security and HIPAA compliance obligations.
Travel: Low travel

Pay and Benefits

Salary: USD 191,000 - 226,000 annually
Equity and Bonus:Equity
Perks:Paid LeaveHealth InsuranceDentalVision401kEquity

Key Requirements

  • •4+ years operating production cloud infrastructure at scale in an SRE, DevOps, or platform engineering role.
  • •Deep expertise with Kubernetes and Terraform in a cloud-first environment (AWS preferred).
  • •Experience with production observability: defining SLOs, building monitoring/alerting, and leading incident response and blameless post-incident reviews.
  • •Strong software engineering fundamentals in Python or Go for infrastructure automation (Kubernetes API experience is a plus).
  • •Experience optimizing cloud cost/performance and operating in security-conscious or regulated environments (HIPAA, SOC 2 is a plus).
Experience:4+ yearsCloud infrastructureSREDevOpsPlatform engineeringObservabilityInfrastructure-as-codeAI/MLHealthcare
Skills:Incident responseRoot cause analysisAutomation-first mindsetClear communicationAccountability
Languages:English
Tech Stack:AWSKubernetesTerraformIstioPythonGoTypeScriptPostgresNATSDatadogGitLab

Eligibility

Work Authorization:Authorization required. Sponsorship not provided.

Company Brief

Garner Health
Provides employer-sponsored healthcare navigation and care coordination services, combining personalized digital coaching, benefits navigation, and care management to help employees access appropriate care and reduce healthcare costs.
Industry: HealthTech
Website