Lead Site Reliability Engineer

Heidi Health
Melbourne
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureExperience: 7+ yearsSkills: ["Leadership","Incident response leadership","Clear communication","Debugging under pressure","Collaboration"]

Lead an SRE team in the core Platform/SRE group that owns production at Heidi. Stay hands-on with incident response and on-call, drive improvements to operational reliability, observability, and toil reduction, and operate Kubernetes clusters and cloud infrastructure. Partner with engineering and product to improve production readiness, deployments, and change safety, while hiring and scaling the team as Heidi grows.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Heidi Health
Heidi Health
1 week ago

Lead Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

Lead an SRE team in the core Platform/SRE group that owns production at Heidi. Stay hands-on with incident response and on-call, drive improvements to operational reliability, observability, and toil reduction, and operate Kubernetes clusters and cloud infrastructure. Partner with engineering and product to improve production readiness, deployments, and change safety, while hiring and scaling the team as Heidi grows.
Location: Melbourne
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Manager level

Key Responsibilities

  • •Participate in on-call and lead incidents end-to-end, including outside the immediate team, ensuring clear communication during production issues.
  • •Improve operational reliability by identifying recurring issues and reliability risks and driving fixes via alerting, automation, system changes, and process improvements.
  • •Own the production environment by operating and improving Kubernetes clusters, cloud infrastructure, and core platform services.
  • •Strengthen observability by building dashboards, alerts, logs, and traces to surface issues earlier and reduce diagnosis time.
  • •Lead and grow the SRE team by managing daily operations, hiring and onboarding as the company scales, and setting on-call structure and team norms.

Pay and Benefits

Equity and Bonus:Equity
Perks:Learning BudgetWellness StipendHome OfficeParental LeaveFertility SupportRemote WorkEquity

Key Requirements

  • •7+ years in SRE, DevOps, platform, or operations-heavy engineering roles, including experience formally leading or managing a team.
  • •A track record of hiring, coaching, and growing engineers.
  • •Deep experience supporting production systems with on-call rotations and credibility to jump into incidents.
  • •Strong experience operating cloud infrastructure at scale (AWS preferred).
  • •Hands-on experience with Kubernetes/containerized workloads, infrastructure as code, monitoring/alerting, and defining/owning SLOs.
Experience:7+ yearsSREDevOpsPlatformProduction operationsKubernetesCloud infrastructureMonitoring & observability
Skills:LeadershipIncident response leadershipClear communicationDebugging under pressureCollaboration
Tech Stack:AWSKubernetesTerraformDatadogPrometheusPythonBashLogsTracesDashboardsAlertingSLOsError budgetsCapacity planningAutomationRunbooksBlameless post-mortems

Company Brief

Heidi Health
Builds an AI medical scribe and clinical productivity platform that automates documentation, form-filling, and task management to expand clinician capacity across hospitals, GP clinics and specialist services worldwide.
Industry: HealthTech
Company Size: Medium (51 to 250 employees)
Growth: Scaleup
Valuation: USD 250M to 500M
Funding: Series B
Headquarters: Melbourne, Australia
Founded: 2019
Glassdoor
Glassdoor: 4.7
WebsiteLinkedInGlassdoor