Senior AI Infrastructure Engineer

Heidi Health
Melbourne
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Operational ownership","Performance debugging","System design","Incident response","Cross-functional collaboration"]

Build and operate the AI infrastructure powering model-serving at scale, from checkpoint to production. Own model-serving pipelines, request routing, autoscaling, fallbacks, and controlled rollouts/rollbacks across regions. Manage GPU clusters and scheduling, improve inference performance, and support distributed training. Make deployments observable with dashboards and alerts, drive incident traceability, and reduce time lost to failed or slow jobs while keeping compute costs actionable.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Heidi Health
Heidi Health
2 days ago

Senior AI Infrastructure Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 9 hours agoStatus: Live

Job Summary

Build and operate the AI infrastructure powering model-serving at scale, from checkpoint to production. Own model-serving pipelines, request routing, autoscaling, fallbacks, and controlled rollouts/rollbacks across regions. Manage GPU clusters and scheduling, improve inference performance, and support distributed training. Make deployments observable with dashboards and alerts, drive incident traceability, and reduce time lost to failed or slow jobs while keeping compute costs actionable.
Location: Melbourne
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Build and own model-serving infrastructure from checkpoint to production, including repeatable deployment pipelines, routing, autoscaling, and controlled rollouts/rollbacks across regions.
  • •Manage GPU clusters and workload scheduling to improve resource allocation across inference, training, and evaluation while keeping online services responsive.
  • •Improve inference performance by profiling real workloads and optimizing latency, throughput, and memory efficiency using techniques like batching and cache management.
  • •Support distributed training and model iteration by enabling reliable launch/management of training jobs, artifact/checkpoint handling, and moving validated models into serving.
  • •Make deployments observable and incidents traceable with dashboards/alerts for model latency, queueing, errors, GPU health, and workload performance, and drive production reliability and recovery improvements.

Pay and Benefits

Perks:Learning BudgetWellness StipendHome OfficeParental LeaveFertility SupportRemote WorkEquity

Key Requirements

  • •At least 1 year building and operating infrastructure for large language models, including model deployment, inference serving, or distributed training.
  • •Deployed and maintained LLMs or other demanding ML workloads using engines such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, or comparable systems.
  • •Managed GPU workloads using Kubernetes, Slurm, or an equivalent platform, with experience in scheduling, resource allocation, capacity planning, and failure recovery.
  • •Strong performance debugging skills using traces/metrics/profiling to identify compute, memory, communication, and scheduling bottlenecks affecting latency, throughput, and cost.
  • •Proficient in Python and comfortable with backend/systems development in Go, C++, Rust, or a comparable language, plus practical Linux/container and distributed services experience.
Experience:Large language modelsLLM infrastructureHealthcareDistributed trainingGPU workloads
Skills:Operational ownershipPerformance debuggingSystem designIncident responseCross-functional collaboration
Tech Stack:PythonGoC++RustLinuxContainersVLLMSGLangTensorRT-LLMTriton Inference ServerKubernetesSlurmPyTorch FSDPMegatronCUDATritonNCCLRDMA

Company Brief

Heidi Health
Builds an AI medical scribe and clinical productivity platform that automates documentation, form-filling, and task management to expand clinician capacity across hospitals, GP clinics and specialist services worldwide.
Industry: HealthTech
Company Size: Medium (51 to 250 employees)
Growth: Scaleup
Valuation: USD 250M to 500M
Funding: Series B
Headquarters: Melbourne, Australia
Founded: 2019
Glassdoor
Glassdoor: 4.7
WebsiteLinkedInGlassdoor