Machine Learning Infrastructure Engineer, Model Inference

Abridge
San Francisco
Workplace: RemoteFull timeUSD 221,000 - 260,000 annuallyFunction: DevOps, Cloud & InfrastructureSkills: ["Communication","Problem-solving","Teamwork"]

Lead the design, deployment, and optimization of scalable ML infrastructure to power production-grade models. You’ll manage Kubernetes-based inference/training pipelines, API orchestration, and GPU-accelerated workflows, collaborating with ML and product teams to boost performance, scalability, and reliability of AI-driven solutions.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Abridge
Abridge
1 year ago

Machine Learning Infrastructure Engineer, Model Inference

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 8 hours agoStatus: Live

Job Summary

Lead the design, deployment, and optimization of scalable ML infrastructure to power production-grade models. You’ll manage Kubernetes-based inference/training pipelines, API orchestration, and GPU-accelerated workflows, collaborating with ML and product teams to boost performance, scalability, and reliability of AI-driven solutions.
Location: San Francisco
Workplace: Remote
Employment Type: Full time · Permanent
Job Function: DevOps, Cloud & Infrastructure
Seniority: Manager level

Key Responsibilities

  • •Design, deploy and maintain scalable Kubernetes clusters for AI model inference and training.
  • •Develop, optimize, and maintain ML model serving and training infrastructure, ensuring high-performance and low-latency.
  • •Collaborate with ML and product teams to scale backend infrastructure for AI-driven products, focusing on model deployment, throughout optimization, and compute efficiency.
  • •Optimize compute-heavy workflows and enhance GPU utilization for ML workloads.
  • •Build a robust model API orchestration system and collaborate with leadership to define scaling strategies for long-term efficiency.

Pay and Benefits

Salary: USD 221,000 - 260,000 annually
Equity and Bonus:Equity
Perks:Health Insurance401kEquitySabbatical

Key Requirements

  • •Strong experience in building and deploying machine learning models in production environments.
  • •Deep understanding of container orchestration and distributed systems architecture.
  • •Expertise in Kubernetes administration, including custom resource definitions, operators, and cluster management.
  • •Experience developing APIs and managing distributed systems for both batch and real-time workloads.
  • •Excellent communication skills, with the ability to interface between research and product engineering.
Experience:HealthcareAI
Skills:CommunicationProblem-solvingTeamwork
Tech Stack:KubernetesGPUNVIDIA Triton ServerPyTorchTensorFlowTerraformAnsibleDockerGitOpsCUDALLMASR

Company Brief

Abridge
Builds an enterprise-grade generative-AI platform that captures and summarizes patient–clinician conversations in real time to produce structured clinical documentation, integrated with EMRs to reduce clinician documentation burden and enable clinical and revenue workflows.
Industry: HealthTech
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series E+
Headquarters: San Francisco, United States
Founded: 2018
Glassdoor
Glassdoor: 4.7
WebsiteLinkedInGlassdoor