Staff/Principal DevOps Engineer, AI Inference

Lila Sciences
Cambridge
Workplace: OnsiteFull timeUSD 192,000 - 272,000Function: DevOps, Cloud & InfrastructureSkills: ["Collaboration","Incident management","Performance optimization","Automation","Capacity planning"]

Design, implement, and optimize infrastructure for large-scale, low-latency ML inference. Build Kubernetes-based GPU/accelerator scheduling and multi-tenant serving platforms using frameworks like vLLM and Triton. Develop autoscaling, model deployment pipelines, and robust observability for latency and token throughput. Own AWS EKS/EC2 accelerator infrastructure and cost optimization while collaborating with ML and software teams to deliver reliable production inference.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Lila Sciences
Lila Sciences
1 month ago

Staff/Principal DevOps Engineer, AI Inference

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 3 hours agoStatus: Live

Job Summary

Design, implement, and optimize infrastructure for large-scale, low-latency ML inference. Build Kubernetes-based GPU/accelerator scheduling and multi-tenant serving platforms using frameworks like vLLM and Triton. Develop autoscaling, model deployment pipelines, and robust observability for latency and token throughput. Own AWS EKS/EC2 accelerator infrastructure and cost optimization while collaborating with ML and software teams to deliver reliable production inference.
Location: Cambridge
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Sr. Manager level

Key Responsibilities

  • •Design and optimize GPU/accelerator infrastructure for low-latency, high-throughput ML inference at scale.
  • •Build and operate model serving platforms with batching, caching, and request routing across heterogeneous accelerator fleets.
  • •Implement autoscaling and deployment pipelines for ML models, including canary rollouts, A/B testing, versioning, and safe rollback across regions.
  • •Develop infrastructure-as-code using Terraform and Helm for GPU-accelerated EKS clusters and CI/CD for model artifacts.
  • •Improve observability and performance using GPU utilization monitoring, inference latency profiling, token throughput dashboards, and SLO/SLI tracking.

Pay and Benefits

Salary: USD 192,000 - 272,000
Perks:Health InsuranceDentalVisionLife InsuranceDisability InsurancePaid LeaveParental LeaveEducation AssistanceCommuter BenefitsMeal Allowance

Key Requirements

  • •Significant experience operating GPU/accelerator infrastructure at scale in DevOps/SRE/Platform Engineering contexts.
  • •Deep Kubernetes experience for ML workloads, including GPU scheduling, resource quotas, node affinity, and accelerator device management.
  • •Proficiency deploying to AWS with infrastructure-as-code (Terraform, Helm) and managing GPU-based compute (EKS, EC2 P-series/Inf/Trn instances).
  • •Experience with model serving infrastructure, including inference servers and techniques like request batching and KV-cache optimization.
  • •Strong understanding of networking for distributed inference, including NCCL and L4/L7 load balancing, plus proficiency in Python for automation and tooling.
Experience:Machine learningGPUAI inferenceSREPlatform engineeringCloudDistributed systems
Skills:CollaborationIncident managementPerformance optimizationAutomationCapacity planning
Languages:En
Tech Stack:KubernetesVLLMTriton Inference ServerTGITerraformHelmAWSEKSEC2S3EFACUDAIAMPythonRustGoNCCLVPCPrivateLinkL4

Company Brief

Lila Sciences
Develops AI-driven platforms to accelerate drug discovery and biological research by integrating machine learning with chemical and biological data to predict molecular properties, streamline candidate selection, and enable faster therapeutic development.
Industry: Biotech
Company Size: Small (11 to 50 employees)
Growth: Early Stage Startup
Funding: Seed
WebsiteLinkedIn