Staff/Principal DevOps Engineer, AI Inference
Lila Sciences
Cambridge
Workplace: OnsiteFull timeUSD 192,000 - 272,000Function: DevOps, Cloud & InfrastructureSkills: ["Collaboration","Incident management","Performance optimization","Automation","Capacity planning"]Design, implement, and optimize infrastructure for large-scale, low-latency ML inference. Build Kubernetes-based GPU/accelerator scheduling and multi-tenant serving platforms using frameworks like vLLM and Triton. Develop autoscaling, model deployment pipelines, and robust observability for latency and token throughput. Own AWS EKS/EC2 accelerator infrastructure and cost optimization while collaborating with ML and software teams to deliver reliable production inference.

