AI Infrastructure Engineer, Serving Platform
Scale AI
London
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 4+ yearsSkills: ["Problem-solving","Independent work","Collaboration"]Design and build scalable, reliable platforms for serving LLMs on the ML Infrastructure team. Own fault-tolerant, high-performance systems, develop an internal platform for LLM capability discovery, and collaborate with researchers and engineers to integrate and optimize models for production and research. Lead end-to-end projects, run architecture and design reviews, and implement monitoring and observability to maintain system health and performance.

