AI Infrastructure Engineer, Serving Platform

Scale AI
London
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 4+ yearsSkills: ["Problem-solving","Independent work","Collaboration"]

Design and build scalable, reliable platforms for serving LLMs on the ML Infrastructure team. Own fault-tolerant, high-performance systems, develop an internal platform for LLM capability discovery, and collaborate with researchers and engineers to integrate and optimize models for production and research. Lead end-to-end projects, run architecture and design reviews, and implement monitoring and observability to maintain system health and performance.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Scale AI
Scale AI
1 month ago

AI Infrastructure Engineer, Serving Platform

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 1 hour agoStatus: Live

Job Summary

Design and build scalable, reliable platforms for serving LLMs on the ML Infrastructure team. Own fault-tolerant, high-performance systems, develop an internal platform for LLM capability discovery, and collaborate with researchers and engineers to integrate and optimize models for production and research. Lead end-to-end projects, run architecture and design reviews, and implement monitoring and observability to maintain system health and performance.
Location: London
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Build and maintain fault-tolerant, high-performance systems for serving LLMs and other models at scale.
  • •Build an internal platform to enable LLM capability discovery.
  • •Collaborate with researchers and engineers to integrate and optimize models for production and research use cases.
  • •Conduct architecture and design reviews to uphold best practices for scalability and system design.
  • •Develop monitoring and observability solutions and lead cross-functional projects end-to-end.

Key Requirements

  • •4+ years building large-scale, high-performance backend systems.
  • •Strong programming skills in one or more languages such as Python, Go, Rust, or C++.
  • •Experience with LLM serving and routing fundamentals (e.g., rate limiting, token streaming, load balancing, budgets).
  • •Experience with LLM capabilities and concepts such as reasoning, tool calling, and prompt templates.
  • •Experience with containers and orchestration tools such as Docker and Kubernetes.
Experience:4+ yearsLLM servingBackend systemsML infrastructureContainersCloud infrastructure
Skills:Problem-solvingIndependent workCollaboration
Languages:English
Tech Stack:PythonGoRustC++DockerKubernetesAWSGCPTerraformVLLMSGLangTensorRT-LLMText-generation-inferenceLLM servingToken streamingRate limitingLoad balancing

Company Brief

Scale AI
Provides data labeling, annotation, and infrastructure services to accelerate machine learning and AI development. Supplies high-quality training data, tooling, and APIs for customers in autonomous vehicles, mapping, robotics, and enterprise AI applications.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series E+
Headquarters: San Francisco, United States
Founded: 2016
WebsiteLinkedIn