AI Infrastructure Engineer, Sandbox Platform

Scale AI
London
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 4+ yearsSkills: ["Debugging","Root cause analysis","Cross-functional collaboration","Developer experience focus","Urgency"]

Build and evolve an agent sandboxing platform that securely executes code powering agentic workflows across internal and customer-managed environments. You’ll design secure isolation and reproducible execution, optimize cold-start latency and resource usage, and reduce error rates through monitoring, debugging, and proactive fixes. Partner with internal teams to shape tooling and a sandboxing roadmap, and lead end-to-end architecture reviews from design through deployment.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Scale AI
Scale AI
1 month ago

AI Infrastructure Engineer, Sandbox Platform

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 1 hour agoStatus: Live

Job Summary

Build and evolve an agent sandboxing platform that securely executes code powering agentic workflows across internal and customer-managed environments. You’ll design secure isolation and reproducible execution, optimize cold-start latency and resource usage, and reduce error rates through monitoring, debugging, and proactive fixes. Partner with internal teams to shape tooling and a sandboxing roadmap, and lead end-to-end architecture reviews from design through deployment.
Location: London
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Design and build the sandboxing platform, client library, and API surface for secure code execution across containerized and virtualized environments.
  • •Ensure strong isolation, security, and reproducibility of execution across user sessions and workloads.
  • •Optimize cold-start latency, memory footprint, and resource utilisation at scale.
  • •Drive down error rates through systematic debugging, monitoring, and proactive fixes.
  • •Respond to incidents and production issues with root cause analysis and preventive fixes, while leading roadmap and end-to-end projects.

Key Requirements

  • •4+ years building high-performance systems software, including maintaining libraries, SDKs, or developer-facing APIs.
  • •Deep Linux internals knowledge including process isolation, memory management, cgroups, and namespaces.
  • •Experience with containerisation and virtualisation technologies such as Docker, Firecracker, gVisor, QEMU, and Kata Containers.
  • •Proficiency in a systems programming language such as Go, Rust, or C/C++.
  • •Strong debugging skills and the ability to navigate performance and security tradeoffs in production systems.
Experience:4+ yearsInfrastructureDeveloper tools
Skills:DebuggingRoot cause analysisCross-functional collaborationDeveloper experience focusUrgency
Languages:English
Tech Stack:LinuxProcess isolationVirtualisationContainerisationDockerFirecrackerGVisorQEMUKata ContainersGoRustC/C++APIsSDKsKubernetesCgroupsNamespacesClient library

Company Brief

Scale AI
Provides data labeling, annotation, and infrastructure services to accelerate machine learning and AI development. Supplies high-quality training data, tooling, and APIs for customers in autonomous vehicles, mapping, robotics, and enterprise AI applications.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series E+
Headquarters: San Francisco, United States
Founded: 2016
WebsiteLinkedIn