Staff+ Infrastructure Engineer, Cluster Infrastructure

Anthropic
London
Workplace: OnsiteFull timeGBP 325,000 - 485,000 annuallyFunction: DevOps, Cloud & InfrastructureEducation: bachelorsSkills: ["Communication","Leadership","Problem-solving","Mentorship","Collaboration"]

Lead the technical direction for building and operating large-scale compute clusters across cloud providers and on-prem datacenters. Own provisioning, updates, and lifecycle management for cluster infrastructure, drive scalability and fault-tolerance, partner with security and cloud teams, and collaborate with research/product teams to shape long-term compute strategy. Mentorship of engineers and establishment of operational excellence are key.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Anthropic
Anthropic
2 months ago

Staff+ Infrastructure Engineer, Cluster Infrastructure

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 6 hours agoStatus: Live

Job Summary

Lead the technical direction for building and operating large-scale compute clusters across cloud providers and on-prem datacenters. Own provisioning, updates, and lifecycle management for cluster infrastructure, drive scalability and fault-tolerance, partner with security and cloud teams, and collaborate with research/product teams to shape long-term compute strategy. Mentorship of engineers and establishment of operational excellence are key.
Location: London
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Own the technical strategy and roadmap for agent-driven cluster lifecycle management - provisioning, updates and decommissioning
  • •Partner across teams to ensure new compute capacity is ingested on time
  • •Align with partner teams on physical build-out and leverage cloud solutions to deliver high-bandwidth inter-cluster connectivity
  • •Collaborate with security owners to ensure clusters are provisioned secure-by-default
  • •Define and drive strategy on cluster scalability, homogeneity and fault tolerance

Pay and Benefits

Salary: GBP 325,000 - 485,000 annually

Key Requirements

  • •Deep expertise in distributed systems, reliability, and cloud platforms (e.g., Kubernetes, IaC, AWS/GCP/Azure)
  • •Strong proficiency in at least one systems language (e.g., Rust, Go, or Python), IaC proficiency with Terraform
  • •Track record of leading complex, multi-quarter technical initiatives spanning multiple teams or systems
  • •Ability to build alignment across senior stakeholders and communicate effectively at all levels
Experience:Cloud computingDistributed systems
Education:Bachelor's
Skills:CommunicationLeadershipProblem-solvingMentorshipCollaboration
Languages:English
Tech Stack:KubernetesTerraformAWSGCPAzureRustGoPythonIaCMesosBorg-likeIstioEnvoyLinkerdCiliumEBPFNetworkPolicyTemporalArgo Workflows

Eligibility

Work Authorization:Sponsorship available.

Company Brief

Anthropic
Develops large-scale AI systems and safety research to create reliable, steerable, and interpretable AI assistants and models for commercial and research applications.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series C
Headquarters: San Francisco, United States
Founded: 2021
WebsiteLinkedIn