Senior Staff+ Software Engineer, Node Infra

Anthropic
London
Workplace: OnsiteFull timeGBP 325,000 - 485,000 annuallyFunction: Software EngineeringEducation: bachelorsSkills: ["Kubernetes","Terraform","AWS","GCP","Azure","Python","Go","Rust","GPU","TPU","Trainium","IaC"]

Lead the Node Infra team to manage the full lifecycle of accelerator capacity, designing scalable infrastructure across clouds and hardware accelerators (GPUs/TPUs/Trainium). Own the roadmap for node ingestion, health, and automated repair, driving cross-team initiatives and optimizing fleet reliability for large-scale AI research.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Anthropic
Anthropic
2 months ago

Senior Staff+ Software Engineer, Node Infra

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 6 hours agoStatus: Live

Job Summary

Lead the Node Infra team to manage the full lifecycle of accelerator capacity, designing scalable infrastructure across clouds and hardware accelerators (GPUs/TPUs/Trainium). Own the roadmap for node ingestion, health, and automated repair, driving cross-team initiatives and optimizing fleet reliability for large-scale AI research.
Location: London
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Own the technical strategy and roadmap for node lifecycle management - ingestion, bring-up, health checking, and automated repair
  • •Drive cross-team initiatives to build and scale AI clusters across multiple clouds and accelerator families
  • •Design and operate the systems that detect, isolate, and remediate unhealthy hardware automatically, driving up fleet MTBI and minimizing stranded capacity
  • •Define infrastructure architecture, ensuring the hardest problems get solved - whether by you directly or by working through others
  • •Work closely with cloud providers and internal research/inference/product teams to shape long-term compute, data, and infrastructure strategy

Pay and Benefits

Salary: GBP 325,000 - 485,000 annually

Key Requirements

  • •Deep expertise in distributed systems, reliability, and cloud platforms (e.g., Kubernetes, IaC, AWS/GCP/Azure)
  • •Strong proficiency in at least one systems language (e.g., Rust, Go, or Python), IaC proficiency with Terraform
  • •Hands-on experience with machine learning accelerators (GPUs, TPUs, or Trainium)
  • •Track record of leading complex, multi-quarter technical initiatives that span multiple teams or systems
  • •Ability to build alignment across senior stakeholders and communicate effectively at all levels
Experience:Cloud computingDistributed systemsInfrastructureAI infrastructureHyperscale compute
Education:Bachelor's
Skills:KubernetesTerraformAWSGCPAzurePythonGoRustGPUTPUTrainiumIaC
Languages:English
Tech Stack:KubernetesTerraformAWSGCPAzurePythonGoRustGPUsTPUsTrainiumIaC

Company Brief

Anthropic
Develops large-scale AI systems and safety research to create reliable, steerable, and interpretable AI assistants and models for commercial and research applications.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series C
Headquarters: San Francisco, United States
Founded: 2021
WebsiteLinkedIn