Staff+ Software Engineer, Inference Velocity

Anthropic
San Francisco, Seattle, New York
Workplace: OnsiteFull timeUSD 405,000 - 485,000 annuallyFunction: Software EngineeringEducation: bachelorsSkills: ["Rust","Python","CUDA","TPU","Trainium","XLA","Triton","Neural","Kubernetes","CI/CD"]

Lead the technical direction of Anthropic’s Inference Runtime, owning the shared accelerator-agnostic core of the inference serving stack. Drive architecture, release/validation systems, and cross-org collaboration across GPU/TPU/Trainium platforms while mentoring engineers and ensuring scalable, high-performance throughput for millions of users.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Anthropic
Anthropic
3 months ago

Staff+ Software Engineer, Inference Velocity

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 3 hours agoStatus: Live

Job Summary

Lead the technical direction of Anthropic’s Inference Runtime, owning the shared accelerator-agnostic core of the inference serving stack. Drive architecture, release/validation systems, and cross-org collaboration across GPU/TPU/Trainium platforms while mentoring engineers and ensuring scalable, high-performance throughput for millions of users.
Location: San Francisco, Seattle, New York
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Set technical direction for the team, owning the architecture and roadmap for the shared runtime of the inference serving stack.
  • •Own and evolve the accelerator-agnostic runtime itself – its interfaces, internal boundaries, and build structure – including hands-on work in a performance-sensitive Rust and Python codebase.
  • •Keep the platform's expansion cost low by ensuring new models and deployment targets pay only for their own specialization, and edge cases stitch back into the core easily.
  • •Drive efficient accelerator usage – utilization, scheduling, memory management – across GPU, TPU, and Trainium.
  • •Build the runtime's validation surface around partitioned builds, change-scoped testing, and canary/shadow/rollback as first-class mechanisms.

Pay and Benefits

Salary: USD 405,000 - 485,000 annually
Equity and Bonus:Equity
Perks:Paid LeaveParental LeaveRemote WorkEquity

Key Requirements

  • •Deep background in systems engineering or ML infrastructure with hands-on performance profiling, latency and throughput optimization, and systems debugging at scale.
  • •Real depth in at least one accelerator ecosystem (CUDA/GPU, TPU, or Trainium/AWS Neuron) and ability to keep the runtime agnostic across all of them.
  • •Significant software engineering experience with a strong background in high-performance, large-scale distributed systems serving millions of users.
  • •Track record of defining and using engineering metrics to drive improvement (SLOs, escape rates, release times, latency, throughput).
  • •Experience driving technical alignment across organizational boundaries, advocating for your team’s needs while contributing to shared infrastructure.
Experience:AIInferenceML infrastructureAcceleratorDistributed systems
Education:Bachelor's
Skills:RustPythonCUDATPUTrainiumXLATritonNeuralKubernetesCI/CD
Languages:English
Tech Stack:RustPythonCUDATPUTrainiumXLATritonNeuralKubernetesCI/CD

Eligibility

Work Authorization:Sponsorship available.

Company Brief

Anthropic
Develops large-scale AI systems and safety research to create reliable, steerable, and interpretable AI assistants and models for commercial and research applications.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series C
Headquarters: San Francisco, United States
Founded: 2021
WebsiteLinkedIn