Engineering Manager, Inference Infrastructure

Anthropic
San Francisco, New York, Seattle
Workplace: OnsiteFull timeUSD 405,000 - 625,000 annuallyFunction: Software EngineeringEducation: bachelorsSkills: ["Technical leadership","Architectural decision-making","Quantitative modeling","Incident response","Cross-team collaboration"]

Lead a team that builds the control plane coordinating Anthropic’s inference fleet—deciding request placement and capacity to meet throughput, reliability, and latency goals. Own the roadmap from caching and protocols to demand-reactive scaling, partner with inference and performance teams on measurable efficiency wins, and run operational excellence through on-call, incident response, and postmortems. You’ll also shape team structure as the system evolves across clouds, hardware, and serving surfaces.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Anthropic
Anthropic
3 days ago

Engineering Manager, Inference Infrastructure

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 13 hours agoStatus: Live

Job Summary

Lead a team that builds the control plane coordinating Anthropic’s inference fleet—deciding request placement and capacity to meet throughput, reliability, and latency goals. Own the roadmap from caching and protocols to demand-reactive scaling, partner with inference and performance teams on measurable efficiency wins, and run operational excellence through on-call, incident response, and postmortems. You’ll also shape team structure as the system evolves across clouds, hardware, and serving surfaces.
Location: San Francisco, New York, Seattle
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Manager level

Key Responsibilities

  • •Own the technical roadmap for coordinating the inference fleet, including traffic placement, capacity allocation, cache placement, demand reaction, and control-plane/inference-engine protocols.
  • •Partner with product, inference engine, performance, and capacity teams to identify throughput, latency, utilization, and cost improvements and ship measurable results.
  • •Build the team’s quantitative modeling practice to define and validate outcomes before shipping changes.
  • •Set technical strategy for evolving the control plane across heterogeneous hardware, multiple cloud providers, and serving surfaces.
  • •Run operational backbone: on-call rotations, incident response, postmortem reviews, and deploy safety to keep the system reliable as teams ship aggressively.

Pay and Benefits

Salary: USD 405,000 - 625,000 annually
Perks:Paid LeaveParental Leave

Key Requirements

  • •Engineering management experience leading teams on critical-path production infrastructure at scale.
  • •Deep systems background (e.g., load balancing, scheduling, cluster orchestration, autoscaling, cache-coherent distributed state, high-performance networking) to make architectural calls.
  • •Experience shipping performance/efficiency improvements in large-scale systems and explaining impact with measurable numbers, including cost.
  • •Experience running production infrastructure with operational stakes (on-call, incident response, capacity events, deploy discipline).
  • •Results-oriented, impact-driven approach balancing throughput, latency, cost, stability, and launch timelines.
Experience:ML infrastructureDistributed systemsInference serving
Education:Bachelor's
Skills:Technical leadershipArchitectural decision-makingQuantitative modelingIncident responseCross-team collaboration
Languages:English
Tech Stack:Load balancingSchedulingCluster orchestrationAutoscalingCache-coherent distributed stateHigh-performance networkingKV cachingContinuous batchingRequest schedulingPrefill/decode disaggregationKubernetes internalsService meshesLoad balancersCluster schedulersAutoscalers

Company Brief

Anthropic
Develops large-scale AI systems and safety research to create reliable, steerable, and interpretable AI assistants and models for commercial and research applications.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series C
Headquarters: San Francisco, United States
Founded: 2021
WebsiteLinkedIn