Performance Engineer, Inference Engine

Anthropic
San Francisco, New York
Workplace: OnsiteFull timeUSD 350,000 - 850,000 annuallyFunction: Software EngineeringEducation: bachelorsSkills: ["Performance analysis","Hypothesis testing","Rapid learning","Collaboration","Low ego"]

Build and optimize Anthropic’s inference engine—the software layer that manages token paths between accelerator kernels and routing. Improve throughput, cost, reliability, and latency across accelerator and cloud platforms by modeling performance constraints (compute, memory, interconnect) and iterating with observability-driven experiments. Work on high-performance systems spanning host/device coordination and distributed inference at scale, with a plus for transformer familiarity.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Anthropic
Anthropic
5 days ago

Performance Engineer, Inference Engine

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 17 hours agoStatus: Live

Job Summary

Build and optimize Anthropic’s inference engine—the software layer that manages token paths between accelerator kernels and routing. Improve throughput, cost, reliability, and latency across accelerator and cloud platforms by modeling performance constraints (compute, memory, interconnect) and iterating with observability-driven experiments. Work on high-performance systems spanning host/device coordination and distributed inference at scale, with a plus for transformer familiarity.
Location: San Francisco, New York
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Build and optimize the inference engine that manages token batching, model layout across chips, and forward-pass coordination.
  • •Improve throughput, cost, reliability, and latency across accelerator and cloud platforms by modeling where time/bytes are spent.
  • •Develop and use performance observability to identify gaps, measure impact of changes, and iterate.
  • •Manage and optimize model state across requests via caching and reuse strategies.
  • •Partner with safety and related teams to maintain production safety systems while improving efficiency without reducing robustness.

Pay and Benefits

Salary: USD 350,000 - 850,000 annually
Perks:Paid LeaveParental Leave

Key Requirements

  • •A working mental model of LLM inference, including how prefill/decode map to compute, memory, and interconnect on accelerators, plus the host’s role.
  • •Demonstrated ability to ramp quickly in deep, unfamiliar systems and ship consequential changes fast.
  • •Strong systems programming experience (Rust, C++, or similar) with attention to code quality and tests.
  • •Analytical approach to performance: observe/profile, form a hypothesis, test, then measure again after changes.
  • •Comfort with pair programming and feedback; able to take initiative and cover work outside your immediate scope.
Experience:LLM serving
Education:Bachelor's
Skills:Performance analysisHypothesis testingRapid learningCollaborationLow ego
Tech Stack:RustC++LLM inferenceTransformer architectureGPU programmingAccelerator programmingPCIeRDMAHBMObservabilityDistributed systems

Company Brief

Anthropic
Develops large-scale AI systems and safety research to create reliable, steerable, and interpretable AI assistants and models for commercial and research applications.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series C
Headquarters: San Francisco, United States
Founded: 2021
WebsiteLinkedIn