Performance Engineer, Inference Systems

Anthropic
San Francisco, New York, Seattle
Workplace: OnsiteFull timeUSD 350,000 - 850,000 annuallyFunction: Software EngineeringEducation: bachelorsSkills: ["Python","SQL","Pandas","Telemetry","Instrumentation","Profiling","Latency","Throughput","Observability","Root-cause analysis","Data analysis"]

Join Anthropic's Inference System Dynamics team to optimize the end-to-end inference fleet that serves Claude across major cloud platforms. You will investigate tail latencies, improve correctness evaluation, and build observability and tooling, collaborating with kernel, routing, autoscaling, and capacity teams to push the performance frontier while ensuring model output quality.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Anthropic
Anthropic
3 months ago

Performance Engineer, Inference Systems

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 10 hours agoStatus: Live

Job Summary

Join Anthropic's Inference System Dynamics team to optimize the end-to-end inference fleet that serves Claude across major cloud platforms. You will investigate tail latencies, improve correctness evaluation, and build observability and tooling, collaborating with kernel, routing, autoscaling, and capacity teams to push the performance frontier while ensuring model output quality.
Location: San Francisco, New York, Seattle
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Run cross-layer performance investigations across throughput, latency, and reliability, sizing the gap between actual fleet performance and theoretical rooflines, identifying root causes, and quantifying the value of closing them.
  • •Own and improve the correctness evaluation pipeline that validates model output quality across hardware platforms, numerics, and serving configurations, and lead the investigation when it catches a regression.
  • •Build the observability, dashboards, and modeling tools that make throughput, latency, cost, reliability, correctness, and their interactions legible across the stack.
  • •Partner with kernel, serving, routing, autoscaling, and capacity teams to prioritize and land the highest-impact optimizations your analysis surfaces.
  • •Ruthlessly stack-rank a large surface area of opportunities by impact and effort, and say no to the ones that don't make the cut.

Pay and Benefits

Salary: USD 350,000 - 850,000 annually

Key Requirements

  • •Hands-on performance engineering experience: profiling, roofline analysis, latency/throughput optimization, and root-cause investigation in complex production systems.
  • •Proficiency in Python, with the ability to read, instrument, and contribute to large production codebases you didn’t write.
  • •Solid data analysis skills (e.g. SQL, pandas, or similar) sufficient to turn raw telemetry into clear findings.
  • •Ability to communicate quantitative results clearly in writing to influence priorities on teams you don’t manage.
  • •Genuine interest in correctness as an engineering discipline: numerics, evaluation design, regression detection.
Experience:AIMLDistributed systemsProduction systemsInference
Education:Bachelor's
Skills:PythonSQLPandasTelemetryInstrumentationProfilingLatencyThroughputObservabilityRoot-cause analysisData analysis
Languages:English
Tech Stack:PythonSQLPandasGPUTPUInferenceModel serversKernelsRoofline analysisObservabilityDashboards

Company Brief

Anthropic
Develops large-scale AI systems and safety research to create reliable, steerable, and interpretable AI assistants and models for commercial and research applications.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series C
Headquarters: San Francisco, United States
Founded: 2021
WebsiteLinkedIn