Performance Engineer, GPU

Anthropic
San Francisco
Workplace: OnsiteFull timeUSD 315,000 - 560,000 annuallyFunction: Software EngineeringEducation: bachelorsSkills: ["CUDA","Triton","CUTLASS","Flash Attention","Tensor core optimization","PyTorch","JAX","Torch.compile","XLA","Nsight","NCCL","NVLink","INT8","FP8","Quantization","Mixed-precision","Large-scale training","Kernel fusion","Memory bandwidth","Distributed systems"]

Architect and optimize GPU-backed performance engines for large-scale language models, spanning low-level kernel development to multi-node distributed GPU orchestration. You will push CUDA/Triton-based optimizations, improve inference efficiency, and design systems that scale across thousands of GPUs, collaborating with researchers and engineers to deliver measurable performance breakthroughs for state-of-the-art AI models.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Anthropic
Anthropic
11 months ago

Performance Engineer, GPU

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 8 minutes agoStatus: Live

Job Summary

Architect and optimize GPU-backed performance engines for large-scale language models, spanning low-level kernel development to multi-node distributed GPU orchestration. You will push CUDA/Triton-based optimizations, improve inference efficiency, and design systems that scale across thousands of GPUs, collaborating with researchers and engineers to deliver measurable performance breakthroughs for state-of-the-art AI models.
Location: San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Design and implement foundational GPU performance systems for large-scale ML workloads, from kernel-level optimization to orchestration of thousands of GPUs.
  • •Optimize end-to-end training and inference pipelines to improve throughput and efficiency of frontier language models.
  • •Develop custom kernels for quantization formats and mixed-precision techniques; explore kernel fusion to reduce memory bottlenecks.
  • •Design distributed communication strategies for multi-node GPU clusters (NCCL, NVLink, model parallelism) and ensure fault-tolerant operations.
  • •Build performance modeling frameworks and profiling tools to predict and optimize GPU utilization across production workloads.

Pay and Benefits

Salary: USD 315,000 - 560,000 annually

Key Requirements

  • •Deep experience with GPU programming and optimization at scale.
  • •Experience with GPU kernel development: CUDA, Triton, CUTLASS, Flash Attention, tensor core optimization.
  • •ML compilers & frameworks: PyTorch/JAX internals, torch.compile, XLA, custom operators.
  • •Performance engineering: kernel fusion, memory bandwidth optimization, profiling with Nsight.
  • •Production systems: large-scale training infrastructure, fault tolerance, cluster orchestration.
Experience:GPUML systemsDistributed training
Education:Bachelor's
Skills:CUDATritonCUTLASSFlash AttentionTensor core optimizationPyTorchJAXTorch.compileXLANsightNCCLNVLinkINT8FP8QuantizationMixed-precisionLarge-scale trainingKernel fusionMemory bandwidthDistributed systems
Languages:English

Company Brief

Anthropic
Develops large-scale AI systems and safety research to create reliable, steerable, and interpretable AI assistants and models for commercial and research applications.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series C
Headquarters: San Francisco, United States
Founded: 2021
WebsiteLinkedIn