Performance Engineer (Inference, Training & GPU)
San Francisco
Workplace: OnsiteFull timeUSD 200,000 - 300,000 annuallyFunction: Education & TrainingSkills: ["High ownership","Root-cause investigation","Collaboration","Problem-solving","Attention to numerical correctness"]Own performance for large generative world models across training and production serving. Profile and eliminate bottlenecks end-to-end—latency, throughput, batching, caching, scheduling, GPU utilization, and pipeline stalls. Write and tune GPU kernels with CUDA and Triton, optimize mixed/low precision execution (FP8/INT8), and build performance models and observability to make tradeoffs measurable. Partner with researchers to productionize models and protect numerical correctness.
Loading
Loading job details...
Preparing the role view and application actions.

