Performance Engineer (Inference, Training & GPU)

World Labs
San Francisco
Workplace: OnsiteFull timeUSD 200,000 - 300,000 annuallyFunction: Education & TrainingSkills: ["High ownership","Root-cause investigation","Collaboration","Problem-solving","Attention to numerical correctness"]

Own performance for large generative world models across training and production serving. Profile and eliminate bottlenecks end-to-end—latency, throughput, batching, caching, scheduling, GPU utilization, and pipeline stalls. Write and tune GPU kernels with CUDA and Triton, optimize mixed/low precision execution (FP8/INT8), and build performance models and observability to make tradeoffs measurable. Partner with researchers to productionize models and protect numerical correctness.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
World Labs
World Labs
4 months ago

Performance Engineer (Inference, Training & GPU)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 5 hours agoStatus: Live
Reposted: similar role first listed 4 months ago

Job Summary

Own performance for large generative world models across training and production serving. Profile and eliminate bottlenecks end-to-end—latency, throughput, batching, caching, scheduling, GPU utilization, and pipeline stalls. Write and tune GPU kernels with CUDA and Triton, optimize mixed/low precision execution (FP8/INT8), and build performance models and observability to make tradeoffs measurable. Partner with researchers to productionize models and protect numerical correctness.
Location: San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Education & Training

Key Responsibilities

  • •Optimize inference and serving end to end for latency, throughput, batching, caching, and scheduling at production scale.
  • •Write and tune GPU kernels with CUDA and Triton, driving kernel fusion, memory/bandwidth optimization, and low-precision execution.
  • •Optimize training throughput and GPU utilization using parallelism strategies, communication/compute overlap, mixed precision, and removing pipeline stalls.
  • •Build performance models, profiling workflows, and observability to make throughput, latency, cost, and utilization tradeoffs legible.
  • •Own numerical correctness across precision, kernel, and hardware changes while partnering with researchers to productionize models.

Pay and Benefits

Salary: USD 200,000 - 300,000 annually

Key Requirements

  • •Strong performance-engineering fundamentals including profiling, roofline analysis, latency/throughput optimization, and root-cause investigation.
  • •Deep GPU programming and optimization experience with CUDA and/or Triton, including kernel-level tuning and bandwidth/memory hierarchy optimization.
  • •Hands-on experience optimizing inference and serving for large models, including batching, KV/prompt caching, quantization, and low-latency/high-throughput sampling.
  • •Hands-on experience optimizing training performance, including parallelism strategies, distributed communication, mixed/low precision, and utilization.
  • •Working knowledge of ML framework internals, including PyTorch and/or JAX (e.g., torch.compile or XLA), and strong proficiency in Python with ability to work in C++/CUDA (and Rust or Go as needed).
Experience:AI researchMachine learningGenerative AI
Skills:High ownershipRoot-cause investigationCollaborationProblem-solvingAttention to numerical correctness
Tech Stack:CUDATritonFP8INT8PyTorchJAXTorch.compileXLAC++RustGoKV cachingPrompt cachingQuantizationNCCLNVLink

Company Brief

World Labs
World Labs develops artificial intelligence and machine learning solutions, providing platforms and tools to help organizations build, deploy, and scale AI-driven applications across industries, with a focus on model development, MLOps, and customizable AI services.
Industry: AI & Machine Learning
Website