Performance Engineer, GPU
San Francisco
Workplace: OnsiteFull timeUSD 315,000 - 560,000 annuallyFunction: Software EngineeringEducation: bachelorsSkills: ["CUDA","Triton","CUTLASS","Flash Attention","Tensor core optimization","PyTorch","JAX","Torch.compile","XLA","Nsight","NCCL","NVLink","INT8","FP8","Quantization","Mixed-precision","Large-scale training","Kernel fusion","Memory bandwidth","Distributed systems"]Architect and optimize GPU-backed performance engines for large-scale language models, spanning low-level kernel development to multi-node distributed GPU orchestration. You will push CUDA/Triton-based optimizations, improve inference efficiency, and design systems that scale across thousands of GPUs, collaborating with researchers and engineers to deliver measurable performance breakthroughs for state-of-the-art AI models.
Loading
Loading job details...
Preparing the role view and application actions.

