Staff Software Engineer: AI Inference Data Plane
San Francisco
Workplace: RemoteFull time167,200 - 209,000 annuallyFunction: Solutions Engineering & Sales EngineeringSkills: ["Triton","CUDA","GPU","Kernel optimization","Memory bandwidth","FP8","INT8","FP4","FlashAttention-4","TileLang"]Design and implement high-performance GPU kernels (Triton/CUDA) to maximize throughput and minimize latency for inference workloads. Optimize memory bandwidth, apply quantization techniques (FP8/INT8/FP4), and push architectural breakthroughs like FlashAttention-4 into production. Lead performance improvements across a remote, high-skill engineering team and shape the technical roadmap for a cutting-edge inference fleet.
Loading
Loading job details...
Preparing the role view and application actions.

