Staff Engineer, Inference Optimizations
Digital Ocean
Austin
Workplace: RemoteFull timeUSD 191,200 - 239,000Function: Software EngineeringExperience: 5+ yearsSkills: ["Technical leadership","Cross-functional collaboration","Systems design","Technical mentorship","Performance optimization"]Design and lead performance architecture for AI inference services, optimizing benchmarking and throughput/latency at the inference engine and GPU kernel layers. Tackle deep performance bottlenecks including attention-layer optimizations, memory/precision management, and parallelization across multi-node GPU clusters. Drive GenAI-focused innovation with tools like Triton and CUDA, work across NVIDIA/AMD ecosystems, mentor through code/design reviews, and partner with product and TPMs to ship high-performance inference features.

