Staff Engineer, Inference Optimizations
Digital Ocean
Denver
Workplace: RemoteFull timeUSD 191,200 - 239,000 annuallyFunction: Software EngineeringExperience: 5+ yearsSkills: ["Technical leadership","Cross-functional collaboration","Code review","Systems design","Problem-solving"]Design and lead performance optimization for AI inference at the inference engine and GPU kernel layers. Drive benchmarking and architectural decisions to maximize throughput and minimize latency for large-model inference, including attention, memory/precision, kernel fusion, and distributed multi-node parallelization. Act as an IC leader and subject-matter expert across NVIDIA/AMD stacks (CUDA/ROCm/TensorRT/Triton), quantization (FP8/INT8/FP4), and open-source AI integration while partnering with Product/TPMs to ship developer-friendly capabilities.

