Staff Engineer, Inference Optimizations
Digital Ocean
Seattle
Workplace: HybridFull timeUSD 191,200 - 239,000 annuallyFunction: Software EngineeringExperience: 5+ yearsSkills: ["Technical leadership","Cross-functional collaboration","Mentorship","Design reviews","Problem-solving"]Lead AI inference optimization for a high-performance inference fleet, making architectural decisions to maximize throughput and minimize latency for large models. Own performance benchmarking and deep-dive optimization across inference engines and GPU kernel layers, including attention optimizations, memory/precision management, and multi-node GPU parallelization. Drive innovation in GenAI inference using advanced quantization and kernel fusion, serve as a GPU ecosystem subject matter expert, and mentor through code/design reviews while partnering with Product and TPMs.

