Staff Engineer, Inference Optimizations
Digital Ocean
San Francisco
Workplace: RemoteFull timeUSD 191,200 - 239,000 annuallyFunction: Software EngineeringExperience: 5+ yearsSkills: ["Technical leadership","Design and code reviews","Cross-functional collaboration","Systems design","Performance optimization"]Drive performance architecture and deep-dive optimization for AI inference on a high-performance inference fleet. Lead benchmarking and improvements across inference engine and GPU kernel layers to maximize throughput and minimize latency for large models. Build solutions for attention, memory/precision management, and parallelization across multi-node GPU clusters, including quantization and kernel fusion. Collaborate with Product and TPMs and serve as an IC leader in open-source AI performance communities.

