Staff Engineer, Inference Optimizations

Digital Ocean
Seattle
Workplace: HybridFull timeUSD 191,200 - 239,000 annuallyFunction: Software EngineeringExperience: 5+ yearsSkills: ["Technical leadership","Cross-functional collaboration","Mentorship","Design reviews","Problem-solving"]

Lead AI inference optimization for a high-performance inference fleet, making architectural decisions to maximize throughput and minimize latency for large models. Own performance benchmarking and deep-dive optimization across inference engines and GPU kernel layers, including attention optimizations, memory/precision management, and multi-node GPU parallelization. Drive innovation in GenAI inference using advanced quantization and kernel fusion, serve as a GPU ecosystem subject matter expert, and mentor through code/design reviews while partnering with Product and TPMs.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Digital Ocean
Digital Ocean
2 months ago

Staff Engineer, Inference Optimizations

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

Lead AI inference optimization for a high-performance inference fleet, making architectural decisions to maximize throughput and minimize latency for large models. Own performance benchmarking and deep-dive optimization across inference engines and GPU kernel layers, including attention optimizations, memory/precision management, and multi-node GPU parallelization. Drive innovation in GenAI inference using advanced quantization and kernel fusion, serve as a GPU ecosystem subject matter expert, and mentor through code/design reviews while partnering with Product and TPMs.
Location: Seattle
Workplace: Hybrid
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Lead benchmarking and performance optimization strategy across the inference engine and GPU kernel layers to maximize throughput and utilization.
  • •Develop solutions for complex performance issues such as attention layer optimizations, memory and precision management, and advanced parallelization across multi-node GPU clusters.
  • •Implement cutting-edge GenAI inference optimization techniques, including batch size tuning, kernel fusion, and expert gateway router kernel tuning for MoE models.
  • •Serve as subject matter expert on modern GPU families (NVIDIA/AMD) and their software stacks, advising on hardware procurement and software integration.
  • •Lead technical mentorship through high-quality code and design reviews, and partner with Product Management/TPMs to turn hardware constraints into shippable features.

Pay and Benefits

Salary: USD 191,200 - 239,000 annually
Equity and Bonus:Equity
Perks:Flexible TimeEmployee AssistanceLearning Budget

Key Requirements

  • •5+ years of experience in high-performance computing or AI infrastructure, with a proven track record addressing compute utilization and memory bandwidth bottlenecks.
  • •Deep familiarity with Gen AI model families (LLM, VLM, LMM) and their architectural requirements.
  • •Hands-on optimization experience including attention-layer optimizations and parallelization across distributed GPU environments.
  • •Comprehensive understanding of NVIDIA and AMD GPU architectures and software ecosystems (e.g., CUDA, ROCm).
  • •Expert-level Triton or CUDA experience, including contributions to the Triton compiler or custom CUDA kernels for large models.
Experience:5+ yearsAI infrastructureHigh-performance computingGen AILLMOpen source
Skills:Technical leadershipCross-functional collaborationMentorshipDesign reviewsProblem-solving
Languages:English
Tech Stack:AI inference optimizationGPU kernelsTritonCUDAROCmTensorRTOpenAI TritonAMD AITERFlashAttentionRMS NormFP8BF16FP4INT8TFLOPTransformersMoEMulti-node GPUsExpert gateway routerCUDA kernels

Company Brief

Digital Ocean
Provides cloud infrastructure and developer-focused cloud services including scalable droplets, managed databases, Kubernetes, object storage, and networking to simplify deploying and managing applications for developers and small-to-medium businesses.
Industry: Cloud Computing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 250M to 500M
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: New York, United States
Founded: 2011
Glassdoor
Glassdoor: 3.8
WebsiteLinkedIn