Staff Engineer, Inference Optimizations

Digital Ocean
San Francisco
Workplace: RemoteFull timeUSD 191,200 - 239,000 annuallyFunction: Software EngineeringExperience: 5+ yearsSkills: ["Technical leadership","Design and code reviews","Cross-functional collaboration","Systems design","Performance optimization"]

Drive performance architecture and deep-dive optimization for AI inference on a high-performance inference fleet. Lead benchmarking and improvements across inference engine and GPU kernel layers to maximize throughput and minimize latency for large models. Build solutions for attention, memory/precision management, and parallelization across multi-node GPU clusters, including quantization and kernel fusion. Collaborate with Product and TPMs and serve as an IC leader in open-source AI performance communities.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Digital Ocean
Digital Ocean
1 month ago

Staff Engineer, Inference Optimizations

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

Drive performance architecture and deep-dive optimization for AI inference on a high-performance inference fleet. Lead benchmarking and improvements across inference engine and GPU kernel layers to maximize throughput and minimize latency for large models. Build solutions for attention, memory/precision management, and parallelization across multi-node GPU clusters, including quantization and kernel fusion. Collaborate with Product and TPMs and serve as an IC leader in open-source AI performance communities.
Location: San Francisco
Workplace: Remote
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Lead technical strategy for benchmarking and performance optimizations at inference engine and GPU kernel layers to maximize value per TFLOP.
  • •Engineer solutions to complex performance issues including attention optimizations, memory and precision management, and parallelization across multi-node GPU clusters.
  • •Implement cutting-edge optimization techniques for Gen AI, including batch size tuning, kernel fusion, and expert gateway router tuning for MoE models.
  • •Serve as a subject matter expert on modern GPU families (NVIDIA/AMD) and software stacks (CUDA/ROCm/TensorRT/Triton), advising on procurement and integration.
  • •Mentor through high-quality code and design reviews and partner with Product and TPMs to turn hardware limits into shippable features.

Pay and Benefits

Salary: USD 191,200 - 239,000 annually
Equity and Bonus:Equity
Perks:Learning BudgetPaid LeaveEquity

Key Requirements

  • •5+ years experience in high-performance computing or AI infrastructure, with a track record of improving compute utilization and memory bandwidth bottlenecks.
  • •Deep familiarity with the Gen AI landscape (LLM, VLM, LMM) and model-family architectural requirements.
  • •Hands-on experience with attention-layer optimization and parallelization across distributed GPU environments.
  • •Strong understanding of NVIDIA and AMD GPU architectures and their software ecosystems (CUDA, ROCm).
  • •Expert-level Triton or CUDA experience, including contributing to Triton or writing custom CUDA kernels for major LLMs.
Experience:5+ yearsHigh-performance computingAI infrastructureGen AILarge language modelsOpen source
Skills:Technical leadershipDesign and code reviewsCross-functional collaborationSystems designPerformance optimization
Languages:English
Tech Stack:CUDAROCmTensorRTOpenAI TritonTritonAITERAMD MI355XTFLOPFP8BF16INT8FP4FlashAttentionRMS NormGPU kernelsCUDA kernels

Company Brief

Digital Ocean
Provides cloud infrastructure and developer-focused cloud services including scalable droplets, managed databases, Kubernetes, object storage, and networking to simplify deploying and managing applications for developers and small-to-medium businesses.
Industry: Cloud Computing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 250M to 500M
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: New York, United States
Founded: 2011
Glassdoor
Glassdoor: 3.8
WebsiteLinkedIn