Staff Engineer, Inference Optimizations

Digital Ocean
Denver
Workplace: RemoteFull timeUSD 191,200 - 239,000 annuallyFunction: Software EngineeringExperience: 5+ yearsSkills: ["Technical leadership","Cross-functional collaboration","Code review","Systems design","Problem-solving"]

Design and lead performance optimization for AI inference at the inference engine and GPU kernel layers. Drive benchmarking and architectural decisions to maximize throughput and minimize latency for large-model inference, including attention, memory/precision, kernel fusion, and distributed multi-node parallelization. Act as an IC leader and subject-matter expert across NVIDIA/AMD stacks (CUDA/ROCm/TensorRT/Triton), quantization (FP8/INT8/FP4), and open-source AI integration while partnering with Product/TPMs to ship developer-friendly capabilities.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Digital Ocean
Digital Ocean
1 month ago

Staff Engineer, Inference Optimizations

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

Design and lead performance optimization for AI inference at the inference engine and GPU kernel layers. Drive benchmarking and architectural decisions to maximize throughput and minimize latency for large-model inference, including attention, memory/precision, kernel fusion, and distributed multi-node parallelization. Act as an IC leader and subject-matter expert across NVIDIA/AMD stacks (CUDA/ROCm/TensorRT/Triton), quantization (FP8/INT8/FP4), and open-source AI integration while partnering with Product/TPMs to ship developer-friendly capabilities.
Location: Denver
Workplace: Remote
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Lead performance architecture for benchmarking and optimizations across the inference engine and GPU kernel layers to maximize throughput and utilization.
  • •Engineer deep-dive optimizations for complex inference performance issues, including attention-layer optimization, memory/precision management, and distributed parallelization.
  • •Implement advanced optimization techniques to stay at the forefront of the Gen AI landscape, including batch-size performance improvements and kernel fusion opportunities.
  • •Serve as the GPU infrastructure and ecosystem subject matter expert (NVIDIA/AMD) across software stacks such as CUDA, ROCm, TensorRT, and Triton, advising on hardware/software integration.
  • •Lead via high-quality code and design reviews, and collaborate with Product Management and TPMs to translate hardware limits into shippable features.

Pay and Benefits

Salary: USD 191,200 - 239,000 annually
Equity and Bonus:Equity
Perks:Paid LeaveLearning Budget

Key Requirements

  • •5+ years of experience in high-performance computing or AI infrastructure, with a track record of resolving compute utilization and memory bandwidth bottlenecks.
  • •Deep familiarity with Gen AI model families (LLM, VLM, LMM) and their architectural requirements.
  • •Hands-on experience with attention-layer optimization and parallelization strategies across distributed GPU environments.
  • •Comprehensive understanding of NVIDIA and AMD GPU architectures and software ecosystems (CUDA, ROCm, etc.).
  • •Expert-level Triton or CUDA experience, including contributing to Triton or writing custom CUDA kernels for major LLMs.
Experience:5+ yearsHigh-performance computingAI infrastructureGen AILLMsOpen source
Skills:Technical leadershipCross-functional collaborationCode reviewSystems designProblem-solving
Languages:English
Tech Stack:CUDAROCmTensorRTOpenAI TritonTritonCUDA kernelsGPU kernelsGPU architecturesNVIDIA GPUsAMD GPUsAMD MI355XAITerFP8BF16INT8FP4FlashAttentionRMS NormTransformerMoE

Company Brief

Digital Ocean
Provides cloud infrastructure and developer-focused cloud services including scalable droplets, managed databases, Kubernetes, object storage, and networking to simplify deploying and managing applications for developers and small-to-medium businesses.
Industry: Cloud Computing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 250M to 500M
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: New York, United States
Founded: 2011
Glassdoor
Glassdoor: 3.8
WebsiteLinkedIn