Staff Engineer, Inference Optimizations

Digital Ocean
Boston
Workplace: RemoteFull timeUSD 191,200 - 239,000 annuallyFunction: Software EngineeringExperience: 5+ yearsSkills: ["Technical leadership","Code/design review","Cross-functional collaboration"]

Lead performance architecture and deep-dive optimization for DigitalOcean’s AI inference stack, improving throughput and reducing latency across inference engine and GPU kernel layers. Drive benchmarking, kernel-level and attention/memory/precision optimizations, and advanced parallelization for multi-node GPU clusters. Serve as a subject matter expert across NVIDIA/AMD ecosystems (CUDA/ROCm/TensorRT/Triton), guide technical roadmap decisions, and translate hardware limits into shippable, developer-friendly product features.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Digital Ocean
Digital Ocean
1 month ago

Staff Engineer, Inference Optimizations

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

Lead performance architecture and deep-dive optimization for DigitalOcean’s AI inference stack, improving throughput and reducing latency across inference engine and GPU kernel layers. Drive benchmarking, kernel-level and attention/memory/precision optimizations, and advanced parallelization for multi-node GPU clusters. Serve as a subject matter expert across NVIDIA/AMD ecosystems (CUDA/ROCm/TensorRT/Triton), guide technical roadmap decisions, and translate hardware limits into shippable, developer-friendly product features.
Location: Boston
Workplace: Remote
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Lead technical strategy for benchmarking and performance optimizations across inference engine and GPU kernel layers.
  • •Engineer solutions for complex performance issues, including attention layer optimizations, memory/precision management, and parallelization across multi-node GPU clusters.
  • •Implement cutting-edge optimization techniques for GenAI inference, including batch size performance tuning, kernel fusion, and MoE router kernel tuning.
  • •Act as subject matter expert on modern GPU families and software stacks, advising on hardware procurement and software integration.
  • •Provide technical mentorship through high-quality code and design reviews, and collaborate with Product Management and TPMs to ship performance-driven features.

Pay and Benefits

Salary: USD 191,200 - 239,000 annually
Perks:EquityEsppEmployee AssistanceFlexible Time

Key Requirements

  • •5+ years of experience in high-performance computing or AI infrastructure, with a track record of resolving compute utilization and memory bandwidth bottlenecks.
  • •Deep familiarity with the Gen AI landscape (LLM, VLM, LMM) and architectural requirements of major model families.
  • •Hands-on experience with attention-layer optimizations and parallelization strategies across distributed GPU environments.
  • •Comprehensive understanding of NVIDIA and AMD GPU architectures and their software ecosystems (CUDA, ROCm).
  • •Expert-level Triton or CUDA, including experience contributing to Triton or writing custom CUDA kernels for major LLMs.
Experience:5+ yearsHigh-performance computingAI infrastructureGenAIOpen sourceGPU optimization
Skills:Technical leadershipCode/design reviewCross-functional collaboration
Languages:English
Tech Stack:GPUNVIDIAAMDCUDAROCmTensorRTOpenAI TritonTriton compilerCUDA kernelsAMD AITERMI355XFP8BF16FP4INT8FlashAttentionRMS NormTransformerMoEExpert gateway router

Company Brief

Digital Ocean
Provides cloud infrastructure and developer-focused cloud services including scalable droplets, managed databases, Kubernetes, object storage, and networking to simplify deploying and managing applications for developers and small-to-medium businesses.
Industry: Cloud Computing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 250M to 500M
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: New York, United States
Founded: 2011
Glassdoor
Glassdoor: 3.8
WebsiteLinkedIn