Staff Engineer, Inference Optimizations

Digital Ocean
Austin
Workplace: RemoteFull timeUSD 191,200 - 239,000Function: Software EngineeringExperience: 5+ yearsSkills: ["Technical leadership","Cross-functional collaboration","Systems design","Technical mentorship","Performance optimization"]

Design and lead performance architecture for AI inference services, optimizing benchmarking and throughput/latency at the inference engine and GPU kernel layers. Tackle deep performance bottlenecks including attention-layer optimizations, memory/precision management, and parallelization across multi-node GPU clusters. Drive GenAI-focused innovation with tools like Triton and CUDA, work across NVIDIA/AMD ecosystems, mentor through code/design reviews, and partner with product and TPMs to ship high-performance inference features.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Digital Ocean
Digital Ocean
1 month ago

Staff Engineer, Inference Optimizations

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

Design and lead performance architecture for AI inference services, optimizing benchmarking and throughput/latency at the inference engine and GPU kernel layers. Tackle deep performance bottlenecks including attention-layer optimizations, memory/precision management, and parallelization across multi-node GPU clusters. Drive GenAI-focused innovation with tools like Triton and CUDA, work across NVIDIA/AMD ecosystems, mentor through code/design reviews, and partner with product and TPMs to ship high-performance inference features.
Location: Austin
Workplace: Remote
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Lead performance architecture for benchmarking and optimization across the inference engine and GPU kernel layers to maximize throughput and minimize latency.
  • •Engineer solutions for complex inference performance issues including attention layer optimizations, memory and precision management, and advanced parallelization across multi-node GPU clusters.
  • •Implement cutting-edge optimization techniques for GenAI inference, including batch-size performance work and kernel fusion opportunities.
  • •Act as a subject matter expert on modern GPU families (NVIDIA/AMD) and their software stacks, advising on hardware procurement and software integration.
  • •Lead through high-quality code and design reviews while partnering with Product Management and TPMs to translate hardware limits into shippable features.

Pay and Benefits

Salary: USD 191,200 - 239,000
Perks:Paid LeaveLearning Budget

Key Requirements

  • •5+ years of experience in high-performance computing or AI infrastructure with a track record of improving compute utilization and memory bandwidth bottlenecks.
  • •Deep familiarity with Gen AI model families (LLM, VLM, LMM) and their architectural requirements.
  • •Hands-on expertise in attention-layer optimizations and parallelization strategies for distributed GPU environments.
  • •Comprehensive understanding of NVIDIA and AMD GPU architectures and their software ecosystems (e.g., CUDA, ROCm).
  • •Expert-level Triton or CUDA experience, including contributions to Triton or writing custom CUDA kernels for major LLMs.
Experience:5+ yearsHigh-performance computingAI infrastructureGenAILLM inferenceOpen source
Skills:Technical leadershipCross-functional collaborationSystems designTechnical mentorshipPerformance optimization
Languages:English
Tech Stack:TritonCUDAROCmTensorRTOpenAI TritonAMD AITERFlashAttentionRMS NormGLM-5MoEQwen3-235BDeepSeek V3MoE modelsFP8BF16INT8FP4GPU kernelsMulti-node GPU clustersTransformers

Company Brief

Digital Ocean
Provides cloud infrastructure and developer-focused cloud services including scalable droplets, managed databases, Kubernetes, object storage, and networking to simplify deploying and managing applications for developers and small-to-medium businesses.
Industry: Cloud Computing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 250M to 500M
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: New York, United States
Founded: 2011
Glassdoor
Glassdoor: 3.8
WebsiteLinkedIn