Staff Software Engineer: AI Inference Data Plane

Digital Ocean
San Francisco
Workplace: RemoteFull time167,200 - 209,000 annuallyFunction: Solutions Engineering & Sales EngineeringSkills: ["Triton","CUDA","GPU","Kernel optimization","Memory bandwidth","FP8","INT8","FP4","FlashAttention-4","TileLang"]

Design and implement high-performance GPU kernels (Triton/CUDA) to maximize throughput and minimize latency for inference workloads. Optimize memory bandwidth, apply quantization techniques (FP8/INT8/FP4), and push architectural breakthroughs like FlashAttention-4 into production. Lead performance improvements across a remote, high-skill engineering team and shape the technical roadmap for a cutting-edge inference fleet.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Digital Ocean
Digital Ocean
4 months ago

Staff Software Engineer: AI Inference Data Plane

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 6 hours agoStatus: Live

Job Summary

Design and implement high-performance GPU kernels (Triton/CUDA) to maximize throughput and minimize latency for inference workloads. Optimize memory bandwidth, apply quantization techniques (FP8/INT8/FP4), and push architectural breakthroughs like FlashAttention-4 into production. Lead performance improvements across a remote, high-skill engineering team and shape the technical roadmap for a cutting-edge inference fleet.
Location: San Francisco
Workplace: Remote
Employment Type: Full time
Job Function: Solutions Engineering & Sales Engineering

Key Responsibilities

  • •Design and implement high-performance GPU kernels using Triton and CUDA C++
  • •Develop and deploy quantization techniques (FP8, INT8, FP4) to double throughput without loss of accuracy
  • •Optimize memory access patterns (SRAM vs. HBM3e) to remove bottlenecks in long-context attention
  • •Implement architectural breakthroughs like FlashAttention-4 and TileLang into production stack
  • •Guide the technical roadmap for high-performance inference fleet and address bottlenecks in memory bandwidth and compute utilization

Pay and Benefits

Salary: 167,200 - 209,000 annually
Equity and Bonus:Equity
Perks:Remote WorkEquity

Key Requirements

  • •Deep understanding of GPU architectures and programming with CUDA/C++ and Triton
  • •Proven track record designing high-performance GPU kernels and optimizing memory access patterns
  • •Experience with quantization techniques (FP8, INT8, FP4) and throughput optimization
  • •Familiarity with recent AI inference workloads and production pipelines (FlashAttention-4, TileLang)
Experience:AICloudGPUInference
Skills:TritonCUDAGPUKernel optimizationMemory bandwidthFP8INT8FP4FlashAttention-4TileLang
Languages:English
Tech Stack:TritonCUDAFP8INT8FP4FlashAttention-4TileLang

Company Brief

Digital Ocean
Provides cloud infrastructure and developer-focused cloud services including scalable droplets, managed databases, Kubernetes, object storage, and networking to simplify deploying and managing applications for developers and small-to-medium businesses.
Industry: Cloud Computing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 250M to 500M
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: New York, United States
Founded: 2011
Glassdoor
Glassdoor: 3.8
WebsiteLinkedIn