Senior Engineer II, Inference Engine - Serving Engine

Digital Ocean
San Francisco
Workplace: RemoteFull timeUSD 167,200 - 209,000 annuallyFunction: Data Science & Machine LearningExperience: 5+ yearsSkills: ["Go","Python","GRPC","Distributed systems","Microservices","Inference engines","Tensor parallelism","KV cache","Dynamo","Ray Serve","NVlink","XGMI","RoCE","Quantization"]

Lead design and delivery of high-scale data plane services powering AI inference for a multi-tenant cloud. Architect resilient systems, optimize distributed hosting for large generative models, mentor engineers, and collaborate across product and engineering teams to align roadmaps with customer needs, while maintaining platform health and observability.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Digital Ocean
Digital Ocean
4 months ago

Senior Engineer II, Inference Engine - Serving Engine

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 6 hours agoStatus: Live

Job Summary

Lead design and delivery of high-scale data plane services powering AI inference for a multi-tenant cloud. Architect resilient systems, optimize distributed hosting for large generative models, mentor engineers, and collaborate across product and engineering teams to align roadmaps with customer needs, while maintaining platform health and observability.
Location: San Francisco
Workplace: Remote
Employment Type: Full time
Job Function: Data Science & Machine Learning

Key Responsibilities

  • •Act as a technical leader on the AI inference engine team, driving end-to-end design, development, and delivery of data plane components for hosting large generative AI models
  • •Architect and refine system designs for a high-scale, multi-tenant AI inference cloud ensuring availability and resiliency
  • •Implement and optimize distributed inference hosting using techniques like tensor/data parallelism and KV cache optimizations
  • •Collaborate cross-functionally with PMs, customer-facing teams, and other engineers to align roadmaps with customer needs
  • •Coach and mentor junior engineers and maintain high standards of observability, reliability, and operational excellence

Pay and Benefits

Salary: USD 167,200 - 209,000 annually

Key Requirements

  • •5+ years of software engineering experience with distributed systems
  • •Experience with hosting large language or multimodal models in inference engines (e.g., vLLM, SGLang, Modular)
  • •Proficiency in Go or Python and familiarity with gRPC
  • •Strong background in microservices, messaging systems, databases, and infrastructure as code
  • •Ability to design, implement, and operate high-scale, resilient data plane services and define SLOs
Experience:5+ yearsAICloudDistributed systems
Skills:GoPythonGRPCDistributed systemsMicroservicesInference enginesTensor parallelismKV cacheDynamoRay ServeNVlinkXGMIRoCEQuantization
Languages:English
Tech Stack:GoPythonGRPCTensor parallelismNVIDIARay ServeDynamoVLLMSGLangModular

Company Brief

Digital Ocean
Provides cloud infrastructure and developer-focused cloud services including scalable droplets, managed databases, Kubernetes, object storage, and networking to simplify deploying and managing applications for developers and small-to-medium businesses.
Industry: Cloud Computing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 250M to 500M
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: New York, United States
Founded: 2011
Glassdoor
Glassdoor: 3.8
WebsiteLinkedIn