Senior Forward Deployed Engineer I (AI Inference)

Digital Ocean
Bengaluru
Workplace: HybridFull timeFunction: Data Science & Machine LearningSkills: ["Growth mindset","Technical leadership","Customer empathy","Communication","Bias for action"]

Embed with AI founders and strategic AI enterprises to build and deploy low-latency, high-throughput LLM inference on DigitalOcean’s GPU cloud. Own the full lifecycle of production-grade, multi-tenant inference infrastructure—architecting distributed systems, profiling bottlenecks, and fixing KV-cache locality issues. Lead customer-facing debugging and performance work to reduce time-to-first-token and time-per-output-token using Kubernetes-native tooling and GPU-efficiency techniques.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Digital Ocean
Digital Ocean
4 days ago

Senior Forward Deployed Engineer I (AI Inference)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 21 hours agoStatus: Live

Job Summary

Embed with AI founders and strategic AI enterprises to build and deploy low-latency, high-throughput LLM inference on DigitalOcean’s GPU cloud. Own the full lifecycle of production-grade, multi-tenant inference infrastructure—architecting distributed systems, profiling bottlenecks, and fixing KV-cache locality issues. Lead customer-facing debugging and performance work to reduce time-to-first-token and time-per-output-token using Kubernetes-native tooling and GPU-efficiency techniques.
Location: Bengaluru
Workplace: Hybrid
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Mid level

Key Responsibilities

  • •Act as an AI inference lead on the FDE team, driving end-to-end design, development, and delivery of critical AI workloads.
  • •Architect and deploy production-grade, multi-tenant LLM inference engines using Kubernetes-native frameworks.
  • •Embed with external tech leads to debug latency spikes, profile GPU memory utilization, and refactor inference code for high-concurrency workloads.
  • •Solve cluster-scale LLM serving distributed-systems challenges such as prefill/decode disaggregation, KV-cache-aware routing, tiered prefix caching, and MoE parallelism.
  • •Guide customers on compute efficiency using tensor/data parallelism, continuous batching, and quantization to improve tokens-per-second per dollar.
Travel: Medium travel

Pay and Benefits

Perks:Annual BonusEquityEmployee AssistanceLearning Budget

Key Requirements

  • •6+ years working in AI/ML systems with deep knowledge of cluster-scale AI inference serving challenges.
  • •Hands-on experience with forward deployed engineering and AI inference/technical consulting supporting production AI systems.
  • •Experience with inference frameworks such as vLLM, llm-d, SGLang, TensorRT-LLM, or Modular MAX, including internals like continuous batching and paged attention.
  • •Expert proficiency in Python or GoLang, with familiarity with gRPC and experience running critical Kubernetes services at high scale.
  • •Ability to translate business latency SLAs into technical infrastructure solutions and collaborate directly with external engineering teams.
Experience:AI/MLDistributed systemsLLM servingForward deployed engineeringOpen-source ecosystems
Skills:Growth mindsetTechnical leadershipCustomer empathyCommunicationBias for action
Languages:English
Tech Stack:KubernetesLlm-dNVIDIA DynamoRay ServeVLLMSGLangTensorRT-LLMModular MAXPythonGoLangGRPCKV-cacheContinuous batchingPaged attentionTensor parallelismData parallelismQuantizationFP8FP4MoE

Company Brief

Digital Ocean
Provides cloud infrastructure and developer-focused cloud services including scalable droplets, managed databases, Kubernetes, object storage, and networking to simplify deploying and managing applications for developers and small-to-medium businesses.
Industry: Cloud Computing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 250M to 500M
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: New York, United States
Founded: 2011
Glassdoor
Glassdoor: 3.8
WebsiteLinkedIn