Director, Engineering - Inference Serving Engine

Digital Ocean
Bengaluru
Workplace: HybridFull timeFunction: Executive & General ManagementSkills: ["Leadership","Communication","Mentorship","Ownership","Strategic thinking"]

Lead a high-performing engineering team focused on designing, building, and scaling a distributed LLM inference platform across serving, orchestration, and hosting layers. Drive technical strategy, architecture, and delivery for GPU-accelerated workloads, collaborate cross-functionally, and ensure production health and reliability at scale.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Digital Ocean
Digital Ocean
2 months ago

Director, Engineering - Inference Serving Engine

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 3 hours agoStatus: Live

Job Summary

Lead a high-performing engineering team focused on designing, building, and scaling a distributed LLM inference platform across serving, orchestration, and hosting layers. Drive technical strategy, architecture, and delivery for GPU-accelerated workloads, collaborate cross-functionally, and ensure production health and reliability at scale.
Location: Bengaluru
Workplace: Hybrid
Employment Type: Full time
Job Function: Executive & General Management
Seniority: Director level

Key Responsibilities

  • •Team Leadership & Development: Recruit, mentor, and coach engineers on the team, fostering ownership, technical excellence, and continuous improvement.
  • •Execution & Delivery: Own the team's project execution, translating business goals into technical roadmaps, milestones, and on-time delivery.
  • •Cross-Functional Partnership: Collaborate with Product Management and other engineering teams to align priorities, manage dependencies, and communicate progress.
  • •Operational Health: Ensure production health, stability, and on-call rotation of all services owned by the Inference Orchestration team.
  • •Oversee System Design: Guide architecture and implementation of a distributed inference platform optimized for diverse GPU platforms (NVIDIA and AMD).

Pay and Benefits

Perks:EquityLearning BudgetHealth InsuranceRemote Work

Key Requirements

  • •10+ years of software engineering experience, with 6+ years in a technical leadership or management role, ideally within AI/ML infrastructure or cloud platforms
  • •Deep expertise in distributed systems design, modern AI/ML technologies, Kubernetes at scale, and LLM inference, and AI workload orchestration, scheduling, and resource management
  • •Strong knowledge of GPU architectures (NVIDIA and/or AMD), interconnects (e.g., NVLink), and hardware topology and their impact on AI training and inference performance
  • •Experience with container runtimes, system isolation, and security contexts to manage risk in shared infrastructure
  • •Excellent communication skills and ability to translate complex technical requirements into user-focused product features
Experience:AICloudInfrastructureLLMKubernetes
Skills:LeadershipCommunicationMentorshipOwnershipStrategic thinking
Languages:English
Tech Stack:CUDAROCmNVIDIAAMDKubernetesGroveLLMCheckpoint/RestoreNVLink

Company Brief

Digital Ocean
Provides cloud infrastructure and developer-focused cloud services including scalable droplets, managed databases, Kubernetes, object storage, and networking to simplify deploying and managing applications for developers and small-to-medium businesses.
Industry: Cloud Computing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 250M to 500M
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: New York, United States
Founded: 2011
Glassdoor
Glassdoor: 3.8
WebsiteLinkedIn