Distinguished Engineer, Scaled Out Inferencing

NVIDIA
Santa Clara, California
Workplace: RemoteFull timeUSD 320,000 - 488,750 annuallyFunction: Design (Product/UX/UI/Visual)Experience: 16+ yearsEducation: mastersSkills: ["Technical leadership","Cross-functional collaboration","Stakeholder engagement","Communication","Organizational alignment"]

Lead the technical roadmap for scaled-out AI inferencing, architecting high-throughput, low-latency distributed pipelines and model serving strategies for massive scale. Own full-lifecycle model deployment, versioning, and automated scaling across enterprise and cloud environments. Collaborate on hardware-software co-optimization, GPU resource management, and kernel/driver performance tuning, while guiding open-source and ecosystem initiatives to run advanced AI models efficiently on NVIDIA hardware.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
1 hour ago

Distinguished Engineer, Scaled Out Inferencing

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 1 hour agoStatus: Live

Job Summary

Lead the technical roadmap for scaled-out AI inferencing, architecting high-throughput, low-latency distributed pipelines and model serving strategies for massive scale. Own full-lifecycle model deployment, versioning, and automated scaling across enterprise and cloud environments. Collaborate on hardware-software co-optimization, GPU resource management, and kernel/driver performance tuning, while guiding open-source and ecosystem initiatives to run advanced AI models efficiently on NVIDIA hardware.
Location: Santa Clara, California
Workplace: Remote
Employment Type: Full time
Job Function: Design (Product/UX/UI/Visual)
Seniority: Mid level

Key Responsibilities

  • •Architect distributed inference pipelines and define implementation for high-throughput, low-latency systems at massive scale.
  • •Drive hardware-software co-optimization and performance tuning at the kernel/driver level for production-grade model serving.
  • •Guide open-source and ecosystem projects (e.g., Dynamo, TensorRT-LLM, vLLM, SGLang, Linux, Kubernetes, Ray) to enable state-of-the-art inferencing on NVIDIA hardware.
  • •Own full-lifecycle model management, including automated deployment, versioning, and intelligent scaling across cloud and datacenter environments.
  • •Collaborate with customers, infrastructure providers, and partners to ensure high performance and availability of NVIDIA solutions.

Pay and Benefits

Salary: USD 320,000 - 488,750 annually
Equity and Bonus:Equity

Key Requirements

  • •16+ overall years in technical roles with recent long-term focus on AI infrastructure and direct experience in large-scale inference orchestration.
  • •7-10+ years of leadership experience.
  • •BS/MS or higher (or equivalent) in systems/software engineering or related fields.
  • •Proficiency in GPU architecture and hardware acceleration with low-level performance tuning (CUDA, kernels) and cloud-native architectures for multi-tenant model serving.
  • •Proven track record delivering secure, highly available, distributed production systems with operational visibility into resource utilization and performance.
Experience:16+ yearsAI infrastructureDistributed systemsCloud-nativeEnterprise environmentsOpen source
Education:Master's
Skills:Technical leadershipCross-functional collaborationStakeholder engagementCommunicationOrganizational alignment
Tech Stack:CUDALinuxKubernetesRayGPU architectureHardware accelerationTensorRT-LLMDynamoVLLMSGLang

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor