Senior System Software Engineer - Local AI

NVIDIA
India
Workplace: OnsiteFull timeFunction: Software EngineeringExperience: 5+ yearsSkills: ["Debugging","Analytical thinking","Problem-solving","Communication","Collaboration","Multitasking"]

Build and optimize NVIDIA’s local AI inference stack for RTX, RTX Pro, and DGX-class GPUs, focusing on high performance, stability, and scalability. Develop modern inference runtimes and execution stacks spanning frameworks such as Llama.cpp, vLLM, PyTorch, WinML, DXCGC, and TensorRT-RTX. Own end-to-end optimization of models, data pipelines, and runtimes, including quantization, pruning, sparsity, and distillation, and drive production-ready bring-up through debugging and performance analysis.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
1 month ago

Senior System Software Engineer - Local AI

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 2 hours agoStatus: Live

Job Summary

Build and optimize NVIDIA’s local AI inference stack for RTX, RTX Pro, and DGX-class GPUs, focusing on high performance, stability, and scalability. Develop modern inference runtimes and execution stacks spanning frameworks such as Llama.cpp, vLLM, PyTorch, WinML, DXCGC, and TensorRT-RTX. Own end-to-end optimization of models, data pipelines, and runtimes, including quantization, pruning, sparsity, and distillation, and drive production-ready bring-up through debugging and performance analysis.
Location: India
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Partner with NVIDIA software, research, architecture, and product teams to align strategies and technical needs for the AI ecosystem on RTX and DGX PCs.
  • •Build and optimize local AI inference stack for RTX, RTX Pro, and DGX GPUs across hardware architectures.
  • •Design and develop modern inference runtimes and execution stacks across LLM, vision-language, TTS, ASR, and diffusion AI workloads.
  • •Perform end-to-end optimization of AI models, data pipelines, and inference runtimes, applying quantization, pruning, sparsity, and distillation for efficient deployment.
  • •Conduct system-level debugging and performance–accuracy trade-off analysis, establish engineering guidelines for bring-up, and ensure production readiness of new inference backends.

Key Requirements

  • •5+ years of experience with a Bachelor’s, Master’s, or PhD in Computer Science, Software Engineering, Mathematics, or related field, or equivalent experience.
  • •Excellent C++ programming and debugging skills with strong understanding of data structures, algorithms, and machine learning.
  • •Proven experience with AI inferencing pipelines and applications using ML/DL frameworks such as Llama.cpp, vLLM, PyTorch, WinML, DXCGC, and TensorRT.
  • •Deep interest in inference backends and runtime internals, including scheduling, memory management, KV-cache behavior, graph execution, quantization, and hardware-aware optimization.
  • •Strong analytical and problem-solving abilities, plus strong written and oral communication skills for collaboration.
Experience:5+ yearsAIMLDeep learningGenerative AIOpen source
Education:
Skills:DebuggingAnalytical thinkingProblem-solvingCommunicationCollaborationMultitasking
Tech Stack:C++Llama.cppVLLMPyTorchWinMLDXCGCTensorRT-RTXTensorRTCUDAVulkanDirectXRTXDGX

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor