Senior System Software Engineer - LocalAI

NVIDIA
India
Workplace: OnsiteFull timeFunction: Software EngineeringExperience: 5+ yearsEducation: bachelorsSkills: ["Programming","Debugging","Analytical thinking","Problem-solving","Communication"]

Develop on-device AI software for RTX and DGX-class systems, focusing on high-performance local inference with low latency and robust memory and infrastructure. Build and optimize inference stacks and modern runtime execution backends for LLM, vision-language, TTS, ASR, and diffusion workloads. Perform end-to-end optimization across model pipelines and inference runtimes, including quantization, pruning, sparsity, and distillation, while tuning and debugging system performance.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
1 month ago

Senior System Software Engineer - LocalAI

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 2 hours agoStatus: Live

Job Summary

Develop on-device AI software for RTX and DGX-class systems, focusing on high-performance local inference with low latency and robust memory and infrastructure. Build and optimize inference stacks and modern runtime execution backends for LLM, vision-language, TTS, ASR, and diffusion workloads. Perform end-to-end optimization across model pipelines and inference runtimes, including quantization, pruning, sparsity, and distillation, while tuning and debugging system performance.
Location: India
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Partner with NVIDIA software, research, architecture, and product teams to align technical requirements and strategic priorities for the AI ecosystem on RTX and DGX PCs.
  • •Build and optimize local AI inference stacks for RTX, RTX Pro, and DGX GPUs with performance, stability, and scalability across hardware architectures.
  • •Design and develop inference runtimes and execution stacks using llama.cpp, vLLM, PyTorch, WinML, DXCGC, and TensorRT-RTX for LLM, vision-language, TTS, ASR, and diffusion workloads.
  • •Perform end-to-end optimization of AI models, data pipelines, and inference runtimes to maximize performance on current and next-generation GPU architectures, using quantization, pruning, sparsity, and distillation.
  • •Conduct system-level debugging, performance tuning, and performance-accuracy trade-off analysis; establish engineering guidelines to ensure production readiness of models and inference backends.

Key Requirements

  • •5+ years of experience with a Bachelor’s, Master’s, or PhD in Computer Science, Software Engineering, Mathematics, or related field (or equivalent experience).
  • •Strong C++ programming and debugging skills with a foundation in data structures, algorithms, and machine learning.
  • •Proven experience building and optimizing AI inference pipelines/applications using Llama.cpp, vLLM, PyTorch, Windows ML, DXCGC, and TensorRT.
  • •Deep understanding of inference backends and runtime internals (scheduling, memory management, KV-cache behavior, graph execution, quantization, hardware-aware optimization).
  • •Strong analytical/problem-solving skills and excellent written/verbal communication for cross-team collaboration.
Experience:5+ yearsAIMachine learningDeep learningGenerative AIInferenceEdge devices
Education:Bachelor's in Computer Science, Software Engineering, Mathematics (or related field)
Skills:ProgrammingDebuggingAnalytical thinkingProblem-solvingCommunication
Tech Stack:C++Llama.cppVLLMPyTorchWinMLWindows MLDXCGCTensorRTTensorRT-RTXCUDAVulkanDirectXRTXDGXKV-cacheGraph executionQuantizationPruningSparsityDistillation

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor