System Software Engineer - Local AI

NVIDIA
Pune
Full timeFunction: Software EngineeringExperience: 2+ yearsEducation: bachelorsSkills: ["Analytical skills","Problem-solving","Communication","Debugging","Multitasking"]

Build efficient on-device AI software for RTX and DGX systems, focusing on low-latency, high-performance local inference. You’ll develop and optimize inference runtimes and execution stacks, implement end-to-end optimization for model pipelines, and apply techniques like quantization and pruning. Partner across software, research, architecture, and product teams to ensure performance, stability, and production readiness across current and next-generation GPU architectures.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
16 hours ago

System Software Engineer - Local AI

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 16 hours agoStatus: Live

Job Summary

Build efficient on-device AI software for RTX and DGX systems, focusing on low-latency, high-performance local inference. You’ll develop and optimize inference runtimes and execution stacks, implement end-to-end optimization for model pipelines, and apply techniques like quantization and pruning. Partner across software, research, architecture, and product teams to ensure performance, stability, and production readiness across current and next-generation GPU architectures.
Location: Pune
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Partner with NVIDIA software, research, architecture, and product teams to align strategies for the LocalAI ecosystem across RTX and DGX systems.
  • •Build and optimize local AI inference stacks for RTX, RTX Pro, and DGX GPUs with a focus on performance, stability, and scalability.
  • •Architect and develop modern inference runtimes and execution stacks for workloads including LLMs, vision-language, TTS, ASR, and diffusion AI.
  • •Perform end-to-end optimization of AI models, data pipelines, and inference runtimes using techniques such as quantization, pruning, sparsity, and distillation.
  • •Conduct system-level debugging and performance–accuracy trade-off analysis, develop infrastructure for performance sweeps, and establish guidelines for production readiness of new models and backends.

Key Requirements

  • •2+ years of experience with a Bachelor’s, Master’s, or PhD in Computer Science, Software Engineering, Mathematics, or a related field, or equivalent experience.
  • •Excellent C++ programming and debugging skills, with strong knowledge of data structures and algorithms and machine learning.
  • •Proven experience working with AI inferencing pipelines and applications using ML/DL frameworks such as Llama.cpp, vLLM, PyTorch, WinML, DXCGC, and TensorRT.
  • •Deep interest in inference backends and runtime internals, including scheduling, memory management, KV-cache behavior, graph execution, quantization, and hardware-aware optimization.
  • •Strong analytical and problem-solving abilities with effective multitasking and outstanding written and oral communication.
Experience:2+ yearsArtificial intelligenceMachine learningOpen source
Education:Bachelor's in Computer Science, Software Engineering, Mathematics, or a related field
Skills:Analytical skillsProblem-solvingCommunicationDebuggingMultitasking
Tech Stack:C++Llama.cppVLLMPyTorchWinMLDXCGCTensorRT-RTXCUDAVulkanDirectXRTXDGX

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor