Senior Software Engineer - GPU Local AI Platforms

NVIDIA
United States
Workplace: OnsiteFull timeUSD 224,000 - 431,250 annuallyFunction: Software EngineeringExperience: 12+ yearsEducation: bachelorsSkills: ["Strong analytical skills","Ability to form hypotheses","Experiment design","Clear communication"]

Build and optimize NVIDIA’s Local AI software stack for running large language models efficiently on edge AI hardware. You’ll track open-source inference frameworks, map model architectures and algorithms to GPU architecture, analyze multi-node inference performance, and produce validation workflows and developer-facing inference recipes that stay current as frameworks evolve.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
1 month ago

Senior Software Engineer - GPU Local AI Platforms

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 2 hours agoStatus: Live

Job Summary

Build and optimize NVIDIA’s Local AI software stack for running large language models efficiently on edge AI hardware. You’ll track open-source inference frameworks, map model architectures and algorithms to GPU architecture, analyze multi-node inference performance, and produce validation workflows and developer-facing inference recipes that stay current as frameworks evolve.
Location: United States
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Evaluate and track innovations in open-source LLM inference frameworks and identify performance-critical features for NVIDIA edge AI hardware.
  • •Analyze how new model architectures and inference algorithms map to NVIDIA GPU architecture, identifying optimization opportunities and fallback paths.
  • •Characterize multi-node inference behavior, including NCCL/RCCL primitives, topology-aware all-reduce strategies, and parallelism efficiency.
  • •Produce performance analysis reports tying theoretical hardware limits to observed throughput, latency, and utilization.
  • •Own model validation and developer inference recipe workflows for new model releases, including architecture compatibility assessment and performance characterization.

Pay and Benefits

Salary: USD 224,000 - 431,250 annually
Equity and Bonus:Equity

Key Requirements

  • •12+ years of software engineering with depth in GPU computing, ML systems, or high-performance inference.
  • •BS, MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or equivalent experience.
  • •Strong Python or C++ programming, software design, and software engineering skills.
  • •Hands-on GPU kernel development or optimization (CUDA/C++, Triton, or equivalent), including understanding of thread blocks and warp execution.
  • •Working knowledge of LLM inference internals such as attention mechanisms, KV-cache management, continuous batching, quantization formats, and tensor parallelism.
Experience:12+ yearsGPU computingML systemsHigh-performance inferenceOpen-sourceLLM inference
Education:Bachelor's in Computer Science, Computer Engineering, Electrical Engineering
Skills:Strong analytical skillsAbility to form hypothesesExperiment designClear communication
Tech Stack:PythonC++CUDATritonDockerOCINVIDIA Container ToolkitNCCLRCCLCI/CDQuantizationKV-cacheTensor parallelismContinuous batching

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor