Senior Software Engineer - AI Inference Performance

NVIDIA
Santa Clara
Workplace: OnsiteFull timeUSD 184,000 - 356,500 annuallyFunction: Software EngineeringExperience: 6+ yearsEducation: mastersSkills: ["Hands-on","Analytical","Collaborative","Autonomous","Creative"]

Advance LLM/VLM inference performance on NVIDIA GPU-accelerated systems by turning profiler and performance models into production-ready code. Lead end-to-end analysis of prefill/decode workloads, optimize latency and throughput metrics, and benchmark against reproducible regression gates. Profile with NVIDIA Nsight tools and PyTorch Profiler, tune serving and KV-cache strategies, and build performance-critical CUDA/CUTLASS/Triton kernels. Collaborate across teams and improve open-source inference engines.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
2 days ago

Senior Software Engineer - AI Inference Performance

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 6 hours agoStatus: Live

Job Summary

Advance LLM/VLM inference performance on NVIDIA GPU-accelerated systems by turning profiler and performance models into production-ready code. Lead end-to-end analysis of prefill/decode workloads, optimize latency and throughput metrics, and benchmark against reproducible regression gates. Profile with NVIDIA Nsight tools and PyTorch Profiler, tune serving and KV-cache strategies, and build performance-critical CUDA/CUTLASS/Triton kernels. Collaborate across teams and improve open-source inference engines.
Location: Santa Clara
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Lead end-to-end analysis of LLM/VLM inference processes, defining representative prefill and decode workloads and optimizing latency, throughput, and efficiency metrics.
  • •Build and apply speed-of-light and roofline models to identify performance headroom and connect hardware arithmetic intensity, bandwidth, occupancy, memory hierarchy, and communication costs to optimization hypotheses.
  • •Profile and eliminate bottlenecks across host code, CUDA kernels, memory, communication, and scheduling using Nsight Systems/Compute and PyTorch Profiler plus custom instrumentation.
  • •Tune serving hyperparameters and techniques including batching, KV-cache management, quantization, speculative decoding, CUDA Graphs, and model parallelism based on workload and service objectives.
  • •Develop and optimize performance-critical kernels and establish repeatable benchmarks and performance regression gates, including contributions to inference engines such as TensorRT-LLM, vLLM, and SGLang.

Pay and Benefits

Salary: USD 184,000 - 356,500 annually
Equity and Bonus:Equity

Key Requirements

  • •More than 6 years of experience in full-stack LLM/VLM inference performance across models, serving, distributed runtimes, kernels, and hardware, resulting in measurable gains in production or production-representative environments.
  • •Strong programming skills in Python, Rust and/or C++, with hands-on experience with CUDA or another GPU programming environment.
  • •Expertise in speed-of-light analysis, roofline models, microbenchmarks, and tools including NVIDIA Nsight Systems and Nsight Compute, converting profiles into testable hypotheses.
  • •Deep understanding of GPU architecture, including Tensor Cores, memory hierarchy, caches, occupancy, synchronization, and numerical formats across hardware generations.
  • •BS or MS in Computer Science, Computer Engineering, or a related field, or equivalent experience.
Experience:6+ yearsLLM inferenceVLM inferenceGPU accelerationDistributed systemsInference performanceOpen-source inference
Education:Master's in Computer Science, Computer Engineering, or a related field
Skills:Hands-onAnalyticalCollaborativeAutonomousCreative
Tech Stack:PythonRustC++CUDACUDA kernelsNVIDIA GPUNVIDIA Nsight SystemsNVIDIA Nsight ComputePyTorch ProfilerPyTorchCUDA GraphsCUTLASSTritonTensorRT-LLMVLLMSGLangNCCLKV cacheMicrobenchmarksRoofline models

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor