Senior Inference Engineer, GPU Kernel Optimization

NVIDIA
United States
Workplace: OnsiteFull timeUSD 184,000 - 287,500 annuallyFunction: Solutions Engineering & Sales EngineeringEducation: mastersSkills: ["Collaboration"]

Drive performance to the ceiling for LLM inference by building silicon-measured GPU kernel microbenchmarking, end-to-end model performance analysis, and agentic kernel optimization systems. Collaborate across compiler, kernel, hardware, and framework teams to attribute bottlenecks, generate optimization policies, and validate improvements with rigorous CUPTI/NSYS/NCU profiling for production-grade throughput and latency gains.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
1 month ago

Senior Inference Engineer, GPU Kernel Optimization

âś“ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 16 hours agoStatus: Live

Job Summary

Drive performance to the ceiling for LLM inference by building silicon-measured GPU kernel microbenchmarking, end-to-end model performance analysis, and agentic kernel optimization systems. Collaborate across compiler, kernel, hardware, and framework teams to attribute bottlenecks, generate optimization policies, and validate improvements with rigorous CUPTI/NSYS/NCU profiling for production-grade throughput and latency gains.
Location: United States
Workplace: Onsite
Employment Type: Full time
Job Function: Solutions Engineering & Sales Engineering
Seniority: Mid level

Key Responsibilities

  • •Build GPU kernel microbenchmarking to measure competing kernel implementations with real-silicon fidelity across production configuration space.
  • •Perform end-to-end model performance analysis to connect performance evidence to model-level serving economics and surface high-value optimization opportunities.
  • •Develop agentic kernel optimization systems that diagnose performance gaps, explore optimizations across the kernel ecosystem, and validate with silicon measurements.
  • •Collaborate closely with compiler, hardware, kernel, and framework teams to deliver upstream improvements.
  • •Translate findings into production-grade performance gains for NVIDIA’s LLM inference stack.

Pay and Benefits

Salary: USD 184,000 - 287,500 annually
Equity and Bonus:Equity

Key Requirements

  • •Master's or PhD in Computer Science, Computer Engineering, or a related field, or equivalent experience.
  • •6+ years of relevant industry experience.
  • •Experience building or directing agentic AI systems (code generation, automated optimization, or multi-step reasoning workflows).
  • •Strong Python and C++ skills with proven software engineering fundamentals.
  • •Hands-on GPU profiling with CUPTI, NSYS, and NCU and ability to attribute bottlenecks across kernel execution, compiler decisions, and runtime scheduling.
Experience:LLM inferenceGPU performance engineeringAgentic AI
Education:Master's in Computer Science, Computer Engineering, or a related field
Skills:Collaboration
Tech Stack:PythonC++CUPTINSYSNCUCUDACUTLASSTritonPTXSASSTRT-LLMSGLangVLLMLLVMMLIRPtxasFlashInfer

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor