Senior Software Engineer, Matrix Multiplication

NVIDIA
Santa Clara
Workplace: OnsiteFull timeUSD 184,000 - 287,500 annuallyFunction: Software EngineeringExperience: 6+ yearsEducation: mastersSkills: ["Python","C++","CUDA","Triton","CuTile","PyTorch","JAX","TensorFlow","ONNX"]

Develop groundbreaking AI inference software by building libraries, code generators, and GPU kernel technologies for NVIDIA’s hardware. You will design high-performance kernels, create extensible abstractions for LLM serving engines, and contribute to open-source projects, collaborating across frameworks, kernels, and GPU teams to accelerate AI workloads.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
4 months ago

Senior Software Engineer, Matrix Multiplication

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

Develop groundbreaking AI inference software by building libraries, code generators, and GPU kernel technologies for NVIDIA’s hardware. You will design high-performance kernels, create extensible abstractions for LLM serving engines, and contribute to open-source projects, collaborating across frameworks, kernels, and GPU teams to accelerate AI workloads.
Location: Santa Clara
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Innovating and developing new AI systems technologies for efficient inference
  • •Designing, implementing, and optimizing kernels for high impact AI workloads
  • •Designing and implementing extensible abstractions for LLM serving engines
  • •Building efficient just-in-time domain specific compilers and runtimes
  • •Collaborating closely with other engineers at NVIDIA across deep learning frameworks, libraries, kernels, and GPU arch teams

Pay and Benefits

Salary: USD 184,000 - 287,500 annually
Equity and Bonus:Equity
Perks:EquityBenefits

Key Requirements

  • •Master's degree in Computer Science, Electrical Engineering, or related field (or equivalent experience); PhD preferred
  • •6+ years (academic/ industry) experience with ML/DL systems development
  • •Strong experience in developing or using deep learning frameworks (e.g. PyTorch, JAX, TensorFlow, ONNX) and inference engines/runtimes (e.g., vLLM, SGLang, MLC)
  • •Strong Python and C/C++ programming skills
  • •Strong experience in GPU kernel development and performance optimizations (CUDA C/C++, cuTile, Triton)
Experience:6+ yearsAIMLDLInferenceLLM
Education:Master's
Skills:PythonC++CUDATritonCuTilePyTorchJAXTensorFlowONNX
Languages:English
Tech Stack:PythonC++CUDATritonCuTilePyTorchJAXTensorFlowONNX

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor