Senior Software Engineer, CUDA Deep Learning Systems

NVIDIA
Santa Clara, Austin
Workplace: OnsiteFull timeUSD 184,000 - 356,500 annuallyFunction: Software EngineeringExperience: 8+ yearsEducation: mastersSkills: ["Collaboration","Initiative"]

Build and prototype performance-optimized deep learning systems at the intersection of CUDA and distributed training. You’ll design and optimize custom CUDA kernels, architect scalable multi-node computing, and analyze hardware-software bottlenecks across training and inference. Collaborate with AI researchers, hardware/software architects, and compiler and driver experts to improve accelerator utilization and memory and networking efficiency, while developing tooling and writing production-quality, maintainable code.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
2 hours ago

Senior Software Engineer, CUDA Deep Learning Systems

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 2 hours agoStatus: Live

Job Summary

Build and prototype performance-optimized deep learning systems at the intersection of CUDA and distributed training. You’ll design and optimize custom CUDA kernels, architect scalable multi-node computing, and analyze hardware-software bottlenecks across training and inference. Collaborate with AI researchers, hardware/software architects, and compiler and driver experts to improve accelerator utilization and memory and networking efficiency, while developing tooling and writing production-quality, maintainable code.
Location: Santa Clara, Austin
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Explore, research, and prototype systems optimizations for advanced deep learning models using modeling, simulation, and silicon prototyping.
  • •Architect and optimize distributed computing systems that scale from single-node to cluster-scale supercomputing environments.
  • •Design, implement, and optimize custom high-performance CUDA kernels for emerging neural network architectures and workloads.
  • •Analyze hardware-software interactions to identify and resolve performance bottlenecks in training and inference pipelines.
  • •Collaborate with AI researchers, hardware and software architects, kernel/compiler authors, and CUDA driver experts to co-design systems and algorithms; develop profiling and runtime tools.

Pay and Benefits

Salary: USD 184,000 - 356,500 annually
Equity and Bonus:Equity

Key Requirements

  • •BS, MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or a related field (or equivalent experience).
  • •8+ years of relevant industry experience (or equivalent academic experience after degree achievement).
  • •Strong proficiency in C++ and Python.
  • •Solid deep learning fundamentals with a focus on transformers.
  • •Experience with systems programming, computer architecture, and low-level performance optimization including CUDA kernel development and profiling.
Experience:8+ yearsDeep learningAI systemsDistributed computingGenerative AI
Education:Master's
Skills:CollaborationInitiative
Tech Stack:CUDAC++PythonPyTorchJAXTensorRTVLLMSgLangNemoMegatronNCCLMPIUCXTritonXLATorch.compileNVFP4MXFP4FP8INT8

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor