Software Engineer, CUDA Deep Learning Systems

NVIDIA
Santa Clara, Austin
Workplace: OnsiteFull timeUSD 124,000 - 195,500 annuallyFunction: Software EngineeringEducation: bachelorsSkills: ["Collaboration","Initiative","Problem-solving"]

Design and prototype next-generation deep learning systems at the intersection of CUDA and large-scale distributed computing. Build and optimize distributed architectures from single-node to cluster-scale supercomputing, develop custom high-performance CUDA kernels, and diagnose hardware-software bottlenecks in training and inference. Partner with AI, hardware, and compiler experts to co-design accelerator performance improvements, and create profiling tools and runtime systems that enable emerging AI workloads.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
2 hours ago

Software Engineer, CUDA Deep Learning Systems

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 2 hours agoStatus: Live

Job Summary

Design and prototype next-generation deep learning systems at the intersection of CUDA and large-scale distributed computing. Build and optimize distributed architectures from single-node to cluster-scale supercomputing, develop custom high-performance CUDA kernels, and diagnose hardware-software bottlenecks in training and inference. Partner with AI, hardware, and compiler experts to co-design accelerator performance improvements, and create profiling tools and runtime systems that enable emerging AI workloads.
Location: Santa Clara, Austin
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Explore, research, and prototype novel deep learning systems optimizations by modeling and simulation across high-level DL frameworks and low-level CUDA.
  • •Architect and optimize distributed computing systems to scale from single-node to massive cluster-scale environments.
  • •Design, implement, and optimize custom high-performance CUDA kernels tailored to emerging neural network workloads.
  • •Analyze hardware-software interactions to identify and resolve performance bottlenecks in training and inference pipelines.
  • •Develop profiling tools and runtime systems, and collaborate with AI researchers and experts in hardware/software/compilers to co-design accelerator performance improvements.

Pay and Benefits

Salary: USD 124,000 - 195,500 annually
Equity and Bonus:Equity

Key Requirements

  • •BS, MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or related field (or equivalent experience).
  • •2+ years of relevant industry experience (or equivalent academic experience after the degree).
  • •Strong proficiency in C++ and Python programming.
  • •Solid fundamentals of deep learning (with a focus on transformers) and experience profiling/optimizing generative AI models.
  • •Experience in systems programming, computer architecture, and low-level performance optimization, including hands-on CUDA programming and kernel optimization.
Experience:Deep learningGenerative AIDistributed computingSystems programming
Education:Bachelor's
Skills:CollaborationInitiativeProblem-solving
Tech Stack:CUDAC++PythonPyTorchJAXTensorRTVLLMSgLangNemoMegatronNCCLMPIUCXTritonXLATorch.compileNVFP4MXFP4FP8INT8

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor