Senior Performance Software Engineer, Deep Learning Libraries

NVIDIA
Shanghai, Beijing
Workplace: OnsiteFull timeFunction: Software EngineeringExperience: 2+ yearsEducation: mastersSkills: ["Performance analysis","Debugging","Software design","Test design","Optimization"]

Develop optimized deep learning performance code for NVIDIA GPUs, writing highly tuned compute kernels that accelerate core linear algebra and operations used in modern deep learning. Build and improve high-performance code for cuDNN, cuBLAS, and TensorRT, collaborating with CUDA compiler, training/inference performance, and hardware/architecture teams. Focus on GPU efficiency, regression testing, CI/CD, and performance analysis down to GPU hardware.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
1 day ago

Senior Performance Software Engineer, Deep Learning Libraries

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 13 hours agoStatus: Live

Job Summary

Develop optimized deep learning performance code for NVIDIA GPUs, writing highly tuned compute kernels that accelerate core linear algebra and operations used in modern deep learning. Build and improve high-performance code for cuDNN, cuBLAS, and TensorRT, collaborating with CUDA compiler, training/inference performance, and hardware/architecture teams. Focus on GPU efficiency, regression testing, CI/CD, and performance analysis down to GPU hardware.
Location: Shanghai, Beijing
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Write highly tuned compute kernels for core deep learning operations such as matrix multiplies, convolutions, and normalizations.
  • •Follow software engineering best practices, including regression testing and CI/CD flows.
  • •Collaborate with NVIDIA teams including the CUDA compiler team to generate optimal assembly code.
  • •Work with deep learning training and inference performance teams to optimize layers.
  • •Analyze performance to identify bottlenecks, optimize resource utilization, and improve throughput.

Key Requirements

  • •Masters or PhD degree (or equivalent) in Computer Science, Computer Engineering, Applied Math, or a related field.
  • •2+ years of relevant industry experience.
  • •Strong C++ programming and software design skills, including debugging, performance analysis, and test design.
  • •Experience with performance-oriented parallel programming (e.g., OpenMP or pthreads).
  • •Solid understanding of computer architecture, with some assembly programming experience.
Experience:2+ yearsDeep learningGPU programming
Education:Master's in Computer Science, Computer Engineering, Applied Math, or related field
Skills:Performance analysisDebuggingSoftware designTest designOptimization
Tech Stack:C++CUDACuDNNCuBLASTensorRTCUTLASSTensor CoresOpenMPPthreadsAssemblyLLVMTVMTensorFlow MLIRCUDA compilerCI/CDRegression testing

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor