Deep Learning Performance Architect

NVIDIA
Shanghai, Beijing
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningEducation: mastersSkills: ["Communication","Cross-collaboration","Customer-oriented","Creative","Autonomous"]

Build and optimize GPU-accelerated deep learning inference software, developing highly optimized deep learning kernels. Conduct performance analysis, profiling, debugging, and tuning to improve runtime efficiency, and apply architectural knowledge of CPU and GPU. Collaborate with cross-functional teams across automotive, image understanding, and speech understanding, translating the latest algorithms for public release in TensorRT. Occasionally travel to support customers and conference training.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
2 weeks ago

Deep Learning Performance Architect

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 3 hours agoStatus: Live
Reposted: similar role first listed 7 months ago

Job Summary

Build and optimize GPU-accelerated deep learning inference software, developing highly optimized deep learning kernels. Conduct performance analysis, profiling, debugging, and tuning to improve runtime efficiency, and apply architectural knowledge of CPU and GPU. Collaborate with cross-functional teams across automotive, image understanding, and speech understanding, translating the latest algorithms for public release in TensorRT. Occasionally travel to support customers and conference training.
Location: Shanghai, Beijing
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Mid level

Key Responsibilities

  • •Develop highly optimized deep learning kernels for inference.
  • •Perform performance optimization, analysis, and tuning.
  • •Work with cross-collaborative teams across automotive, image understanding, and speech understanding to develop innovative solutions.
  • •Collaborate with the deep learning community to implement the latest algorithms for public release in TensorRT.
  • •Occasionally travel to conferences and customers for technical consultation and training.
Travel: Low travel

Key Requirements

  • •Masters or PhD (or equivalent) in a relevant discipline (CE, CS&E, CS, or AI).
  • •5 years of relevant work experience.
  • •Excellent C/C++ programming and software design skills.
  • •Performance modelling, profiling, debug, and code optimization (or architectural knowledge of CPU and GPU).
  • •GPU programming experience (CUDA or OpenCL) desired; Python experience and SW Agile skills are helpful.
Experience:Deep learningGPU computing
Education:Master's
Skills:CommunicationCross-collaborationCustomer-orientedCreativeAutonomous
Tech Stack:CC++PythonTensorRTCUDAOpenCLGPUCPU

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor