Deep Learning Kernel Software Performance Architect
NVIDIA
Shanghai, Beijing
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningExperience: 2+ yearsEducation: mastersSkills: ["Communication","Collaboration","Problem-solving","Debugging","Cross-team coordination"]Optimize GPU kernel performance for modern data-center deep learning workloads by building automated, data-driven workflows that detect, explain, and prevent performance regressions. You’ll perform end-to-end performance analysis, debugging, and optimization, while developing Python-heavy automation and scalable regression infrastructure. Partner with kernel, compiler, chip architecture, SWQA, and infrastructure teams to align performance checks with release needs and improve throughput for key operators like GEMM, Attention, and MoE.

