Deep Learning Compiler Engineer - CUDA
NVIDIA
Shanghai, Beijing
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningExperience: 2+ yearsEducation: mastersSkills: ["Problem-solving","Oral communication"]Design and implement the DSL and core compiler for a tile-aware GPU programming model, targeting emerging GPU architectures. Drive continuous compiler architecture innovations to optimize performance, and investigate next-generation GPU architectures to deliver solutions across the DSL and compiler stack. Perform performance analysis for AI/LLM workloads and integrate with AI/ML frameworks, collaborating closely with architecture and engineering teams.

