Deep Learning Compiler Engineer - CUDA

NVIDIA
Shanghai, Beijing
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningExperience: 2+ yearsEducation: mastersSkills: ["Problem-solving","Oral communication"]

Design and implement the DSL and core compiler for a tile-aware GPU programming model, targeting emerging GPU architectures. Drive continuous compiler architecture innovations to optimize performance, and investigate next-generation GPU architectures to deliver solutions across the DSL and compiler stack. Perform performance analysis for AI/LLM workloads and integrate with AI/ML frameworks, collaborating closely with architecture and engineering teams.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
1 month ago

Deep Learning Compiler Engineer - CUDA

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 11 hours agoStatus: Live
Reposted: similar role first listed 7 months ago

Job Summary

Design and implement the DSL and core compiler for a tile-aware GPU programming model, targeting emerging GPU architectures. Drive continuous compiler architecture innovations to optimize performance, and investigate next-generation GPU architectures to deliver solutions across the DSL and compiler stack. Perform performance analysis for AI/LLM workloads and integrate with AI/ML frameworks, collaborating closely with architecture and engineering teams.
Location: Shanghai, Beijing
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Mid level

Key Responsibilities

  • •Design and implement the DSL and core compiler for a tile-aware GPU programming model.
  • •Innovate and iterate on the compiler’s core architecture to consistently optimize performance.
  • •Investigate next-generation GPU architectures and provide solutions within the DSL/compiler stack.
  • •Analyze performance on emerging AI/LLM workloads and integrate with AI/ML frameworks.

Key Requirements

  • •Master’s or PhD (or equivalent) in a relevant discipline such as CE, CS&E, CS, or AI.
  • •2+ years of relevant work experience.
  • •Excellent C/C++ programming and software engineering skills.
  • •Strong fundamental knowledge of computer architecture.
  • •Strong compiler background, including MLIR/TVM/Triton/LLVM (desired).
Experience:2+ yearsAI/LLMHPCCompiler
Education:Master's
Skills:Problem-solvingOral communication
Languages:English
Tech Stack:CUDACC++DSLMLIRTVMTritonLLVMGPU architectureAI/ML frameworksLLM algorithmsMulti-GPU distributed communicationKernel programmingHPC

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor