Senior Deep Learning Compiler Engineer - XLA

NVIDIA
Santa Clara, Austin, Redmond
Workplace: RemoteFull timeUSD 152,000 - 287,500 annuallyFunction: Data Science & Machine LearningExperience: 4+ yearsSkills: ["Independent work","Performance analysis","Debugging","Software design","Interpersonal skills"]

Develop compiler optimization algorithms for deep learning workloads, focusing on inference and training performance for the JAX framework and the OpenXLA compiler on NVIDIA GPUs. Collaborate with deep-learning framework teams and GPU hardware architecture to improve performance through graph transformations, partitioning, tensor sharding, tuning, and analysis. Build production-grade features, including code-generation for GPU backends using tools like MLIR and LLVM.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
2 days ago

Senior Deep Learning Compiler Engineer - XLA

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 23 hours agoStatus: Live
Reposted: similar role first listed 5 months ago

Job Summary

Develop compiler optimization algorithms for deep learning workloads, focusing on inference and training performance for the JAX framework and the OpenXLA compiler on NVIDIA GPUs. Collaborate with deep-learning framework teams and GPU hardware architecture to improve performance through graph transformations, partitioning, tensor sharding, tuning, and analysis. Build production-grade features, including code-generation for GPU backends using tools like MLIR and LLVM.
Location: Santa Clara, Austin, Redmond
Workplace: Remote
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Mid level

Key Responsibilities

  • •Develop compiler optimization algorithms for deep learning workloads, improving inference and training performance on NVIDIA GPUs.
  • •Implement compiler techniques for deep learning network graphs, including graph partitioning and tensor sharding for distributed training and inference.
  • •Perform performance tuning, analysis, and debugging to evaluate improvements at scale.
  • •Create GPU backend code generation using open-source compilers such as MLIR and LLVM (including OpenAI Triton).
  • •Collaborate with deep-learning framework and GPU hardware teams to design and deliver AI compiler features for next-generation GPUs.

Pay and Benefits

Salary: USD 152,000 - 287,500 annually
Perks:Equity

Key Requirements

  • •Bachelors, Masters, or Ph.D. in Computer Science, Computer Engineering, or a related field (or equivalent experience).
  • •4+ years of relevant work or research experience in performance analysis and compiler optimizations.
  • •Excellent C/C++ programming and software design skills, including debugging, performance analysis, and test design.
  • •Strong foundation in CPU/GPU or high-performance accelerator architecture, including knowledge of high-performance computing and distributed programming.
  • •Experience with XLA, TVM, MLIR, LLVM, and/or OpenAI Triton is a strong plus.
Experience:4+ yearsDeep learning
Education:
Skills:Independent workPerformance analysisDebuggingSoftware designInterpersonal skills
Tech Stack:JAXOpenXLAXLAOpenAI TritonMLIRLLVMTritonTVMPyTorchTensorFlowCUDAOpenCLCUDA programmingGPU programmingC/C++

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor