Senior AI Kernel Engineer

Modular
United States, Canada
Workplace: RemoteFull timeUSD 198,000 - 286,000 annuallyFunction: Data Science & Machine LearningExperience: 5+ yearsSkills: ["Problem-solving","Independent work","Mentoring","Performance tuning","Cross-functional collaboration"]

Lead the design and optimization of high-performance GPU kernels for large-scale AI inference on GPUs and emerging custom accelerators. Own performance-critical code paths, drive architectural decisions, and translate complex workloads into efficient production implementations across single- and multi-GPU, heterogeneous environments. Collaborate with compiler, runtime, model, and infrastructure teams using profilers and microbenchmarks, and mentor engineers to improve kernel fusion, scheduling, and long-term inference performance.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Modular
Modular
7 months ago

Senior AI Kernel Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: Just nowStatus: Live

Job Summary

Lead the design and optimization of high-performance GPU kernels for large-scale AI inference on GPUs and emerging custom accelerators. Own performance-critical code paths, drive architectural decisions, and translate complex workloads into efficient production implementations across single- and multi-GPU, heterogeneous environments. Collaborate with compiler, runtime, model, and infrastructure teams using profilers and microbenchmarks, and mentor engineers to improve kernel fusion, scheduling, and long-term inference performance.
Location: United States, Canada
Workplace: Remote
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Mid level

Key Responsibilities

  • •Design, implement, and optimize performance-critical kernels for AI inference workloads (e.g., GEMM, attention, communication, fusion).
  • •Lead kernel-level optimization across single-GPU, multi-GPU, and heterogeneous hardware environments.
  • •Make trade-offs between latency, throughput, memory footprint, and numerical precision.
  • •Analyze performance using profilers, hardware counters, and microbenchmarks; translate insights into improvements.
  • •Collaborate with compiler and runtime teams to influence code generation, scheduling, and kernel fusion; mentor engineers and contribute to performance roadmaps.
Travel: Low travel

Pay and Benefits

Salary: USD 198,000 - 286,000 annually
Equity and Bonus:Equity
Perks:Health Insurance401kPaid LeaveEquity

Key Requirements

  • •5+ years of experience in performance-critical systems or kernel development (or equivalent depth of expertise).
  • •Strong proficiency in C/C++ and low-level programming.
  • •Extensive hands-on GPU kernel programming experience (CUDA, HIP, or equivalent).
  • •Deep understanding of GPU architecture, including memory hierarchies, synchronization, and execution models.
  • •Proven track record delivering measurable performance improvements in production systems.
Experience:5+ yearsAI inferenceGPU kernel developmentLow-level systemsPerformance optimization
Skills:Problem-solvingIndependent workMentoringPerformance tuningCross-functional collaboration
Tech Stack:C/C++CUDAHIPGPU kernel programmingPTXAssembly-level tuningTritonTensor CoresGEMMTransformersLLMsDiffusionProfilerHardware countersMicrobenchmarks

Company Brief

Modular
Provides infrastructure and developer tools to build, deploy, and scale large AI models and foundation-model applications, including model hosting, orchestration, and SDKs to accelerate AI product development.
Industry: AI & Machine Learning
Website