Software Engineer - GPU Kernels

Baseten
San Francisco, New York, Toronto, Canada, Montreal
Workplace: OnsiteFull timeUSD 185,000 - 250,000 annuallyFunction: Healthcare (Clinical, Medical, Wellness)Experience: 1-5 yearsEducation: bachelorsSkills: ["Problem-solving","Communication","Collaboration"]

Design and optimize high-performance GPU kernels for core ML operations, including matrix multiplications, attention, and mixture-of-experts routing. Implement CUDA/C++/PTX with architecture-aware techniques, optimize memory usage and overlap, and collaborate with research teams to productionize advancements for Baseten’s Model Performance platform serving millions of users.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Baseten
Baseten
1 year ago

Software Engineer - GPU Kernels

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 8 hours agoStatus: Live

Job Summary

Design and optimize high-performance GPU kernels for core ML operations, including matrix multiplications, attention, and mixture-of-experts routing. Implement CUDA/C++/PTX with architecture-aware techniques, optimize memory usage and overlap, and collaborate with research teams to productionize advancements for Baseten’s Model Performance platform serving millions of users.
Location: San Francisco, New York, Toronto, Canada, Montreal
Workplace: Onsite
Employment Type: Full time
Job Function: Healthcare (Clinical, Medical, Wellness)

Key Responsibilities

  • •Design and implement high-performance GPU kernels for key ML operations, including matrix multiplications, attention mechanisms, and mixture-of-experts routing.
  • •Write and optimize code using CUDA, PTX assembly, and architecture-specific techniques.
  • •Apply advanced performance optimization methods such as memory coalescing, warp-level programming, tensor core acceleration, and compute/memory overlap.
  • •Identify and resolve performance bottlenecks using profiling tools like Nsight Systems, Nsight Compute, and Torch Profiler.
  • •Collaborate with research teams to productionize theoretical advancements.

Pay and Benefits

Salary: USD 185,000 - 250,000 annually
Equity and Bonus:Equity
Perks:Health InsuranceDentalVisionPaid Parental401kEquity

Key Requirements

  • •1–5 years of experience in CUDA development.
  • •Strong understanding of GPU architecture and programming paradigms (memory hierarchy, thread/block/grid, synchronization).
  • •Proficient in C++ and GPU performance profiling tools.
  • •Knowledge of CUDA C++ API, memory access patterns, bandwidth optimization, numerical precision and quantization, and modern GPU features.
  • •Experience with modern GPU features (e.g., tensor cores, async operations).
Experience:1-5 yearsAIMachine learningHigh-performance computing
Education:Bachelor's
Skills:Problem-solvingCommunicationCollaboration
Languages:English
Tech Stack:CUDAC++PTXNsight SystemsNsight ComputeTorch ProfilerCUDA C++ APIMemory access patternsTensor coresCutlassTritonThrustCUB

Company Brief

Baseten
Baseten provides an inference-first ML infrastructure platform that lets engineering and ML teams deploy, serve, and scale machine-learning models with optimized performance, autoscaling, and GPU-backed hosting for production AI applications. ([crunchbase.com](https://www.crunchbase.com/organization/baseten?utm_source=openai))
Industry: AI & Machine Learning
Company Size: Medium (51 to 250 employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series E+
Headquarters: San Francisco, United States
Founded: 2019
Glassdoor
Glassdoor: 5.0
WebsiteLinkedInGlassdoor