Research Engineer, Infrastructure, Kernels

Thinking Machines Lab
San Francisco
Workplace: HybridFull timeUSD 350,000 - 475,000 annuallyFunction: Research & Scientific (R&D)Education: bachelorsSkills: ["Debugging","Collaboration","Initiative","Performance profiling","Technical communication"]

Design and optimize high-performance ML kernels that power large-scale language model training, including attention, matrix multiplication, gating, and normalization. Build compute primitives to reduce memory bottlenecks and improve compute efficiency, then profile performance across GPU/accelerator generations. Collaborate with research teams and systems architects to align kernel optimizations with model goals, maintain reusable kernel libraries and benchmarks, and help ensure scalability, stability, and reproducibility. Document and share technical insights via papers and talks.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Thinking Machines Lab
Thinking Machines Lab
1 day ago

Research Engineer, Infrastructure, Kernels

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 9 hours agoStatus: Live

Job Summary

Design and optimize high-performance ML kernels that power large-scale language model training, including attention, matrix multiplication, gating, and normalization. Build compute primitives to reduce memory bottlenecks and improve compute efficiency, then profile performance across GPU/accelerator generations. Collaborate with research teams and systems architects to align kernel optimizations with model goals, maintain reusable kernel libraries and benchmarks, and help ensure scalability, stability, and reproducibility. Document and share technical insights via papers and talks.
Location: San Francisco
Workplace: Hybrid
Employment Type: Full time
Job Function: Research & Scientific (R&D)

Key Responsibilities

  • •Design and implement custom ML kernels for core LLM operations (attention, matrix multiplication, gating, normalization) optimized for GPU/accelerator architectures.
  • •Create compute primitives to reduce memory bandwidth bottlenecks and improve kernel compute efficiency.
  • •Collaborate with research teams to align kernel-level optimizations with model architecture and algorithmic goals.
  • •Develop and maintain reusable kernel libraries and performance benchmarks for internal model training.
  • •Contribute to infrastructure stability and scalability, ensuring reproducibility, consistent low-precision behavior, and high compute utilization.

Pay and Benefits

Salary: USD 350,000 - 475,000 annually
Perks:Health InsuranceDentalVisionPaid LeaveRelocation

Key Requirements

  • •Bachelor’s degree or equivalent experience in computer science, electrical engineering, statistics, machine learning, physics, robotics, or related fields.
  • •Strong engineering skills to build performant, maintainable code and debug complex codebases.
  • •Understanding of deep learning frameworks (e.g., PyTorch, JAX) and their underlying system architectures.
  • •Proficiency in CUDA, CuTe, Triton, or other GPU programming frameworks.
  • •Ability to analyze, profile, and optimize compute-intensive workloads.
Experience:Machine learningDeep learningLarge-scale language modelsML systemsGPU programming
Education:Bachelor's
Skills:DebuggingCollaborationInitiativePerformance profilingTechnical communication
Tech Stack:CUDACuTeTritonPyTorchJAXFP8INT8XLATVM

Company Brief

Thinking Machines Lab
Develops enterprise AI solutions, custom large language models, and ML platforms to help organizations deploy intelligent applications. Services include data engineering, model development, and AI consulting for scale and production readiness.
Industry: AI & Machine Learning
Company Size: Medium (51 to 250 employees)
Growth: Growth Stage Startup
Headquarters: Mumbai, India
Website