Member of Technical Staff, Kernels

Inception Labs
San Mateo
Workplace: OnsiteFull timeFunction: Healthcare (Clinical, Medical, Wellness)Skills: ["Problem-solving","Collaboration"]

Design and optimize high-performance ML kernels and the distributed compute stack to power large-scale language model training and inference. Develop CUDA-based kernels, enable low-precision arithmetic, and improve kernel efficiency and infrastructure stability for scalable, reproducible ML systems on modern GPUs.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Inception Labs
Inception Labs
6 months ago

Member of Technical Staff, Kernels

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 3 hours agoStatus: Live

Job Summary

Design and optimize high-performance ML kernels and the distributed compute stack to power large-scale language model training and inference. Develop CUDA-based kernels, enable low-precision arithmetic, and improve kernel efficiency and infrastructure stability for scalable, reproducible ML systems on modern GPUs.
Location: San Mateo
Workplace: Onsite
Employment Type: Full time
Job Function: Healthcare (Clinical, Medical, Wellness)

Key Responsibilities

  • •Design and implement custom ML kernels (CUDA, CuTe, Triton) for core dLLM operations such as attention, matrix multiplication, gating, and normalization, optimized for modern GPU architectures.
  • •Design compute primitives to reduce memory bandwidth bottlenecks and improve kernel efficiency.
  • •Contribute to infrastructure stability and scalability, ensuring reproducibility, consistency across precision formats, and high utilization of compute resources.
  • •Optimize kernels for high utilization of compute resources on distributed GPU clusters.
  • •Collaborate on scaling compute foundations powering large-scale language model training and inference.

Pay and Benefits

Perks:EquityHealth InsuranceDentalVisionCommuter BenefitsMeal Allowance

Key Requirements

  • •BS/MS/PhD in Computer Science, Engineering, or a related field (or equivalent experience).
  • •Proficiency in CUDA, CuTe, Triton, or other GPU programming frameworks.
  • •Understanding of ML frameworks (PyTorch, TensorFlow) from a systems perspective.
  • •Background in performance optimization and profiling of ML systems.
  • •Experience implementing low-precision formats (FP8, INT8, block floating point) or contributing to related compiler stacks (XLA, TVM).
Experience:AI/ML infrastructureGPU computingHigh-performance computing
Skills:Problem-solvingCollaboration
Tech Stack:CUDACuTeTritonPyTorchTensorFlowFP8INT8Block floating pointXLATVMDockerKubernetesPythonC++RustGo

Company Brief

Inception Labs
Develops artificial intelligence solutions and research-driven products, focusing on machine learning models, AI tools, and enterprise AI integrations to help organizations automate workflows, extract insights from data, and build intelligent applications.
Industry: AI & Machine Learning
Website