Research Engineer - ML Infrastructure

Chai Discovery
San Francisco, California
Workplace: OnsiteFull timeFunction: Research & Scientific (R&D)Experience: 4+ yearsSkills: ["Ownership","High velocity","Collaboration"]

Build and optimize the distributed ML training stack that powers the models researchers ship. You’ll profile large GPU-cluster training runs to locate bottlenecks across compute, communication, and storage, then develop tooling for monitoring throughput, utilization, and uptime. Partner with research scientists to scale new architectures and recipes, and ensure reliability through fault tolerance, checkpointing, and deterministic orchestration. Optimize workloads via parallelism, quantization, and custom CUDA/Triton kernels.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Chai Discovery
Chai Discovery
1 day ago

Research Engineer - ML Infrastructure

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 15 hours agoStatus: Live

Job Summary

Build and optimize the distributed ML training stack that powers the models researchers ship. You’ll profile large GPU-cluster training runs to locate bottlenecks across compute, communication, and storage, then develop tooling for monitoring throughput, utilization, and uptime. Partner with research scientists to scale new architectures and recipes, and ensure reliability through fault tolerance, checkpointing, and deterministic orchestration. Optimize workloads via parallelism, quantization, and custom CUDA/Triton kernels.
Location: San Francisco, California
Workplace: Onsite
Employment Type: Full time
Job Function: Research & Scientific (R&D)
Seniority: Mid level

Key Responsibilities

  • •Architect, debug, and optimize the distributed ML training stack across model, layer, and kernel levels to remove runtime and reliability bottlenecks.
  • •Profile end-to-end training runs to identify bottlenecks across compute, communication, and storage, and build monitoring tooling for throughput, utilization, and uptime.
  • •Optimize ML workloads using parallelism strategies, quantization, and custom CUDA/Triton kernels.
  • •Collaborate with research scientists to ensure new model architectures and training recipes scale efficiently.
  • •Own training-stack reliability, including fault tolerance, checkpointing, and deterministic orchestration for long-running jobs.

Key Requirements

  • •4+ years of industry experience working within AI/ML infrastructure teams.
  • •Proficiency in Python and PyTorch or JAX.
  • •Strong software systems design skills across the stack from model code to kernels.
  • •Experience orchestrating GPU clusters and large-scale model training.
  • •Experience optimizing ML workloads using parallelism, quantization, and CUDA/Triton kernels.
Experience:4+ yearsAI/MLMachine learningInfrastructureDistributed systems
Skills:OwnershipHigh velocityCollaboration
Tech Stack:PythonPyTorchJAXCUDATritonGPU clustersDistributed ML trainingCheckpointingFault toleranceQuantizationParallelism

Company Brief

Chai Discovery
Builds foundation AI models to predict and design interactions between biochemical molecules, enabling de novo antibody and drug design to accelerate pharmaceutical discovery and development.
Industry: Biotech
Company Size: Small (11 to 50 employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series B
Headquarters: San Francisco, United States
Founded: 2024
WebsiteLinkedIn