Research Engineer - ML Infrastructure
San Francisco, California
Workplace: OnsiteFull timeFunction: Research & Scientific (R&D)Experience: 4+ yearsSkills: ["Ownership","High velocity","Collaboration"]Build and optimize the distributed ML training stack that powers the models researchers ship. You’ll profile large GPU-cluster training runs to locate bottlenecks across compute, communication, and storage, then develop tooling for monitoring throughput, utilization, and uptime. Partner with research scientists to scale new architectures and recipes, and ensure reliability through fault tolerance, checkpointing, and deterministic orchestration. Optimize workloads via parallelism, quantization, and custom CUDA/Triton kernels.
Loading
Loading job details...
Preparing the role view and application actions.

