ML Infra Engineer
San Francisco
Workplace: OnsiteFull timeFunction: Software EngineeringSkills: ["Communication","Ownership","Problem-solving"]You will design, build, and maintain large-scale ML training infrastructure, owning systems for training and inference, scaling JAX/PyTorch training across TPU/GPU clusters, and optimizing performance. You’ll collaborate with researchers and engineers to turn ideas into experiments and production runs, balancing researcher flexibility with system reliability in a fast-paced ML infrastructure team.
Loading
Loading job details...
Preparing the role view and application actions.

