ML Infra Engineer, Modeling
San Francisco
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Cross-functional communication","Ownership mindset","Debugging","Performance optimization"]Own and scale training and inference infrastructure for large-scale model development. Build and maintain reusable, efficient JAX training pipelines, orchestration, scheduling, checkpointing, and monitoring. Collaborate with researchers to scale distributed training across TPU/GPU clusters, improving performance through profiling and optimizations. Help evolve core training code so research experiments can reliably become production-grade training runs.
Loading
Loading job details...
Preparing the role view and application actions.

