Member of Technical Staff - GPU Infrastructure Engineer
San Francisco
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Root cause analysis","Communication","Prioritization"]Own the reliability and operation of GPU cluster compute environments powering foundation model training and research. Debug issues across compute, storage, networking, schedulers, and distributed workloads, while improving CPU/GPU/storage utilization through tooling and automation. Build monitoring and validation, onboard and migrate workloads across GPU providers and hardware platforms, and contribute to the longer-term training infrastructure architecture.

