Research Engineer Graduate (AI Training Systems Reliability & Performance - Seed Infra) - 2026 Start (PhD)
Seattle
Workplace: OnsiteFull timeFunction: Education & TrainingEducation: phdSkills: ["Problem-solving","Collaboration","Analytical skills"]Build reliability and performance improvements for large-scale AI training systems spanning pre-training, fine-tuning, evaluation, and inference. Develop observability, profiling, and debugging tools for distributed ML workloads, and identify bottlenecks across GPU, networking, and storage. Work on multi-GPU and multi-node distributed training frameworks, collaborate with model and infrastructure teams, and support incident analysis to improve operational stability.
Loading
Loading job details...
Preparing the role view and application actions.

