Research Engineer, ML Platform
Palo Alto
Workplace: HybridFull timeFunction: Research & Scientific (R&D)Experience: 4+ yearsSkills: ["Developer experience focus","Problem diagnosis","Reliability mindset","Ownership","Comfort in ambiguous environments"]Build and operate the ML platform powering large-scale training, evaluation, and batch inference. You’ll develop APIs and tooling for distributed GPU workloads, design workload orchestration (scheduling, quotas, priorities, preemption, placement), and improve heterogeneous GPU capacity utilization. Partnering across the ML lifecycle, you’ll add observability and reliability mechanisms, create self-service workflows for researchers, and participate in on-call to troubleshoot production systems.
Loading
Loading job details...
Preparing the role view and application actions.

