Machine Learning Performance Engineer - Offboard Training & Inference
Sunnyvale
Workplace: OnsiteFull timeUSD 215,000 - 285,000 annuallyFunction: Data Science & Machine LearningSkills: ["Profiling","Roofline analysis","Debugging","Analytical problem-solving","Collaboration"]Own performance optimization for large-scale distributed training and high-throughput batch inference over petabyte-scale autonomy logs. Profile end to end, close the gap between theoretical accelerator capability and real workload throughput, and improve cluster goodput by reducing GPU idle time from data pipeline, I/O, scheduling gaps, stragglers, and recovery issues. Build benchmarking/observability tooling and collaborate across accelerator and ML infrastructure teams to reduce training time-to-result and offline processing cost.
Loading
Loading job details...
Preparing the role view and application actions.

