Machine Learning Ops Engineer, Global SRE
San Jose
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningEducation: bachelorsSkills: ["Communication","Ownership","Drive","Troubleshooting"]Own the reliability of machine learning systems by defining SLOs for online model serving and improving the stability and success of offline training workflows. Drive GPU-based model training rollouts across Non-China regions, support stability for AIGC workloads, and manage ML resource planning including cost, budget, and efficiency for both offline and online pipelines. Collaborate with teams to troubleshoot production issues and ensure efficient ML operations.

