Site Reliability Engineer, Machine Learning Systems - Singapore
ByteDance
Singapore
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningExperience: 1+ yearsEducation: bachelorsSkills: ["Logical analysis","Responsibility","Learning ability","Communication","Team spirit"]Build and operate massively distributed ML training and inference systems that power large models. Own reliability across multi–data center, multi-region, and multi-cloud deployments, including disaster recovery, cluster governance, and business-service stability. Manage compute and storage resources to improve cost efficiency, and create monitoring/management tools for ML infrastructure. Join a global on-call roster supporting steady operations.

