Site Reliability Engineer - Machine Learning Systems (Singapore)
ByteDance
Singapore
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningExperience: 1+ yearsEducation: bachelorsSkills: ["Logical analysis","Communication","Self-driven","Team spirit","Responsibility","Learning ability","Documentation"]Build and operate highly reliable, scalable ML training and inference systems for large models. Own production stability for offline tasks across multi–data center, multi-region, and multi-cloud environments, including disaster recovery, cluster governance, and on-call support. Manage compute/storage resources, costs, and capacity planning, and develop monitoring and operational tools for ML infrastructure. Collaborate with a global ML systems team spanning the US, China, and Singapore.

