Site Reliability Engineer - Machine Learning Systems (Singapore)

ByteDance
Singapore
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningExperience: 1+ yearsEducation: bachelorsSkills: ["Logical analysis","Communication","Self-driven","Team spirit","Responsibility","Learning ability","Documentation"]

Build and operate highly reliable, scalable ML training and inference systems for large models. Own production stability for offline tasks across multi–data center, multi-region, and multi-cloud environments, including disaster recovery, cluster governance, and on-call support. Manage compute/storage resources, costs, and capacity planning, and develop monitoring and operational tools for ML infrastructure. Collaborate with a global ML systems team spanning the US, China, and Singapore.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Site Reliability Engineer - Machine Learning Systems (Singapore)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 30 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Build and operate highly reliable, scalable ML training and inference systems for large models. Own production stability for offline tasks across multi–data center, multi-region, and multi-cloud environments, including disaster recovery, cluster governance, and on-call support. Manage compute/storage resources, costs, and capacity planning, and develop monitoring and operational tools for ML infrastructure. Collaborate with a global ML systems team spanning the US, China, and Singapore.
Location: Singapore
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning

Key Responsibilities

  • •Ensure ML systems run efficiently for large model deployment, training, evaluation, and inference.
  • •Maintain stability of offline tasks/services across multi–data center, multi-region, and multi-cloud scenarios.
  • •Plan and manage resource usage, including computing/storage capacity and cost/budget.
  • •Own global system disaster recovery, cluster machine governance, and stability and utilization improvements.
  • •Build monitoring and management tools for ML infrastructure and provide on-call support.

Key Requirements

  • •Bachelor's degree or above in Computer Science, computer engineering, or a related field.
  • •Strong proficiency in at least one programming language (Go, Python, or Shell) in a Linux environment.
  • •1+ years of relevant operation and maintenance experience with Kubernetes and containers.
  • •Hands-on experience operating and maintaining large-scale ML distributed systems.
  • •Experience operating and maintaining GPU servers; ability to analyze and abstract business logic and communicate effectively.
Experience:1+ yearsMachine learningDistributed systemsAI large modelsLLMAIGCGPU infrastructure
Education:Bachelor's in Computer Science, computer engineering or related fields
Skills:Logical analysisCommunicationSelf-drivenTeam spiritResponsibilityLearning abilityDocumentation
Tech Stack:GoPythonShellLinuxKubernetesContainersGPUNPURDMAStorageMulti-cloudDisaster recovery

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn