Site Reliability Engineer, Machine Learning Systems - Singapore

ByteDance
Singapore
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningExperience: 1+ yearsEducation: bachelorsSkills: ["Logical analysis","Responsibility","Learning ability","Communication","Team spirit"]

Build and operate massively distributed ML training and inference systems that power large models. Own reliability across multi–data center, multi-region, and multi-cloud deployments, including disaster recovery, cluster governance, and business-service stability. Manage compute and storage resources to improve cost efficiency, and create monitoring/management tools for ML infrastructure. Join a global on-call roster supporting steady operations.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Site Reliability Engineer, Machine Learning Systems - Singapore

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 30 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Build and operate massively distributed ML training and inference systems that power large models. Own reliability across multi–data center, multi-region, and multi-cloud deployments, including disaster recovery, cluster governance, and business-service stability. Manage compute and storage resources to improve cost efficiency, and create monitoring/management tools for ML infrastructure. Join a global on-call roster supporting steady operations.
Location: Singapore
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Mid level

Key Responsibilities

  • •Ensure ML systems operate efficiently for large model deployment, training, evaluation, and inference.
  • •Maintain stability of offline tasks/services across multi-data center, multi-region, and multi-cloud scenarios.
  • •Manage compute/storage resource planning and budgeting, improving resource utilization and operation efficiency.
  • •Drive global system disaster recovery and cluster governance to keep business services stable.
  • •Build monitoring and management software tools, products, and systems for ML infrastructure and services, including on-call support.

Key Requirements

  • •Bachelor’s degree or above in Computer Science, computer engineering, or related fields.
  • •Strong proficiency in at least one programming language (Go/Python/Shell) in a Linux environment.
  • •Hands-on experience with Kubernetes and containers, plus 1+ year of relevant operations/maintenance experience.
  • •Experience operating and maintaining large-scale ML distributed systems.
  • •Experience operating and maintaining GPU servers and the ability to analyze, abstract, and split business logic effectively.
Experience:1+ yearsMachine learningLarge-scale systems
Education:Bachelor's in Computer Science, computer engineering or related fields
Skills:Logical analysisResponsibilityLearning abilityCommunicationTeam spirit
Tech Stack:GoPythonShellLinuxKubernetesContainersGPUNPURDMAStorageMulti-cloudLLMAIGCAGI

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn