Site Reliability Engineer - ARK Large Model Platform (Singapore)

ByteDance
Singapore
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsEducation: bachelorsSkills: []

Build and operate the ARK Large Model Platform on VolcanoEngine, focusing on stability across control and data planes for large-scale model systems. You will implement DevOps practices, develop observability for monitoring reliability and performance, and manage super large-scale clusters to ensure efficient operations and maintenance. The role also involves researching solutions for large model deployments and supporting a cloud-native infrastructure stack used for high-performance inference.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Site Reliability Engineer - ARK Large Model Platform (Singapore)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 33 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Build and operate the ARK Large Model Platform on VolcanoEngine, focusing on stability across control and data planes for large-scale model systems. You will implement DevOps practices, develop observability for monitoring reliability and performance, and manage super large-scale clusters to ensure efficient operations and maintenance. The role also involves researching solutions for large model deployments and supporting a cloud-native infrastructure stack used for high-performance inference.
Location: Singapore
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Develop and support the ARK Large Model Platform on VolcanoEngine, including solutions for large model implementations and applications.
  • •Manage and oversee stability across control and data aspects of large-scale model systems using effective DevOps practices.
  • •Develop and enhance observability systems to monitor stability, ensuring high reliability and performance.
  • •Handle super large-scale cluster management and ensure efficient operations and maintenance for large model systems.

Key Requirements

  • •B. Sc or higher in Computer Science or a related field from an accredited institution.
  • •Minimum 5 years of R&D experience in cloud computing or large-scale model systems.
  • •Proficiency in cloud-native technologies and understanding of the relevant technology stack.
  • •Expertise in Golang, Python, or Java, used professionally.
  • •Familiarity with cloud-native technologies for log collection, monitoring, and alerting.
Experience:5+ yearsCloud computingLarge-scale model systemsR&DMachine learning
Education:Bachelor's in Computer Science
Tech Stack:VolcanoEngineDevOpsObservabilityLog collectionMonitoringAlertingGolangPythonJavaTerraformInfrastructure as codeCloud-native

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn