Site Reliability Engineer, ARK Large Model Platform (Singapore)

ByteDance
Singapore
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureEducation: bachelorsSkills: []

Build and evolve the ARK Large Model Platform on VolcanoEngine, applying DevOps practices to ensure stability across control and data planes for large-scale model systems. Drive reliability and performance by developing observability capabilities for monitoring and alerting, and manage super-large clusters to support efficient operation and maintenance. Collaborate on systematic solutions to deploy large-model implementations across industries while reducing infrastructure costs and meeting growing user demand.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Site Reliability Engineer, ARK Large Model Platform (Singapore)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 29 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Build and evolve the ARK Large Model Platform on VolcanoEngine, applying DevOps practices to ensure stability across control and data planes for large-scale model systems. Drive reliability and performance by developing observability capabilities for monitoring and alerting, and manage super-large clusters to support efficient operation and maintenance. Collaborate on systematic solutions to deploy large-model implementations across industries while reducing infrastructure costs and meeting growing user demand.
Location: Singapore
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Own ARK Large Model Platform development on VolcanoEngine and research systematic solutions for large-model implementations and applications.
  • •Manage and oversee stability of control and data aspects of large-scale model systems using effective DevOps practices.
  • •Develop and enhance observability systems to monitor stability, ensuring high reliability and performance.
  • •Handle super large-scale cluster management to ensure efficient operation and maintenance.
  • •Work to reduce IT cost of large-model applications while meeting growing user demand for intelligent interaction.

Key Requirements

  • •B. Sc. or higher in Computer Science or a related field, with R&D experience in cloud computing or large-scale model systems.
  • •Proficiency in cloud-native technologies and understanding of the relevant technology stack.
  • •Expertise in one of: Golang, Python, or Java, used proficiently in a professional setting.
  • •Familiarity with cloud-native technologies for log collection, monitoring, and alerting.
  • •Experience constructing and maintaining stability systems and operating/maintaining large-scale systems; infrastructure as code (Terraform) is highly desirable.
Experience:Cloud computingLarge-scale model systems
Education:Bachelor's in Computer Science
Tech Stack:VolcanoEngineDevOpsObservabilityLog collectionMonitoringAlertingCloud-native technologiesGolangPythonJavaInfrastructure as codeTerraform

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn