Cloud Site Relibility Engineer - DCS

ByteDance
Singapore
Workplace: OnsiteFull timeFunction: Data Analytics & Business IntelligenceEducation: bachelorsSkills: ["Communication","Collaboration","Problem-solving","Ownership","Cross-team execution"]

Build and operate hyperscale datacenter and infrastructure systems across public and private clouds. Develop automation, monitoring, and operational tooling, and create standardized cloud AMIs/images aligned with global compliance. Own reliability operations including on-call incident response for cloud, OS, network, performance, and drive continuous improvements across the infrastructure lifecycle.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Cloud Site Relibility Engineer - DCS

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 30 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Build and operate hyperscale datacenter and infrastructure systems across public and private clouds. Develop automation, monitoring, and operational tooling, and create standardized cloud AMIs/images aligned with global compliance. Own reliability operations including on-call incident response for cloud, OS, network, performance, and drive continuous improvements across the infrastructure lifecycle.
Location: Singapore
Workplace: Onsite
Employment Type: Full time
Job Function: Data Analytics & Business Intelligence
Seniority: Mid level

Key Responsibilities

  • •Design, build, scale, and operate global infrastructure spanning public and private clouds.
  • •Develop tools, automation frameworks, visualizations, and monitoring systems to optimize infrastructure operations.
  • •Create, manage, and standardize cloud AMIs/images for multi-environment use while meeting global compliance standards.
  • •Participate in technical operations and on-call rotations to address incidents across cloud, OS, network, performance, and reliability.
  • •Improve the end-to-end infrastructure lifecycle from ideation and design through development, deployment, user support, and refinement.

Key Requirements

  • •Bachelor’s degree or above in Computer Science, Software Engineering, Information Security, or related fields.
  • •2+ years of experience in Linux operations, SRE, or DevOps.
  • •Proficiency in at least one programming language (Go, Python, or C++) with platform/system tooling and automation skills.
  • •Strong fundamentals in Linux OS, computer networks, storage systems, GPU systems, and databases, including troubleshooting and root-cause analysis.
  • •Experience with reliability practices such as monitoring/alerting, capacity/change management, canary/gray releases, incident response, and postmortems.
Experience:SREDevOpsCloud infrastructureHyperscale datacenters
Education:Bachelor's in Computer Science, Software Engineering, Information Security, or a related field
Skills:CommunicationCollaborationProblem-solvingOwnershipCross-team execution
Tech Stack:GoPythonC++LinuxAWSAzureGCPOCIDockerKubernetesKVMQEMUGPUCUDAMIGCgroupsNamespacesAMIsMonitoringAlerting

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn