Tech Lead Cloud Site Reliability Engineer - DCS Cloud

ByteDance
Seattle
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsSkills: ["Communication","Collaboration","Ownership","Problem-solving"]

Lead the design, build, scaling, and operation of ByteDance’s global infrastructure across public and private clouds. You’ll create automation, monitoring, and optimization tools, and manage standardized cloud AMIs/images for multi-environment compliance. Participate in on-call rotations to handle incidents across cloud, OS, network, performance, and reliability, and continuously improve the infrastructure lifecycle from ideation through deployment and refinement.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Tech Lead Cloud Site Reliability Engineer - DCS Cloud

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 30 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Lead the design, build, scaling, and operation of ByteDance’s global infrastructure across public and private clouds. You’ll create automation, monitoring, and optimization tools, and manage standardized cloud AMIs/images for multi-environment compliance. Participate in on-call rotations to handle incidents across cloud, OS, network, performance, and reliability, and continuously improve the infrastructure lifecycle from ideation through deployment and refinement.
Location: Seattle
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Design, build, scale, and operate ByteDance’s global infrastructure spanning public and private clouds.
  • •Develop tools, automation frameworks, visualizations, and monitoring systems to streamline operations and optimize infrastructure.
  • •Create, manage, and standardize cloud AMIs/images across multiple environments to meet global compliance standards.
  • •Participate in technical operations and on-call rotations to address incidents across cloud, OS, network, performance, and reliability.
  • •Drive improvements across the infrastructure lifecycle from ideation and design through development, deployment, user support, and continuous refinement.

Key Requirements

  • •Bachelor’s degree or above in Computer Science, Software Engineering, Information Security, or a related field.
  • •5+ years of experience in Linux operations, SRE, or DevOps.
  • •Proficiency in at least one programming language such as Go, Python, or C++ with strong engineering for platform development and automation.
  • •Deep understanding of Linux OS principles, computer networks, storage systems, GPU systems, and databases, with troubleshooting and root-cause analysis skills.
  • •Experience with reliability practices including monitoring/alerting, capacity management, change management, canary/gray releases, incident response, and postmortems.
Experience:5+ yearsSREDevOpsLinuxPublic cloud
Skills:CommunicationCollaborationOwnershipProblem-solving
Tech Stack:LinuxGoPythonC++AWSAzureGCPOCIAMIsCloud imagesKVM/QEMUKubernetesDockerContainerdCgroupsNamespacesGPU clustersCUDAMIG

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn