Tech Lead Cloud Site Reliability Engineer - DCS Cloud

ByteDance
San Jose
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsEducation: bachelorsSkills: ["Communication","Collaboration","Ownership"]

Lead the design, build, scaling, and operation of ByteDance’s global infrastructure across public and private clouds. Build automation frameworks, monitoring, and visualization to optimize hyper-scale operations. Create and standardize cloud AMIs/images aligned with global compliance, and participate in on-call rotations to resolve incidents across cloud, OS, networking, performance, and reliability. Drive end-to-end improvements from ideation through continuous refinement.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Tech Lead Cloud Site Reliability Engineer - DCS Cloud

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 30 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Lead the design, build, scaling, and operation of ByteDance’s global infrastructure across public and private clouds. Build automation frameworks, monitoring, and visualization to optimize hyper-scale operations. Create and standardize cloud AMIs/images aligned with global compliance, and participate in on-call rotations to resolve incidents across cloud, OS, networking, performance, and reliability. Drive end-to-end improvements from ideation through continuous refinement.
Location: San Jose
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Sr. Manager level

Key Responsibilities

  • •Design, build, scale, and operate global infrastructure spanning large-scale systems across public and private clouds.
  • •Develop tools, automation frameworks, visualizations, and monitoring systems to streamline operations and optimize infrastructure.
  • •Create, manage, and standardize cloud AMIs/images across multiple environments while aligning with global compliance standards.
  • •Participate in technical operations and on-call rotations to address incidents related to cloud, OS, network, performance, and reliability.
  • •Drive improvements across the full infrastructure lifecycle from ideation and design through development, deployment, user support, and continuous refinement.

Key Requirements

  • •Bachelor’s degree or above in Computer Science, Software Engineering, Information Security, or a related field.
  • •5+ years of experience in Linux operations, SRE, or DevOps.
  • •Proficiency in at least one programming language such as Go, Python, or C++, with strong engineering capabilities in platform development and automation.
  • •Deep understanding of Linux OS principles, computer networks, storage systems, GPU systems, and databases, with troubleshooting and root-cause analysis skills.
  • •Experience with core reliability practices including monitoring/alerting, capacity and change management, canary/gray releases, incident response, and postmortems.
Experience:5+ yearsSREDevOpsCloud infrastructure
Education:Bachelor's
Skills:CommunicationCollaborationOwnership
Tech Stack:GoPythonC++LinuxSREDevOpsPublic cloudPrivate cloudOCIAWSAzureGCPAMIsKVMQEMUDockerKubernetesContainerdCgroupsNamespaces

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn