Cloud Site Reliability Engineer - DCS Cloud

ByteDance
San Jose
Full timeFunction: DevOps, Cloud & InfrastructureExperience: 2+ yearsEducation: bachelorsSkills: ["Communication","Collaboration","Ownership","Troubleshooting","Root-cause analysis"]

Build, scale, and operate ByteDance’s global infrastructure across public and private clouds. Develop automation, monitoring, and visualization tooling to streamline operations and improve reliability. Create and standardize cloud AMIs/images aligned with compliance requirements, and participate in on-call rotations to respond to incidents spanning cloud, OS, networking, performance, and reliability. Drive end-to-end infrastructure lifecycle improvements from design through continuous refinement.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Cloud Site Reliability Engineer - DCS Cloud

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 29 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Build, scale, and operate ByteDance’s global infrastructure across public and private clouds. Develop automation, monitoring, and visualization tooling to streamline operations and improve reliability. Create and standardize cloud AMIs/images aligned with compliance requirements, and participate in on-call rotations to respond to incidents spanning cloud, OS, networking, performance, and reliability. Drive end-to-end infrastructure lifecycle improvements from design through continuous refinement.
Location: San Jose
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Design, build, scale, and operate ByteDance’s global infrastructure across public and private clouds.
  • •Develop tools, automation frameworks, visualizations, and monitoring systems to optimize global infrastructure operations.
  • •Create, manage, and standardize cloud AMIs/images across multiple environments aligned with global compliance standards.
  • •Participate in technical operations and on-call rotations to address incidents related to cloud, OS, network, performance, and reliability.
  • •Drive improvements across the full infrastructure lifecycle from ideation and design through development, deployment, support, and continuous refinement.

Key Requirements

  • •Bachelor’s degree or above in Computer Science, Software Engineering, Information Security, or a related field.
  • •2+ years of experience in Linux operations, SRE, or DevOps.
  • •Proficiency in Go, Python, or C++ with strong platform development, system tooling, and automation capabilities.
  • •Strong fundamentals in Linux, computer networks, storage systems, GPU systems, and databases, with troubleshooting and root-cause analysis skills.
  • •Familiarity with reliability practices such as monitoring/alerting, capacity management, change management, canary/gray releases, incident response, and postmortems.
Experience:2+ yearsLinux operationsSREDevOpsPublic cloudGPU clustersCloud-nativeOpen source
Education:Bachelor's in Computer Science, Software Engineering, Information Security, or a related field
Skills:CommunicationCollaborationOwnershipTroubleshootingRoot-cause analysis
Tech Stack:GoPythonC++LinuxAWSAzureGCPOCICloudAMIsAutomationMonitoringCanary/gray releasesIncident responsePostmortemGPU systemsCUDAMIGKVMQEMU

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn