Cloud Site Reliability Engineer - DCS Cloud

ByteDance
Seattle
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 2+ yearsEducation: bachelorsSkills: ["Communication","Collaboration","Ownership"]

Build and run ByteDance’s global infrastructure across public and private clouds, scaling large-scale systems that power fast growth. Develop automation, monitoring, and visualization tools to optimize operations and reliability. Own the creation and standardization of cloud AMIs/images for compliant use across environments. Participate in technical operations and on-call incident response for cloud, OS, networking, performance, and reliability, driving improvements end to end.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Cloud Site Reliability Engineer - DCS Cloud

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 30 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Build and run ByteDance’s global infrastructure across public and private clouds, scaling large-scale systems that power fast growth. Develop automation, monitoring, and visualization tools to optimize operations and reliability. Own the creation and standardization of cloud AMIs/images for compliant use across environments. Participate in technical operations and on-call incident response for cloud, OS, networking, performance, and reliability, driving improvements end to end.
Location: Seattle
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Design, build, scale, and operate ByteDance’s global infrastructure spanning public and private clouds.
  • •Develop tools, automation frameworks, visualizations, and monitoring systems to streamline operations and optimize infrastructure.
  • •Create, manage, and standardize cloud AMIs/images across multiple environments aligned with global compliance standards.
  • •Respond to incidents through technical operations and on-call rotations covering cloud, OS, network, performance, and reliability.
  • •Drive infrastructure lifecycle improvements from ideation and design through development, deployment, support, and continuous refinement.

Key Requirements

  • •Bachelor’s degree or above in Computer Science, Software Engineering, Information Security, or a related field.
  • •2+ years of experience in Linux operations, SRE, or DevOps.
  • •Proficiency in Go, Python, or C++ with strong platform development, system tooling, and automation engineering.
  • •Strong computer science fundamentals including Linux OS, computer networks, storage systems, GPU systems, databases, troubleshooting, and root-cause analysis.
  • •Familiarity with reliability practices such as monitoring/alerting, capacity management, change management, canary/gray releases, incident response, and postmortems.
Experience:2+ years
Education:Bachelor's
Skills:CommunicationCollaborationOwnership
Tech Stack:LinuxGoPythonC++AWSAzureGCPOCIDockerKubernetesContainerdCgroupsNamespacesKVMQEMUGPUCUDAMIGMonitoringAlerting

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn