Site Reliability Engineer, Hybrid Cloud Operation and Delivery - Data Infrastructure

ByteDance
Singapore
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsEducation: bachelorsSkills: ["Ownership","Autonomy","Open collaboration"]

Build and run reliable hybrid-cloud infrastructure for IaaS/PaaS/SaaS and AI workloads. Own delivery activities such as cloud platform planning, software deployment, and resource expansion, then operate cloud environments for internal and external customers with on-call support and change management. Partner with R&D to improve high-availability architecture, disaster recovery, and monitoring, and standardize serviceability acceptance for faster, more efficient rollouts.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Site Reliability Engineer, Hybrid Cloud Operation and Delivery - Data Infrastructure

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 30 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Build and run reliable hybrid-cloud infrastructure for IaaS/PaaS/SaaS and AI workloads. Own delivery activities such as cloud platform planning, software deployment, and resource expansion, then operate cloud environments for internal and external customers with on-call support and change management. Partner with R&D to improve high-availability architecture, disaster recovery, and monitoring, and standardize serviceability acceptance for faster, more efficient rollouts.
Location: Singapore
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Deliver hybrid-cloud products by handling cloud platform planning, software deployment, and resource expansion, and collaborate with R&D teams to complete delivery.
  • •Operate cloud platform environments for internal and external customers, including daily alarm handling, on-call support, change management, and ensuring stability during important event periods.
  • •Work with R&D to improve cloud product stability, enhancing high availability architecture, disaster recovery, and alarm monitoring based on large-scale operational experience.
  • •Improve hybrid-cloud serviceability by contributing to standardized SOW for O&M and delivery for new product versions and building SRE serviceability acceptance standards.

Key Requirements

  • •Bachelor's or Master's degree in Computer Science or a related major with at least 5 years of relevant experience.
  • •Solid knowledge of computer software and understanding of Linux, networking, and middleware principles.
  • •Familiar with one or more programming languages such as Shell, Python, Go, or Java, including building scripts/tools.
  • •Experience operating and maintaining systems such as virtual machines, containers, Kubernetes (K8s), load balancing, middleware, and AI models.
  • •Preferred experience operating data center (IDC) equipment such as switches and GPU servers, and working with cloud platform vendors.
Experience:5+ years
Education:Bachelor's in Computer Science
Skills:OwnershipAutonomyOpen collaboration
Tech Stack:LinuxShellPythonGoJavaVirtual machinesContainersKubernetes (K8s)Load balancingMiddlewareAI models

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn