Site Reliability Engineer, System - System Service Global

ByteDance
Singapore
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureEducation: bachelorsSkills: ["Self-motivated","Collaboration","Troubleshooting","Continuous improvement","Incident response"]

Own reliability for ByteDance’s non-China data center infrastructure services, spanning core foundational components like DNS, NTP, DHCP, NAT, APT repositories, and Kerberos. Build high-availability deployment architectures with fault tolerance and disaster recovery, define SLOs/SLIs, and lead incident response with blameless post-mortems. Partner with network, security, and application teams while automating host management to reduce toil and improve operational efficiency.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Site Reliability Engineer, System - System Service Global

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 30 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Own reliability for ByteDance’s non-China data center infrastructure services, spanning core foundational components like DNS, NTP, DHCP, NAT, APT repositories, and Kerberos. Build high-availability deployment architectures with fault tolerance and disaster recovery, define SLOs/SLIs, and lead incident response with blameless post-mortems. Partner with network, security, and application teams while automating host management to reduce toil and improve operational efficiency.
Location: Singapore
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Manage and maintain large-scale host infrastructure across ByteDance’s non-China data centers, including OS lifecycle management, configuration standardization, and fleet health monitoring.
  • •Ensure reliability and availability of core foundational services such as DNS, NTP, DHCP, NAT, APT repository management, and Kerberos authentication.
  • •Design and implement deployment architectures for foundational services with high availability, fault tolerance, and disaster recovery across regions.
  • •Define and enforce SLOs/SLIs, lead incident response, perform root cause analysis, and run post-mortem reviews to improve reliability.
  • •Collaborate with network, security, and application teams and identify automation opportunities to reduce toil and increase operational efficiency.

Key Requirements

  • •Bachelor’s degree or higher in Electrical Engineering, Computer Engineering, Computer Science, or related majors.
  • •Solid experience managing large-scale Linux hosts, including OS deployment, configuration management, patching, and fleet operations.
  • •Hands-on knowledge of foundational data center services: DNS (BIND/PowerDNS), NTP, DHCP, NAT, APT repository management, and Kerberos.
  • •Proficiency with DevOps tooling such as Ansible/Salt/Puppet and CI/CD pipelines.
  • •Familiarity with SRE practices (SLO/SLI, error budgets, blameless post-mortems) and high-availability/disaster recovery design patterns.
Experience:Production environmentsData centers
Education:Bachelor's
Skills:Self-motivatedCollaborationTroubleshootingContinuous improvementIncident response
Tech Stack:LinuxDNSBINDPowerDNSNTPDHCPNATAPTKerberosAnsibleSaltPuppetCI/CDPythonGoBash

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn