Site Reliability Engineer - System Service Global

ByteDance
San Jose
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureEducation: high_schoolSkills: ["Troubleshooting","Incident response","Root cause analysis","Continuous improvement","Collaboration"]

Own the reliability and operations of ByteDance’s non-China data center foundational services, including DNS, NTP, DHCP, NAT, APT repositories, and Kerberos authentication. Manage large-scale Linux host infrastructure with OS lifecycle, configuration standardization, and fleet health monitoring. Design high-availability and disaster recovery deployment architectures, set SLOs, and lead incident response with blameless post-mortems. Improve automation to reduce toil and boost operational efficiency.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Site Reliability Engineer - System Service Global

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 30 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Own the reliability and operations of ByteDance’s non-China data center foundational services, including DNS, NTP, DHCP, NAT, APT repositories, and Kerberos authentication. Manage large-scale Linux host infrastructure with OS lifecycle, configuration standardization, and fleet health monitoring. Design high-availability and disaster recovery deployment architectures, set SLOs, and lead incident response with blameless post-mortems. Improve automation to reduce toil and boost operational efficiency.
Location: San Jose
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Manage and maintain large-scale host infrastructure across ByteDance’s non-China data centers, including OS lifecycle management, configuration standardization, and fleet-wide health monitoring.
  • •Ensure reliability and availability of core foundational services including DNS, NTP, DHCP, NAT, APT repository, and Kerberos authentication.
  • •Design and implement deployment architectures for foundational services with high availability, fault tolerance, and disaster recovery across regions.
  • •Develop and enforce SLOs; lead incident response, root cause analysis, and post-mortem reviews to drive reliability improvements.
  • •Identify automation opportunities across host management and service operations to reduce toil and improve operational efficiency.

Key Requirements

  • •Bachelor’s degree or higher in Electrical Engineering, Computer Engineering, Computer Science or related majors.
  • •Solid experience managing large-scale Linux host infrastructure, including OS deployment, configuration management, patching, and fleet operations.
  • •Hands-on knowledge of core data center foundational services: DNS (BIND/PowerDNS), NTP, DHCP, NAT, APT repository management, and Kerberos.
  • •Proficiency with DevOps tooling, including configuration management tools (Ansible, Salt, Puppet) and CI/CD pipelines.
  • •Familiarity with SRE principles such as SLO/SLI definition, error budget management, and blameless post-mortems.
Experience:Data centersInfrastructureDevOpsSRE
Education:High School in Electrical Engineering, Computer Engineering, Computer Science or related majors
Skills:TroubleshootingIncident responseRoot cause analysisContinuous improvementCollaboration
Tech Stack:LinuxDNSBINDPowerDNSNTPDHCPNATAPTKerberosAnsibleSaltPuppetCI/CDSLOSLIPythonGoBash

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn