Site Reliability Engineer - System Service Global
ByteDance
San Jose
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureEducation: high_schoolSkills: ["Troubleshooting","Incident response","Root cause analysis","Continuous improvement","Collaboration"]Own the reliability and operations of ByteDance’s non-China data center foundational services, including DNS, NTP, DHCP, NAT, APT repositories, and Kerberos authentication. Manage large-scale Linux host infrastructure with OS lifecycle, configuration standardization, and fleet health monitoring. Design high-availability and disaster recovery deployment architectures, set SLOs, and lead incident response with blameless post-mortems. Improve automation to reduce toil and boost operational efficiency.

