Senior Site Reliability Engineer - Data Infrastructure (San Jose)
ByteDance
San Jose
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsEducation: bachelorsSkills: ["Incident response","Blameless postmortems","Problem-solving","Communication","Ownership"]Own reliability for core data services powering the company’s products, operating a massive distributed environment and responding to production incidents. Define and maintain SLOs and error budgets, drive capacity and cost optimization, and raise operational quality through runbooks, monitoring, and change management. Build pragmatic automation (including AI orchestration) to reduce toil and improve deployment safety, while collaborating across time zones and mentoring other SREs.

