Senior Site Reliability Engineer - Data Infrastructure
Seattle
Full timeFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsEducation: bachelorsSkills: ["Incident response","Post-incident reviews","Communication","Problem-solving","Ownership","Mentorship"]Own reliability for core data services powering TikTok products, engineering resilience, scalability, and efficiency across a distributed infrastructure stack. Respond to and resolve production incidents, lead blameless postmortems, and define SLOs and error budgets for critical data systems. Drive capacity and cost optimization, automate operational toil (including AI orchestration), and uphold production readiness via runbooks and change management. Also lead data center and AI infrastructure work and mentor junior SREs.

