Site Reliability Engineer, TikTok Generalized Arch USTO
San Jose
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureEducation: bachelorsSkills: ["Problem-solving","Communication","Ownership","Responsibility","Incident handling"]Ensure the stability and reliability of TikTok’s core services by responding to production incidents and improving incident handling efficiency. Define and maintain system quality SLAs, identify system risks, and drive reliability, scalability, and performance. Contribute to disaster recovery initiatives, capacity planning, and contingency planning. Build and document operational best practices, tools, and frameworks, with opportunities to apply AI-powered automation to SRE and operations workflows.

