Tech Lead Site Reliability Engineer, TikTok Generalized Arch USTO

TikTok
San Jose
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsEducation: bachelorsSkills: ["Problem-solving","Communication","Ownership","Responsibility"]

Lead Site Reliability Engineering for TikTok’s Generalized Architecture US Tech and Operations team, ensuring core services run stable and cost-effectively at global scale. Own reliability through incident response, SLAs, and risk management, and help strengthen resilience via disaster recovery design and capacity planning. Build operational tools, best practices, and technical documentation, leveraging programming and AI-powered automation to improve SRE efficiency.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
TikTok
TikTok
1 month ago

Tech Lead Site Reliability Engineer, TikTok Generalized Arch USTO

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 18 hours agoStatus: Live

Job Summary

Lead Site Reliability Engineering for TikTok’s Generalized Architecture US Tech and Operations team, ensuring core services run stable and cost-effectively at global scale. Own reliability through incident response, SLAs, and risk management, and help strengthen resilience via disaster recovery design and capacity planning. Build operational tools, best practices, and technical documentation, leveraging programming and AI-powered automation to improve SRE efficiency.
Location: San Jose
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Sr. Manager level

Key Responsibilities

  • •Ensure stability and reliability of TikTok’s core services through fast production incident response and mechanisms to improve incident handling efficiency.
  • •Define and maintain system quality SLAs using continuous data operations; identify and manage system risks to improve reliability, scalability, and performance.
  • •Support disaster recovery initiatives including risk assessment, disaster recovery design, capacity planning, and contingency plan development to strengthen resilience.
  • •Develop and maintain best practices, tools, and frameworks for operations and maintenance.
  • •Provide guidance on system architecture design and component selection and produce technical and operational documentation.

Key Requirements

  • •Bachelor’s degree (or above) in Computer Science or a related field.
  • •Solid computer science and software engineering foundation, including operating systems (especially Linux), storage systems, and network I/O principles.
  • •Proficiency in one or more programming languages such as Python, Go, Java, PHP, C, or C++.
  • •Strong problem-solving skills with a systematic approach, effective communication, and strong ownership.
  • •5+ years of relevant experience in a large-scale internet or cloud-based environment.
Experience:5+ yearsLarge-scale internetCloud-based environment
Education:Bachelor's in Computer Science or related field
Skills:Problem-solvingCommunicationOwnershipResponsibility
Tech Stack:PythonGoJavaPHPCC++Linux

Company Brief

TikTok
Short-form video platform that lets users create, share, and discover entertainment content through algorithmic recommendations. It also offers advertising and creator tools for brands, influencers, and businesses.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Headquarters: Singapore, Singapore
Founded: 2016
WebsiteLinkedIn