Senior Site Reliability Engineer - Data Infrastructure

TikTok
Seattle
Full timeFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsEducation: bachelorsSkills: ["Incident response","Post-incident reviews","Communication","Problem-solving","Ownership","Mentorship"]

Own reliability for core data services powering TikTok products, engineering resilience, scalability, and efficiency across a distributed infrastructure stack. Respond to and resolve production incidents, lead blameless postmortems, and define SLOs and error budgets for critical data systems. Drive capacity and cost optimization, automate operational toil (including AI orchestration), and uphold production readiness via runbooks and change management. Also lead data center and AI infrastructure work and mentor junior SREs.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
TikTok
TikTok
1 month ago

Senior Site Reliability Engineer - Data Infrastructure

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 8 hours agoStatus: Live

Job Summary

Own reliability for core data services powering TikTok products, engineering resilience, scalability, and efficiency across a distributed infrastructure stack. Respond to and resolve production incidents, lead blameless postmortems, and define SLOs and error budgets for critical data systems. Drive capacity and cost optimization, automate operational toil (including AI orchestration), and uphold production readiness via runbooks and change management. Also lead data center and AI infrastructure work and mentor junior SREs.
Location: Seattle
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Act as incident commander for critical production issues; lead triage, resolution, and blameless postmortems with follow-up actions.
  • •Define, negotiate, and maintain SLOs for critical data services and champion error budgets to balance reliability and feature work.
  • •Drive capacity planning, performance tuning, and resource management to improve scalability and control costs.
  • •Design and build automation to eliminate operational toil; leverage AI Agents to improve deployment safety and operational efficiency.
  • •Lead operational standards for production readiness (runbooks, monitoring, alerting) and manage complex change vetting; also lead data center and AI infrastructure construction/optimization.
  • •Mentor junior SREs and serve as a reliability subject matter expert across teams.

Key Requirements

  • •Bachelor’s degree in Computer Science (or related technical field) or equivalent practical experience.
  • •5+ years of experience in Site Reliability Engineering, Production Engineering, or similar roles.
  • •Strong proficiency in a programming/scripting language for automation (Go, Python, or Bash).
  • •Deep understanding of Linux/Unix, networking fundamentals (TCP/IP, DNS), and distributed systems.
  • •Experience managing large-scale data infrastructure (e.g., MySQL, Redis, Kafka, Flink) and leading incident response for high-impact events.
Experience:5+ yearsSaaSData infrastructure
Education:Bachelor's in Computer Science
Skills:Incident responsePost-incident reviewsCommunicationProblem-solvingOwnershipMentorship
Tech Stack:KubernetesRedisMySQLMessage QueueLinuxUnixTCP/IPDNSGoPythonBashKafkaFlinkSLO/SLAError budgetsRunbooksMonitoringAlertingAI AgentsAI orchestration

Company Brief

TikTok
Short-form video platform that lets users create, share, and discover entertainment content through algorithmic recommendations. It also offers advertising and creator tools for brands, influencers, and businesses.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Headquarters: Singapore, Singapore
Founded: 2016
WebsiteLinkedIn