Site Reliability Engineer - Compute Platform

TikTok
Seattle
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureEducation: bachelorsSkills: ["Problem-solving","Critical thinking"]

Join the Compute Platform SRE team to ensure reliability for TikTok’s major data warehouse products, services, and query engines. You’ll manage SLAs and respond to outages, optimize performance, lead incident troubleshooting and postmortems, and automate infrastructure provisioning and scaling. Partner with product and development teams to embed reliability into the software lifecycle and forecast capacity and demand for upcoming initiatives.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
TikTok
TikTok
1 month ago

Site Reliability Engineer - Compute Platform

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 2 hours agoStatus: Live

Job Summary

Join the Compute Platform SRE team to ensure reliability for TikTok’s major data warehouse products, services, and query engines. You’ll manage SLAs and respond to outages, optimize performance, lead incident troubleshooting and postmortems, and automate infrastructure provisioning and scaling. Partner with product and development teams to embed reliability into the software lifecycle and forecast capacity and demand for upcoming initiatives.
Location: Seattle
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Ensure reliability of TikTok’s major data warehouse products, services, and query engines (e.g., ClickHouse, Spark, Presto, Doris).
  • •Uphold service level agreements (SLAs) by meeting service level objectives and promptly responding to outages.
  • •Analyze performance and reliability patterns to prevent disruptions and improve application performance with development teams.
  • •Lead incident management: troubleshoot and resolve incidents, coordinate cross-functional responses, and conduct postmortems.
  • •Automate infrastructure provisioning, scaling, and management to reduce manual work and improve service quality.

Key Requirements

  • •Bachelor’s degree or above in Computer Science, Engineering, or a related field.
  • •In-depth understanding of Linux, computer networking, and databases.
  • •Proficiency with SRE/DevOps toolsets, system monitoring tools, and container orchestration (e.g., Kubernetes).
  • •Experience with open-source or commercial data technologies such as ClickHouse, Hadoop, Doris, Spark, and Presto.
  • •Strong coding skills in at least one language such as Python, Shell, Java, or Go.
Experience:Big DataData warehousingSRE/DevOpsData platforms
Education:Bachelor's
Skills:Problem-solvingCritical thinking
Tech Stack:LinuxComputer networkingDatabasesKubernetesClickHouseHadoopDorisSparkPrestoPythonShellJavaGo

Company Brief

TikTok
Short-form video platform that lets users create, share, and discover entertainment content through algorithmic recommendations. It also offers advertising and creator tools for brands, influencers, and businesses.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Headquarters: Singapore, Singapore
Founded: 2016
WebsiteLinkedIn