Site Reliability Engineer, Compute Platform

TikTok
San Jose
Full timeFunction: DevOps, Cloud & InfrastructureEducation: bachelorsSkills: ["Problem-solving","Critical thinking","Written communication","Verbal communication","Ownership"]

Join a newly established Compute Platform SRE team that ensures reliability for TikTok’s major data warehouse products, services, and query engines. You’ll uphold SLAs, lead incident response and postmortems, and continuously improve performance by analyzing reliability signals. Work closely with product and development teams to embed reliability into the software lifecycle, automate provisioning and scaling, and plan capacity based on growth and upcoming initiatives.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
TikTok
TikTok
1 month ago

Site Reliability Engineer, Compute Platform

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 2 hours agoStatus: Live

Job Summary

Join a newly established Compute Platform SRE team that ensures reliability for TikTok’s major data warehouse products, services, and query engines. You’ll uphold SLAs, lead incident response and postmortems, and continuously improve performance by analyzing reliability signals. Work closely with product and development teams to embed reliability into the software lifecycle, automate provisioning and scaling, and plan capacity based on growth and upcoming initiatives.
Location: San Jose
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Ensure the reliability of major data warehouse products, services, and query engines, including ClickHouse, Spark, Presto, and Doris.
  • •Uphold SLAs by meeting service level objectives and responding promptly to outages and issues.
  • •Analyze performance and reliability patterns to identify bottlenecks and implement proactive measures to prevent disruptions.
  • •Lead incident management, troubleshooting, resolution, and postmortems, coordinating cross-functionally.
  • •Automate infrastructure provisioning, scaling, and management to reduce manual effort and improve service quality.

Key Requirements

  • •Bachelor’s degree or above in Computer Science, Engineering, or a related field.
  • •In-depth understanding of Linux, computer networking, and databases.
  • •Proficiency with SRE/DevOps tools, system monitoring tools, and container orchestration such as Kubernetes.
  • •Experience or familiarity with technologies including ClickHouse, Hadoop, Doris, Spark, and Presto, along with Kubernetes.
  • •Strong coding skills in at least one scripting or programming language (e.g., Python, Shell, Java, Go).
Experience:Big data
Education:Bachelor's in Computer Science, Engineering, or a related field
Skills:Problem-solvingCritical thinkingWritten communicationVerbal communicationOwnership
Tech Stack:LinuxComputer networkingDatabasesKubernetesClickHouseHadoopDorisSparkPrestoSREDevOpsSystem monitoringContainer orchestrationPythonShellJavaGo

Company Brief

TikTok
Short-form video platform that lets users create, share, and discover entertainment content through algorithmic recommendations. It also offers advertising and creator tools for brands, influencers, and businesses.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Headquarters: Singapore, Singapore
Founded: 2016
WebsiteLinkedIn