Site Reliability Engineer - Data Infrastructure

TikTok
Seattle
Full timeFunction: DevOps, Cloud & InfrastructureExperience: 2+ yearsEducation: bachelorsSkills: ["Problem-solving","Communication","Proactive attitude","Continuous learning"]

Own the reliability and performance of TikTok’s core data services that power product teams. You’ll respond to production alerts and incidents, run change-controlled deployments and maintenance, and strengthen observability through better monitoring and instrumentation. The role includes rotational on-call coverage and cross-time-zone collaboration, plus building automation using scripting (Python/Go/Bash) and AI agents. You’ll also support data center and AI infrastructure operations.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
TikTok
TikTok
1 month ago

Site Reliability Engineer - Data Infrastructure

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 8 hours agoStatus: Live

Job Summary

Own the reliability and performance of TikTok’s core data services that power product teams. You’ll respond to production alerts and incidents, run change-controlled deployments and maintenance, and strengthen observability through better monitoring and instrumentation. The role includes rotational on-call coverage and cross-time-zone collaboration, plus building automation using scripting (Python/Go/Bash) and AI agents. You’ll also support data center and AI infrastructure operations.
Location: Seattle
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Serve as first responder for production alerts and incidents, executing runbooks to mitigate issues and escalate when needed.
  • •Perform operational excellence and change management tasks including deployments, configuration changes, and system maintenance.
  • •Automate repetitive manual work using scripting (Python/Go/Bash) and AI agents to reduce toil and improve consistency.
  • •Improve observability by refining monitoring dashboards, tuning alert thresholds, and ensuring systems are instrumented.
  • •Support daily operations, construction, and maintenance of data center environments and AI infrastructure.

Key Requirements

  • •Bachelor’s degree in Computer Science (or related technical field) or equivalent practical experience.
  • •2+ years of experience in an SRE, DevOps, Systems Administration, or similar role.
  • •Experience with at least one scripting language such as Python, Bash, or Go.
  • •Solid understanding of Linux operating systems and networking concepts.
  • •Familiarity with container technologies like Docker and Kubernetes.
Experience:2+ yearsSRE
Education:Bachelor's
Skills:Problem-solvingCommunicationProactive attitudeContinuous learning
Tech Stack:KubernetesRedisMySQLMessage QueuePythonGoBashLinuxDockerPostgreSQLPrometheusGrafanaELK StackAI Agents

Company Brief

TikTok
Short-form video platform that lets users create, share, and discover entertainment content through algorithmic recommendations. It also offers advertising and creator tools for brands, influencers, and businesses.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Headquarters: Singapore, Singapore
Founded: 2016
WebsiteLinkedIn