Site Reliability Engineer - Data Infrastructure

TikTok
San Jose
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 2+ yearsEducation: bachelorsSkills: ["Problem-solving","Communication","Proactive attitude","Desire to learn"]

Own the reliability, scalability, and efficiency of core data services powering TikTok products. This SRE role responds to production incidents via runbooks, performs change-controlled deployments and maintenance, and improves observability through dashboards, alert tuning, and instrumentation. You’ll automate repetitive operations with scripting and AI augmentation, and support daily operations and upkeep of data center and AI infrastructure for large-scale data processing.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
TikTok
TikTok
1 month ago

Site Reliability Engineer - Data Infrastructure

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 18 hours agoStatus: Live

Job Summary

Own the reliability, scalability, and efficiency of core data services powering TikTok products. This SRE role responds to production incidents via runbooks, performs change-controlled deployments and maintenance, and improves observability through dashboards, alert tuning, and instrumentation. You’ll automate repetitive operations with scripting and AI augmentation, and support daily operations and upkeep of data center and AI infrastructure for large-scale data processing.
Location: San Jose
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Entry level

Key Responsibilities

  • •Respond to production alerts and incidents as a first responder, executing runbooks and escalating when needed.
  • •Perform operational excellence work like deployments, configuration changes, and maintenance under change control.
  • •Automate repetitive manual tasks using scripting and AI Agents to reduce toil and improve consistency.
  • •Improve observability by refining monitoring dashboards, tuning alert thresholds, and ensuring systems are instrumented to detect and diagnose issues.
  • •Support daily operations, construction, and maintenance of data center and AI infrastructure for large-scale data processing.

Key Requirements

  • •Bachelor’s degree in Computer Science or a related technical field, or equivalent practical experience.
  • •2+ years of experience in an SRE, DevOps, Systems Administration, or similar role.
  • •Experience with at least one scripting language (e.g., Python, Bash, Go).
  • •Solid understanding of Linux operating systems and networking concepts.
  • •Familiarity and/or hands-on experience with container technologies and data/observability tooling is preferred.
Experience:2+ years
Education:Bachelor's
Skills:Problem-solvingCommunicationProactive attitudeDesire to learn
Tech Stack:KubernetesRedisMySQLMessage QueuePythonGoBashLinuxDockerPostgreSQLPrometheusGrafanaELK StackAI Agents

Company Brief

TikTok
Short-form video platform that lets users create, share, and discover entertainment content through algorithmic recommendations. It also offers advertising and creator tools for brands, influencers, and businesses.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Headquarters: Singapore, Singapore
Founded: 2016
WebsiteLinkedIn