Senior Site Reliability Engineer, Global E-Commerce

TikTok
San Jose
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsEducation: bachelorsSkills: ["Communication","Cross-functional collaboration","Ownership","Accountability"]

Build and improve the reliability of TikTok’s Global E-commerce services. You’ll strengthen disaster recovery readiness, manage server and compute capacity through planning and dynamic scaling, and enhance service monitoring for fast detection and resolution of failures. Working with business and engineering stakeholders, you’ll help drive stability governance and operate large-scale, cloud-native production systems across the U.S.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
TikTok
TikTok
1 month ago

Senior Site Reliability Engineer, Global E-Commerce

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 23 hours agoStatus: Live

Job Summary

Build and improve the reliability of TikTok’s Global E-commerce services. You’ll strengthen disaster recovery readiness, manage server and compute capacity through planning and dynamic scaling, and enhance service monitoring for fast detection and resolution of failures. Working with business and engineering stakeholders, you’ll help drive stability governance and operate large-scale, cloud-native production systems across the U.S.
Location: San Jose
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Ensure data center disaster recovery capabilities under normal operations, including contingency planning, drills, and capacity assurance.
  • •Manage and plan server and compute resources, including resource restructuring, capacity planning, and dynamic scaling.
  • •Establish and enhance service monitoring systems for timely alerting and rapid issue identification and resolution.
  • •Partner with business stakeholders to conduct ongoing stability governance.
  • •Operate and improve reliability for large-scale production systems supporting TikTok Global E-commerce services in the U.S.

Key Requirements

  • •Bachelor’s degree in Computer Science, Engineering, or related field, or equivalent practical experience.
  • •5+ years of experience in Site Reliability Engineering, infrastructure, or production engineering roles.
  • •Proficiency in at least one programming language such as Go, Python, or Java.
  • •Strong understanding of Linux systems, networking fundamentals, and distributed systems architecture.
  • •Experience operating services in cloud-native or large-scale production environments.
Experience:5+ yearsE-commerceInternet platformsCloud-nativeLarge-scale productionLarge-scale distributed systems
Education:Bachelor's
Skills:CommunicationCross-functional collaborationOwnershipAccountability
Tech Stack:GoPythonJavaLinuxNetworkingDistributed systems architectureCloud-native

Company Brief

TikTok
Short-form video platform that lets users create, share, and discover entertainment content through algorithmic recommendations. It also offers advertising and creator tools for brands, influencers, and businesses.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Headquarters: Singapore, Singapore
Founded: 2016
WebsiteLinkedIn