Site Reliability Engineer (Multiple Positions)

TikTok
San Jose
Workplace: OnsiteFull timeUSD 226,138 - 316,800 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 1-3 yearsEducation: mastersSkills: ["Problem-solving","Troubleshooting","Incident response","Documentation","Monitoring"]

Provide site reliability engineering to keep TikTok’s large-scale, fault-tolerant systems highly available. You’ll support reliability, scalability, and release-cycle improvements across infrastructure services—covering deployment, monitoring, incident response, and blameless postmortems. Measure availability and latency, build automation and monitoring tools, and help establish best practices. Share on-call responsibility and troubleshoot issues across a wide set of services.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
TikTok
TikTok
11 hours ago

Site Reliability Engineer (Multiple Positions)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 11 hours agoStatus: Live

Job Summary

Provide site reliability engineering to keep TikTok’s large-scale, fault-tolerant systems highly available. You’ll support reliability, scalability, and release-cycle improvements across infrastructure services—covering deployment, monitoring, incident response, and blameless postmortems. Measure availability and latency, build automation and monitoring tools, and help establish best practices. Share on-call responsibility and troubleshoot issues across a wide set of services.
Location: San Jose
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Provide site reliability engineering support to ensure the highest availability of large-scale, fault-tolerant systems.
  • •Improve reliability, scalability, and the release cycle of infrastructure services from inception through design, development, deployment, support, and refinement.
  • •Measure and monitor availability, latency, and overall service health.
  • •Practice sustainable user support, incident response, and blameless postmortems.
  • •Build tools, automations, visualizations, and monitors; establish design and best practices; share on-call responsibility and troubleshoot across services and functional areas.
Travel: Medium travel

Pay and Benefits

Salary: USD 226,138 - 316,800 annually

Key Requirements

  • •Must have a Master’s degree (or foreign equivalent) in Computer Science, Engineering, Information Systems, Mathematics, or related field and 1 year of related work experience; OR a Bachelor’s degree (or foreign equivalent) and 3 years of related work experience.
  • •Provide functionality and reliability support for critical site components by measuring and monitoring availability, latency, and overall system health, including performance tuning and troubleshooting.
  • •Monitor system activity and resolve system issues.
  • •Coordinate and monitor data services operations, including SLA management and system deployment.
  • •Analyze error logs to identify issues, work with service owners to resolve them, and create and maintain runbooks for alerts, troubleshooting, and resolution.
Experience:1-3 yearsSite reliability engineeringInfrastructure operationsCloud infrastructureData services
Education:Master's in Computer Science, Engineering (any), Information Systems, Mathematics, or a related field
Skills:Problem-solvingTroubleshootingIncident responseDocumentationMonitoring

Company Brief

TikTok
Short-form video platform that lets users create, share, and discover entertainment content through algorithmic recommendations. It also offers advertising and creator tools for brands, influencers, and businesses.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Headquarters: Singapore, Singapore
Founded: 2016
WebsiteLinkedIn