Site Reliability Engineer - Cloud Infrastructure

TikTok
Dublin
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureEducation: bachelorsSkills: ["Reliability","Troubleshooting","Capacity planning","Performance testing","Fault diagnosis"]

Build and run reliable, scalable infrastructure for production systems that support TikTok’s global user base. You’ll ensure high availability and performance, troubleshoot issues, and drive capacity planning, performance testing, anomaly analysis, and fault diagnosis. The role also focuses on automating and improving operation platforms, researching system architectures and technologies, optimizing IT costs, and designing data protection plans for new IDC setups.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
TikTok
TikTok
1 month ago

Site Reliability Engineer - Cloud Infrastructure

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 1 hour agoStatus: Live

Job Summary

Build and run reliable, scalable infrastructure for production systems that support TikTok’s global user base. You’ll ensure high availability and performance, troubleshoot issues, and drive capacity planning, performance testing, anomaly analysis, and fault diagnosis. The role also focuses on automating and improving operation platforms, researching system architectures and technologies, optimizing IT costs, and designing data protection plans for new IDC setups.
Location: Dublin
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Ensure stability of core infrastructure with a focus on high availability, reliability, performance, and capacity.
  • •Troubleshoot technical issues and collaborate with engineering teams on capacity planning, performance testing, anomaly analysis, and fault diagnosis/resolution.
  • •Research and evaluate large-scale system architectures and technologies to improve systems and processes.
  • •Design and implement O&M platforms for efficient, automated, and intelligent maintenance.
  • •Develop delivery standards and capacity assessments to optimize IT costs and establish data protection plans for new IDC setups.

Key Requirements

  • •Bachelor’s or Master’s degree in Computer Science or a related major.
  • •Solid knowledge of computer software and Linux, including storage and network IO fundamentals.
  • •Familiarity with one or more programming languages such as Python, Go, or Java, with understanding of design patterns and coding principles.
  • •Experience with storage and systems including KV, Table, Graph, Redis, MySQL, MongoDB, MQ, and Kafka.
  • •Experience with computing & big data and systems such as Kubernetes, Docker/containers, AIops, Spark, Flink, FaaS, RPC framework, and service mesh.
Education:Bachelor's in Computer Science or related major
Skills:ReliabilityTroubleshootingCapacity planningPerformance testingFault diagnosis
Tech Stack:LinuxPythonGoJavaElasticsearchRedisMySQLMongoDBKafkaKubernetesDockerContainersAIopsSparkFlinkFunction as a serviceRPC frameworkService MeshKVTable

Company Brief

TikTok
Short-form video platform that lets users create, share, and discover entertainment content through algorithmic recommendations. It also offers advertising and creator tools for brands, influencers, and businesses.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Headquarters: Singapore, Singapore
Founded: 2016
WebsiteLinkedIn