Site Reliability Engineer - Infrastructure

TikTok
Sydney
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 3+ yearsEducation: mastersSkills: ["Problem-solving","Ownership","Self-governance","Independence","Openness"]

Own reliability and efficiency for core infrastructure spanning storage, computing, and databases. Tackle large-scale performance issues through troubleshooting, bottleneck analysis, and high-availability architecture upgrades. Build automation, monitoring, and SOA governance platforms, and optimize CPU and systems costs through delivery standards and budgeting. Design and implement data protection and support new IDC setup to meet compliance requirements.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
TikTok
TikTok
1 month ago

Site Reliability Engineer - Infrastructure

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

Own reliability and efficiency for core infrastructure spanning storage, computing, and databases. Tackle large-scale performance issues through troubleshooting, bottleneck analysis, and high-availability architecture upgrades. Build automation, monitoring, and SOA governance platforms, and optimize CPU and systems costs through delivery standards and budgeting. Design and implement data protection and support new IDC setup to meet compliance requirements.
Location: Sydney
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Ensure reliability and efficiency of core infrastructure, including capacity and stability, and establish reliability standards and recovery SOP.
  • •Troubleshoot and locate technical issues, perform bottleneck analysis, and manage high-availability architecture transformation and upgrades.
  • •Build automated operation solutions for large-scale systems and partner with system development teams for iteration.
  • •Design and implement software platforms and monitoring frameworks for efficient, automated, and intelligent SOA governance.
  • •Design and set up new IDC and implement a data protection plan to meet standard requirements.

Key Requirements

  • •Solid basic knowledge of computer software.
  • •Understanding of Linux operating system, storage, and network I/O principles.
  • •Familiarity with one or more programming languages such as Python, Go, and Java.
  • •Knowledge of design patterns and coding principles.
  • •Experience with storage systems (e.g., KV, Table, Graph, Redis, MySQL, MongoDB, MQ, Kafka) and large-scale computing/big data technologies (e.g., Kubernetes, Docker/containers).
Experience:3+ yearsInfrastructureStorage systemsBig dataKubernetesCloud infrastructure
Education:Master's in Computer Science or related major
Skills:Problem-solvingOwnershipSelf-governanceIndependenceOpenness
Tech Stack:LinuxPythonGoJavaRedisMySQLMongoDBKafkaKubernetesDockerContainersAIopsSparkFlinkFunction as a serviceRPC FrameworkService Mesh

Company Brief

TikTok
Short-form video platform that lets users create, share, and discover entertainment content through algorithmic recommendations. It also offers advertising and creator tools for brands, influencers, and businesses.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Headquarters: Singapore, Singapore
Founded: 2016
WebsiteLinkedIn