Site Reliability Engineer Graduate (Global SRE- GMPT) - 2027 Start

TikTok
Singapore
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureEducation: bachelorsSkills: ["Analytical thinking","Problem-solving","Ownership","Communication","Collaboration"]

Build and improve globally distributed, fault-tolerant ad infrastructure reliability. Own reliability for ByteDance’s global advertising systems, help design reliability architecture, and ensure high availability. Develop, implement, and continuously optimize disaster recovery solutions and incident response. Plan and scale infrastructure capacity while improving utilization through performance optimization. Create reliability platforms and engineering tools, including monitoring, alerting, SLO management, and resource governance.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
TikTok
TikTok
14 hours ago

Site Reliability Engineer Graduate (Global SRE- GMPT) - 2027 Start

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 14 hours agoStatus: Live

Job Summary

Build and improve globally distributed, fault-tolerant ad infrastructure reliability. Own reliability for ByteDance’s global advertising systems, help design reliability architecture, and ensure high availability. Develop, implement, and continuously optimize disaster recovery solutions and incident response. Plan and scale infrastructure capacity while improving utilization through performance optimization. Create reliability platforms and engineering tools, including monitoring, alerting, SLO management, and resource governance.
Location: Singapore
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Graduate level

Key Responsibilities

  • •Own the reliability of ByteDance’s global advertising systems and ensure high availability at large scale.
  • •Participate in the design of reliability architecture for advertising systems.
  • •Design, implement, validate, and continuously optimize global disaster recovery solutions and lead DR preparedness and incident response.
  • •Manage and plan infrastructure capacity to scale efficiently with business growth and improve utilization via performance optimization.
  • •Develop reliability platforms and engineering tools, including monitoring, alerting, SLO management, risk governance, and resource management.

Key Requirements

  • •Completing or recently completed a Bachelor's or Master’s degree in Computer Science or a related discipline.
  • •Expertise in Unix/Linux operating systems and IP networking.
  • •Experience programming in at least one: Python, Go, C, C++, or Java.
  • •Strong analytical and problem-solving skills with a sense of ownership and excellent communication and collaboration.
  • •Preferred: experience in SRE for ads/recommendation systems and designing, analyzing, and troubleshooting large-scale distributed systems.
Experience:AdsRecommendation systemsDistributed systems
Education:Bachelor's in Computer Science
Skills:Analytical thinkingProblem-solvingOwnershipCommunicationCollaboration
Tech Stack:UnixLinuxIP networkingPythonGoCC++JavaMonitoringAlertingSLO managementDisaster recoveryIncident response

Company Brief

TikTok
Short-form video platform that lets users create, share, and discover entertainment content through algorithmic recommendations. It also offers advertising and creator tools for brands, influencers, and businesses.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Headquarters: Singapore, Singapore
Founded: 2016
WebsiteLinkedIn