Site Reliability Engineer Intern (Global SRE- GMPT) - 2027 Start

TikTok
Singapore
Workplace: OnsiteInternshipFunction: DevOps, Cloud & InfrastructureSkills: ["Analytical thinking","Problem-solving","Ownership","Communication","Collaboration"]

Build and improve large-scale, globally distributed, fault-tolerant advertising systems as part of the Ads Infrastructure SRE team. You’ll take ownership of reliability, help design reliability architecture, and develop global disaster recovery solutions with incident response readiness. The role also involves capacity planning and performance optimization, plus creating reliability platforms and engineering tools for monitoring, alerting, SLO management, risk governance, and resource management.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
TikTok
TikTok
14 hours ago

Site Reliability Engineer Intern (Global SRE- GMPT) - 2027 Start

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 14 hours agoStatus: Live

Job Summary

Build and improve large-scale, globally distributed, fault-tolerant advertising systems as part of the Ads Infrastructure SRE team. You’ll take ownership of reliability, help design reliability architecture, and develop global disaster recovery solutions with incident response readiness. The role also involves capacity planning and performance optimization, plus creating reliability platforms and engineering tools for monitoring, alerting, SLO management, risk governance, and resource management.
Location: Singapore
Workplace: Onsite
Employment Type: Internship
Job Function: DevOps, Cloud & Infrastructure
Seniority: Intern level

Key Responsibilities

  • •Own the reliability of TikTok’s global advertising systems and help ensure high availability.
  • •Design reliability architecture for large-scale advertising systems.
  • •Design, implement, validate, and continuously optimize global disaster recovery (DR) solutions and incident response.
  • •Manage and plan infrastructure capacity, scaling resources efficiently and improving utilization through performance optimization.
  • •Develop reliability platforms and engineering tools for monitoring, alerting, SLO management, risk governance, and resource management.

Key Requirements

  • •Currently pursuing an Undergraduate or Master's in Computer Science or a related discipline.
  • •Expertise in Unix/Linux operating systems and IP networking.
  • •Programming experience in at least one of Python, Go, C, C++, or Java.
  • •Strong analytical and problem-solving skills, with a sense of ownership.
  • •Excellent communication and collaboration skills.
Experience:SREAds infrastructureLarge-scale systemsCloud platformsRecommendationSearch
Education:
Skills:Analytical thinkingProblem-solvingOwnershipCommunicationCollaboration
Tech Stack:Unix/LinuxIP networkingPythonGoCC++JavaDisaster recoveryIncident responseMonitoringAlertingSLO managementRisk governanceResource management

Company Brief

TikTok
Short-form video platform that lets users create, share, and discover entertainment content through algorithmic recommendations. It also offers advertising and creator tools for brands, influencers, and businesses.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Headquarters: Singapore, Singapore
Founded: 2016
WebsiteLinkedIn