Senior Site Reliability Engineer - Traffic Infrastructure

ByteDance
San Jose
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 3+ yearsEducation: bachelorsSkills: ["Analytical skills","Communication","Responsibility","Team spirit"]

Build and maintain ByteDance’s Network-Traffic Infrastructure stability and production operations for edge services outside China. You’ll deliver facilities, components, and releases, design stability assurance with observability, troubleshooting, and automated/self-healing issue resolution, and implement an emergency response system for ticketing, risk governance, and long-term optimization. Work across infrastructure, Kubernetes, edge computing, and cloud networking to improve scale, performance, and cost.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Senior Site Reliability Engineer - Traffic Infrastructure

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 30 days agoStatus: Live

Job Summary

Build and maintain ByteDance’s Network-Traffic Infrastructure stability and production operations for edge services outside China. You’ll deliver facilities, components, and releases, design stability assurance with observability, troubleshooting, and automated/self-healing issue resolution, and implement an emergency response system for ticketing, risk governance, and long-term optimization. Work across infrastructure, Kubernetes, edge computing, and cloud networking to improve scale, performance, and cost.
Location: San Jose
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Operate and maintain ByteDance’s Network-Traffic Infrastructure while ensuring stability.
  • •Deliver and operate/maintain the production system, including facilities, components, and releases, improving efficiency of delivery and operations.
  • •Design and implement a stability assurance system covering observability, troubleshooting (root cause analysis and impact assessment), and issue resolution (manual and self-healing).
  • •Design and implement an emergency response system including ticket processing, emergency response, risk governance, and long-term optimization.

Key Requirements

  • •Bachelor’s degree or above in computer science or a related field, with at least 3 years of relevant experience in R&D, system operation/maintenance, or SRE.
  • •Solid understanding of infrastructure architecture and experience with Kubernetes, edge computing, cloud networking, load balancing, and microservice architecture.
  • •Strong analytical skills, excellent communication abilities, and a strong sense of responsibility and team spirit.
  • •Preferred: Practical operations and maintenance and stability assurance experience with Kubernetes, cloud computing, edge computing, and cloud networking.
  • •Preferred: Experience with high availability, reliability assurance, and emergency response systems for infrastructure or distributed systems.
Experience:3+ yearsR&DSystem operationsSREKubernetesEdge computingCloud networking
Education:Bachelor's in Computer science
Skills:Analytical skillsCommunicationResponsibilityTeam spirit
Tech Stack:KubernetesEdge computingCloud networkingLoad BalanceMicroservice architectureCloud computing

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn