Site Reliability Engineer (Multiple Positions)

ByteDance
San Jose
Workplace: OnsiteFull timeUSD 212,800 - 387,600 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 2-5 yearsEducation: bachelorsSkills: []

Design and develop highly available infrastructure platforms for large-scale distributed production systems. Build automation for provisioning, deployment orchestration, configuration management, and operational workflows. Use observability to monitor performance and reliability, identify bottlenecks, and implement reliability frameworks. Lead incident response with root-cause analysis, implement disaster recovery and failover across multi-region environments, and mentor junior SREs.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
12 hours ago

Site Reliability Engineer (Multiple Positions)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 12 hours agoStatus: Live

Job Summary

Design and develop highly available infrastructure platforms for large-scale distributed production systems. Build automation for provisioning, deployment orchestration, configuration management, and operational workflows. Use observability to monitor performance and reliability, identify bottlenecks, and implement reliability frameworks. Lead incident response with root-cause analysis, implement disaster recovery and failover across multi-region environments, and mentor junior SREs.
Location: San Jose
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Design and develop highly available infrastructure platforms supporting large-scale distributed production systems.
  • •Build automation systems for infrastructure provisioning, deployment orchestration, configuration management, and operational workflows.
  • •Monitor system performance, availability, and reliability using observability tools and identify performance bottlenecks and reliability issues.
  • •Develop reliability engineering frameworks and implement disaster recovery architectures, failover strategies, and business continuity solutions across multi-region environments.
  • •Oversee incident response, conduct root cause analysis, implement long-term remediation, participate in technical rotations, and mentor junior Site Reliability Engineers.

Pay and Benefits

Salary: USD 212,800 - 387,600 annually

Key Requirements

  • •Master’s degree in Computer Science/Engineering/IT/Mathematics (or related) with 2 years of related work OR Bachelor’s with 5 years of post-bachelor’s progressive related work.
  • •Experience across all phases of the software development lifecycle, including requirements, design, development, testing, deployment, and maintenance of back-end and cloud-native projects.
  • •Design scalable and highly available software systems for cloud environments.
  • •Process and analyze large quantities of logs and data using analytical tools to build dashboards/alerts using search processing languages.
  • •Perform Linux administration, including monitoring and debugging, and develop Docker containers deployed and managed in Kubernetes.
Experience:2-5 years
Education:Bachelor's in Computer Science, Engineering (any), Information Technology, Mathematics, or a related field
Tech Stack:LinuxDockerKubernetes

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn