Site Reliability Engineer, Compute Platform

ByteDance
San Jose
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureEducation: bachelorsSkills: ["Problem-solving","Critical thinking","Written communication","Verbal communication","Collaboration"]

Build and operate the reliability layer for major Big Data services and products, including data warehouse products and query engines. Own SLAs, respond to outages, and run incident management with troubleshooting and postmortems. Continuously optimize performance by analyzing reliability patterns, automating infrastructure provisioning and scaling, and collaborating with product and development teams. Forecast capacity and demand to support growth.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Site Reliability Engineer, Compute Platform

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 30 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Build and operate the reliability layer for major Big Data services and products, including data warehouse products and query engines. Own SLAs, respond to outages, and run incident management with troubleshooting and postmortems. Continuously optimize performance by analyzing reliability patterns, automating infrastructure provisioning and scaling, and collaborating with product and development teams. Forecast capacity and demand to support growth.
Location: San Jose
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Ensure reliability of major TikTok data warehouse products, services, and query engines (e.g., ClickHouse, Spark, Presto, Doris).
  • •Uphold SLAs by meeting service level objectives and responding promptly to outages and issues.
  • •Continuously optimize performance by analyzing reliability patterns, preventing disruptions, and working with development teams to improve application performance and resource utilization.
  • •Lead incident management, troubleshooting, and postmortems while coordinating with cross-functional teams to mitigate service-impacting events.
  • •Automate infrastructure provisioning, scaling, and management; plan capacity and demand; and collaborate with product and development teams on reliability considerations.

Key Requirements

  • •Bachelor’s degree or above in Computer Science, Engineering, or a related field.
  • •In-depth understanding of Linux, computer networking, and databases.
  • •Proficiency with SRE/DevOps open-source toolsets, system monitoring tools, and container orchestration (e.g., Kubernetes).
  • •Experience or familiarity with open-source or commercial technologies such as ClickHouse, Hadoop, Doris, Spark, and Presto.
  • •Strong coding skills in at least one language (e.g., Python, Shell, Java, Go).
Experience:Big DataData warehouses
Education:Bachelor's
Skills:Problem-solvingCritical thinkingWritten communicationVerbal communicationCollaboration
Tech Stack:LinuxNetworkingDatabasesSREDevOpsSystem monitoringKubernetesClickHouseHadoopDorisSparkPrestoPythonShellJavaGo

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn