Site Reliability Engineer - Big Data Computer Platform

ByteDance
Singapore
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureEducation: bachelorsSkills: ["Problem-solving","Critical thinking","Written communication","Verbal communication","Ownership"]

Join the Compute Platform SRE team to ensure reliability for ByteDance’s major big data warehouse products, services, and query engines. You’ll drive incident management and postmortems, analyze performance patterns to prevent bottlenecks, and collaborate with development teams to optimize application performance. Build automation and toolchains for deployment and reliability assurance, including auto-failure detection and disaster drills, while forecasting infrastructure needs and staying current on industry best practices.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Site Reliability Engineer - Big Data Computer Platform

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 30 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Join the Compute Platform SRE team to ensure reliability for ByteDance’s major big data warehouse products, services, and query engines. You’ll drive incident management and postmortems, analyze performance patterns to prevent bottlenecks, and collaborate with development teams to optimize application performance. Build automation and toolchains for deployment and reliability assurance, including auto-failure detection and disaster drills, while forecasting infrastructure needs and staying current on industry best practices.
Location: Singapore
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Ensure reliability of major data warehouse products, services, and query engines (e.g., ClickHouse, Spark, Presto, Doris).
  • •Meet service level objectives and respond promptly to outages or system issues.
  • •Analyze service performance and reliability patterns to identify bottlenecks and implement proactive prevention measures.
  • •Lead incident troubleshooting and resolution, run postmortems, and coordinate cross-functionally to mitigate service-impacting events.
  • •Build and automate end-to-end deployment and reliability assurance toolchains, including scaling/provisioning automation, auto-healing capabilities, chaotic engineering, and disaster drills.

Key Requirements

  • •Bachelor’s degree or above in Computer Science, Engineering, or a related field.
  • •At least 2 years of experience in the SRE domain.
  • •At least 2 years of experience with Linux, computer networking, and databases.
  • •Proficiency with SRE/DevOps open-source toolsets, system monitoring tools, and container orchestration (e.g., Kubernetes).
  • •Familiarity with technologies such as ClickHouse, Hadoop, Doris, Spark, and Presto, plus 2+ years coding in Python, Shell, Java, Go, or similar languages.
Education:Bachelor's in Computer Science, Engineering, or a related field
Skills:Problem-solvingCritical thinkingWritten communicationVerbal communicationOwnership
Tech Stack:LinuxComputer networkingDatabasesKubernetesSystem monitoringClickHouseHadoopDorisSparkPrestoPythonShellJavaGo

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn