Cloud Native Observability Management Platform Engineer - Data Infrastructure

ByteDance
Singapore
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsEducation: bachelorsSkills: ["Clean code","Unit testing","Quality engineering","Stability engineering","High-concurrency coding"]

Build and optimize cloud-native observability management for a time-series observability platform, covering cluster lifecycle from creation and configuration to HA, disaster recovery, backup, and audit. Develop automated testing frameworks to improve function, performance, and stability. Partner with engineering efforts to drive reliability for ByteDance and Volcengine observability services using k8s, Docker, and microservice technologies.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Cloud Native Observability Management Platform Engineer - Data Infrastructure

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 31 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Build and optimize cloud-native observability management for a time-series observability platform, covering cluster lifecycle from creation and configuration to HA, disaster recovery, backup, and audit. Develop automated testing frameworks to improve function, performance, and stability. Partner with engineering efforts to drive reliability for ByteDance and Volcengine observability services using k8s, Docker, and microservice technologies.
Location: Singapore
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Deliver and optimize user-oriented operations products and SRE-oriented O&M products for observability services.
  • •Develop a cloud-native intelligent management platform for observability time-series engine, including cluster lifecycle management (creation, configuration, HA, disaster recovery, backup, audit, etc.).
  • •Build and optimize automated testing frameworks for observability (function, performance, stability).
  • •Drive quality engineering and stability construction for observability services.
  • •Support high-quality, cost-effective, high-performance time-series services for ByteDance and Volcengine.

Key Requirements

  • •5+ years of software engineering experience in large-scale production environments.
  • •Bachelor's degree in Computer Science or a related field.
  • •Proficient in Go/Java and experience with microservice frameworks such as gin/kitex/Spring.
  • •In-depth understanding of observability and familiarity with tools like Prometheus, Grafana, ELK Stack, or Jaeger.
  • •Strong server-side coding practices, including unit testing and working with high concurrency; MySQL/PostgreSQL and O&M experience are a plus.
Experience:5+ yearsLarge-scale productionObservabilityTime-seriesMicroservices
Education:Bachelor's in Computer Science or related field
Skills:Clean codeUnit testingQuality engineeringStability engineeringHigh-concurrency coding
Tech Stack:GoJavaGinKitexSpringK8sKubernetesDockerMicroservicesPrometheusGrafanaELK StackJaegerMySQLPostgreSQLMySQL/PostgreSQLUnit testingYarnMesosRaft

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn