Senior Site Reliability Engineer - Data Infrastructure (San Jose)

ByteDance
San Jose
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsEducation: bachelorsSkills: ["Incident response","Blameless postmortems","Problem-solving","Communication","Ownership"]

Own reliability for core data services powering the company’s products, operating a massive distributed environment and responding to production incidents. Define and maintain SLOs and error budgets, drive capacity and cost optimization, and raise operational quality through runbooks, monitoring, and change management. Build pragmatic automation (including AI orchestration) to reduce toil and improve deployment safety, while collaborating across time zones and mentoring other SREs.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Senior Site Reliability Engineer - Data Infrastructure (San Jose)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 29 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Own reliability for core data services powering the company’s products, operating a massive distributed environment and responding to production incidents. Define and maintain SLOs and error budgets, drive capacity and cost optimization, and raise operational quality through runbooks, monitoring, and change management. Build pragmatic automation (including AI orchestration) to reduce toil and improve deployment safety, while collaborating across time zones and mentoring other SREs.
Location: San Jose
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Act as incident commander for critical production issues and lead blameless post-incident reviews with follow-up actions.
  • •Define, negotiate, and maintain SLOs/SLA and use error budgets for reliability tradeoffs.
  • •Drive capacity planning, performance tuning, and resource management to scale efficiently and stay within budget.
  • •Design and build automation and AI orchestration to eliminate operational toil and improve deployment safety.
  • •Uphold production operations standards (runbooks, monitoring, alerting) and lead data center and AI infrastructure construction, maintenance, and optimization.

Key Requirements

  • •Bachelor’s degree in Computer Science or related technical field, or equivalent practical experience.
  • •5+ years of experience in Site Reliability Engineering, Production Engineering, or a similar role.
  • •Strong proficiency in a programming or scripting language (Go, Python, or Bash) for automation and tools.
  • •Deep understanding of Linux/Unix, networking fundamentals (TCP/IP, DNS), and distributed systems.
  • •Extensive hands-on experience managing large-scale data infrastructure (e.g., MySQL, Redis, Kafka, Flink) and production Kubernetes operations.
Experience:5+ yearsSREProduction engineeringDistributed systemsData infrastructureContainer orchestrationKubernetes
Education:Bachelor's in Computer Science
Skills:Incident responseBlameless postmortemsProblem-solvingCommunicationOwnership
Tech Stack:KubernetesRedisMySQLMessage QueueLinuxUnixTCP/IPDNSGoPythonBashKafkaFlink

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn