Production System Engineer, Infrastructure Engineering

ByteDance
Singapore
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsEducation: bachelorsSkills: ["Troubleshooting","Root-cause analysis","Incident response","Cross-team collaboration","Communication"]

Build and run ByteDance hyperscale data center production systems by owning the server fleet lifecycle—from OS installation and deployment to operations, monitoring, and decommissioning. Create automation and monitoring tools to improve reliability, latency, scalability, and operability across core IDCs and CDN/edge. Triage incidents, perform root-cause analysis, and work on-call across regions while collaborating with infrastructure, platform, and operations stakeholders.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Production System Engineer, Infrastructure Engineering

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 30 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Build and run ByteDance hyperscale data center production systems by owning the server fleet lifecycle—from OS installation and deployment to operations, monitoring, and decommissioning. Create automation and monitoring tools to improve reliability, latency, scalability, and operability across core IDCs and CDN/edge. Triage incidents, perform root-cause analysis, and work on-call across regions while collaborating with infrastructure, platform, and operations stakeholders.
Location: Singapore
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Enhance the stability, efficiency, effectiveness, and scalability of data center and server operations, platform, and services.
  • •Improve the end-to-end server fleet lifecycle from system design/consultation through deployment, operations, and retirement.
  • •Develop and deploy automation tools to improve server reliability, scalability, and operability.
  • •Create monitoring solutions to improve availability, latency, and data center/server/network health.
  • •Troubleshoot production issues in high-pressure environments, perform root-cause analysis, and support incident response and postmortems, including on-call coverage.

Key Requirements

  • •Bachelor's degree in Computer Science, Electronic Engineering, or a related technical field (or equivalent practical experience).
  • •At least 5 years of experience in server operations, including Linux system administration and deep understanding of Linux kernels, drivers, and modules.
  • •Scripting proficiency in Bash and Python to automate system operations such as configuration, performance tuning, and security management.
  • •Experience planning, delivering, and operating large-scale data centers in different countries.
  • •Experience building and maintaining hardware, network, or service monitoring software for 10,000+ servers, plus customization and lifecycle management of operational/maintenance tools.
Experience:5+ yearsData centers
Education:Bachelor's
Skills:TroubleshootingRoot-cause analysisIncident responseCross-team collaborationCommunication
Languages:English
Tech Stack:LinuxBashPythonFlaskJavaScriptNode.jsSQLRedisAnsible

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn