Production System Engineer, Infrastructure Engineering

ByteDance
Singapore
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 3+ yearsEducation: bachelorsSkills: ["Troubleshooting","Incident response","Root-cause analysis","Cross-team collaboration","Automation mindset"]

Build and operate hyperscale data center production systems by owning the server fleet lifecycle—from OS installation and deployment to operation and retirement. Improve stability, scalability, availability, and latency through automation, monitoring, and time-sensitive disaster recovery with root-cause analysis and postmortems. Collaborate across infrastructure, platform, and operations teams to design solutions for Core IDCs and CDN/Edge, including on-call incident support across regions.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Production System Engineer, Infrastructure Engineering

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 29 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Build and operate hyperscale data center production systems by owning the server fleet lifecycle—from OS installation and deployment to operation and retirement. Improve stability, scalability, availability, and latency through automation, monitoring, and time-sensitive disaster recovery with root-cause analysis and postmortems. Collaborate across infrastructure, platform, and operations teams to design solutions for Core IDCs and CDN/Edge, including on-call incident support across regions.
Location: Singapore
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Enhance stability, efficiency, effectiveness, and scalability of data center and server operations, platforms, and services.
  • •Contribute to the full server fleet lifecycle, from system introduction and launch reviews through deployment, operation, and retirement.
  • •Develop and deploy automation tools and solutions to improve server reliability, scalability, and operability.
  • •Create monitoring tools to improve availability, latency, and overall health of infrastructure, servers, and networks.
  • •Troubleshoot production issues during disaster recovery, perform root-cause analysis, run preventive measures, and support sustainable incident response and postmortems; participate in on-call across regions.

Key Requirements

  • •Bachelor’s degree in Computer Science, Electronic Engineering (or equivalent practical experience).
  • •Minimum 3 years in server operations, including Linux system administration and strong knowledge of Linux kernels, drivers, and modules.
  • •Ability to script with Bash and Python to automate system operations (configuration, performance tuning, security).
  • •3+ years supporting planning, delivery, and operation of large-scale data centers across different countries.
  • •Experience developing/maintaining monitoring software for 10,000+ servers and/or tailoring operation/maintenance tools for new server hardware.
Experience:3+ yearsDatacentersServer operations
Education:Bachelor's
Skills:TroubleshootingIncident responseRoot-cause analysisCross-team collaborationAutomation mindset
Tech Stack:LinuxLinux kernelDriversModulesBashPythonRESTful APIsFlaskJavaScriptNode.jsSQLRedisAnsibleGPU servers

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn