Production System Engineer

ByteDance
London
Workplace: OnsiteFull timeFunction: Manufacturing & Production OperationsExperience: 5+ yearsEducation: bachelorsSkills: ["Cross-team collaboration","Root-cause analysis","Incident response"]

Own production server lifecycles for large-scale data center and infrastructure services, from deployment and operation through decommissioning and recycling. Enhance server lifecycle processes via system design consultation, launch reviews, and continuous operations improvements. Build automation and monitoring tools to improve reliability, scalability, availability, and latency. Lead complex troubleshooting and root-cause analysis for production incidents, and collaborate with cross-functional teams to design solutions for Core IDCs and CDN/edge.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Production System Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 31 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Own production server lifecycles for large-scale data center and infrastructure services, from deployment and operation through decommissioning and recycling. Enhance server lifecycle processes via system design consultation, launch reviews, and continuous operations improvements. Build automation and monitoring tools to improve reliability, scalability, availability, and latency. Lead complex troubleshooting and root-cause analysis for production incidents, and collaborate with cross-functional teams to design solutions for Core IDCs and CDN/edge.
Location: London
Workplace: Onsite
Employment Type: Full time
Job Function: Manufacturing & Production Operations
Seniority: Mid level

Key Responsibilities

  • •Improve stability, efficiency, effectiveness, and scalability of data center and server operations, platform, and services globally.
  • •Enhance the full server fleet lifecycle from system design consultation to launch reviews, deployment, operation, and retirement.
  • •Develop and deploy automation tools and solutions to improve reliability, scalability, and operability of datacenter servers.
  • •Build monitoring tools to improve availability, latency, and datacenter infrastructure, server, and network health.
  • •Troubleshoot production incidents, perform root-cause analysis, and establish preventive measures as part of incident response and on-call support.

Key Requirements

  • •Bachelor’s degree in Computer Science, Electronic Engineering, a relevant technical field, or equivalent practical experience.
  • •5+ years of experience in server operations with in-depth Linux system administration knowledge, including kernels, drivers, and modules.
  • •Ability to script in Bash and Python to automate system configuration, performance tuning, and security management in Linux.
  • •In-depth understanding of server hardware and the ability to troubleshoot and diagnose complex faults.
  • •Over 5 years participating in planning, implementation, and operation of large-scale data centers across different countries.
Experience:5+ yearsData centers
Education:Bachelor's
Skills:Cross-team collaborationRoot-cause analysisIncident response
Tech Stack:LinuxLinux kernelBashPythonGPU serversCDNEdgeCloudNetwork health

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn