Production System Engineer Graduate (Server Management) - 2027 Start

ByteDance
San Jose
Workplace: OnsiteFull timeFunction: Manufacturing & Production OperationsEducation: bachelorsSkills: ["Analytical thinking","Troubleshooting","Quick learning","Communication","Cross-functional collaboration"]

Join the Server Management DevOps team to support end-to-end lifecycle operations for large-scale server fleets across ByteDance’s self-built data centers. Work with Linux production environments to deploy, validate, monitor, and troubleshoot CPU/GPU infrastructure, while building automation using Python, Bash, and Go. Gain exposure to AI infrastructure and GPU reliability improvements, analyze operational metrics, and collaborate cross-functionally on infrastructure tooling and AI-assisted operations.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 day ago

Production System Engineer Graduate (Server Management) - 2027 Start

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 13 hours agoStatus: Live

Job Summary

Join the Server Management DevOps team to support end-to-end lifecycle operations for large-scale server fleets across ByteDance’s self-built data centers. Work with Linux production environments to deploy, validate, monitor, and troubleshoot CPU/GPU infrastructure, while building automation using Python, Bash, and Go. Gain exposure to AI infrastructure and GPU reliability improvements, analyze operational metrics, and collaborate cross-functionally on infrastructure tooling and AI-assisted operations.
Location: San Jose
Workplace: Onsite
Employment Type: Full time
Job Function: Manufacturing & Production Operations
Seniority: Graduate level

Key Responsibilities

  • •Assist with deployment, validation, monitoring, maintenance, and lifecycle management for large-scale CPU/GPU server fleets.
  • •Develop scripts, tools, and automation solutions to reduce manual operations and improve infrastructure efficiency.
  • •Work in Linux-based production environments to troubleshoot operating system, hardware, storage, networking, and performance issues.
  • •Contribute to operational tooling and reliability improvements for GPU and AI infrastructure platforms.
  • •Analyze server health, hardware failures, and operational metrics to identify trends, risks, and improvement opportunities.

Key Requirements

  • •Completing or recently completed a Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering, Information Technology, or a related technical field.
  • •Experience in systems engineering, infrastructure operations, DevOps, Site Reliability Engineering, or related technical roles (or equivalent hands-on project experience).
  • •Solid understanding of Linux (Debian or Ubuntu preferred) administration and troubleshooting.
  • •Programming/scripting experience in Python, Bash, Go, or another modern language.
  • •Understanding of operating systems, computer architecture, networking fundamentals, and storage systems.
Experience:Infrastructure operationsDevOpsSite reliability engineeringAI infrastructure
Education:Bachelor's in Computer Science, Computer Engineering, Electrical Engineering, Information Technology, or related technical field
Skills:Analytical thinkingTroubleshootingQuick learningCommunicationCross-functional collaboration
Tech Stack:LinuxDebianUbuntuPythonBashGoNVIDIA GPUCUDATCP/IPDNSDHCPVLANRoutingMonitoringObservabilityLoggingTelemetryGitDockerKubernetes

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn