Senior Production System Engineer - San Jose

ByteDance
New York
Full timeFunction: Manufacturing & Production OperationsEducation: bachelorsSkills: ["Systems-thinking","Technical leadership","Communication"]

Own the end-to-end production readiness and lifecycle management of ByteDance’s large-scale GPU infrastructure across self-built data centers. Lead evaluation, qualification, integration, and rollout of next-generation platforms (e.g., NVIDIA GB200/GB300 NVL72, Vera Rubin), driving fleet onboarding, monitoring, incident response, and long-term reliability. Partner across hardware, networking, storage, power, cooling, and vendors while building automation, telemetry, and AI-assisted operations tooling for scalable GPU operations.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Senior Production System Engineer - San Jose

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 30 days agoStatus: Live

Job Summary

Own the end-to-end production readiness and lifecycle management of ByteDance’s large-scale GPU infrastructure across self-built data centers. Lead evaluation, qualification, integration, and rollout of next-generation platforms (e.g., NVIDIA GB200/GB300 NVL72, Vera Rubin), driving fleet onboarding, monitoring, incident response, and long-term reliability. Partner across hardware, networking, storage, power, cooling, and vendors while building automation, telemetry, and AI-assisted operations tooling for scalable GPU operations.
Location: New York
Employment Type: Full time
Job Function: Manufacturing & Production Operations
Seniority: Mid level

Key Responsibilities

  • •Lead evaluation, qualification, integration, and production rollout of next-generation GPU platforms (e.g., GB200/GB300 NVL72, B200/B300, Vera Rubin, rack-scale AI infrastructure).
  • •Define production readiness plans and launch criteria across server hardware, firmware, BMC, OS, drivers, GPU software stacks, networking, storage, security, telemetry, and operational tooling.
  • •Drive rack-scale integration and fleet operations with hardware, data center, network, storage, power, cooling, and vendor teams to improve availability, utilization, serviceability, and lifecycle management.
  • •Diagnose complex Linux, hardware, firmware, PCIe, NVLink/NVSwitch, network, storage, and memory issues; develop qualification, burn-in, benchmarking, health-check, and regression-testing frameworks.
  • •Build automation and actionable observability telemetry for provisioning, configuration, monitoring, fault detection, remediation, repair, and lifecycle operations; provide on-call leadership during critical incidents and coordinate root-cause analysis and preventive actions.
Travel: Low travel

Key Requirements

  • •5+ years in production systems, infrastructure engineering, Site Reliability Engineering, DevOps, hardware systems engineering, or large-scale data center operations.
  • •Hands-on experience introducing and productionizing large-scale GPU infrastructure, including hardware NPI and/or fleet onboarding through qualification, integration, deployment, and post-launch reliability.
  • •Deep Linux administration and troubleshooting with strong understanding of server architecture and management technologies (kernels, drivers, BIOS/UEFI, BMC, Redfish, firmware, PCIe, NVMe, NICs, DPUs, hardware telemetry).
  • •Hands-on experience deploying or supporting distributed AI training/inference workloads in containerized and orchestrated environments, including CUDA, GPU drivers, NCCL, NVLink/NVSwitch, RDMA, InfiniBand/RoCE, or high-performance Ethernet.
  • •Proficiency in Python, Go, Bash (or similar) for production-grade infrastructure automation, plus experience applying AI/LLMs to engineering workflows like incident analysis and automated remediation.
Experience:AI infrastructureData centersDevOpsSite Reliability EngineeringHardware lifecycle management
Education:Bachelor's
Skills:Systems-thinkingTechnical leadershipCommunication
Languages:English
Tech Stack:LinuxNVIDIA GB200NVIDIA GB300NVL72HGXDGX B200DGX B300Vera RubinCUDANCCLNVLinkNVSwitchRDMAInfiniBandRoCEHigh-performance EthernetPythonGoBashKernels

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn