Software Engineer - Compute Infrastructure (Orchestration & Scheduling)

ByteDance
Seattle
Workplace: OnsiteFull timeFunction: Software EngineeringExperience: 2+ yearsEducation: bachelorsSkills: ["Communication","Teamwork"]

Build and evolve ByteDance’s Kubernetes- and serverless-based compute infrastructure that powers large-scale clusters for AI/LLM workloads worldwide. You’ll enhance hyper-scale cluster management, design unified scheduling for diverse containers and VMs, and develop AI-assisted scheduling to improve performance and resource efficiency across CPU, GPU, memory, network, and power. Contribute high-quality code and open-source capabilities for next-gen ML training and inference platforms.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Software Engineer - Compute Infrastructure (Orchestration & Scheduling)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 30 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Build and evolve ByteDance’s Kubernetes- and serverless-based compute infrastructure that powers large-scale clusters for AI/LLM workloads worldwide. You’ll enhance hyper-scale cluster management, design unified scheduling for diverse containers and VMs, and develop AI-assisted scheduling to improve performance and resource efficiency across CPU, GPU, memory, network, and power. Contribute high-quality code and open-source capabilities for next-gen ML training and inference platforms.
Location: Seattle
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Enhance Kubernetes-based cluster platforms for performance, scalability, and resilience across ByteDance’s global infrastructure.
  • •Design and maintain unified scheduling for diverse workloads including containers, VMs, online services, offline computing, and AI/ML CPU/GPU workloads.
  • •Develop intelligent scheduling using AI models to optimize workload performance and resource utilization across heterogeneous resources and global data centers.
  • •Design and drive next-gen compute platforms for fast, reliable, and cost-effective ML and LLM training/inference.
  • •Write high-quality, maintainable code and stay current with open-source and research advancements in AI, ML, systems, and serverless technologies.

Key Requirements

  • •B.S./M.S. in Computer Science, Computer Engineering, or related field with 2+ years of relevant experience; Ph.D. new graduates with strong publications may be considered.
  • •Solid understanding of Unix/Linux environments, distributed and parallel systems, high-performance networking, and large-scale software systems.
  • •Proven experience designing, architecting, and building cloud and ML infrastructure for resource management, allocation, job scheduling, and monitoring.
  • •Familiarity with container and orchestration technologies such as Docker and Kubernetes.
  • •Proficiency in at least one major programming language such as Python, Go, C++, Rust, or Java.
Experience:2+ yearsAI/MLLLMCloud infrastructureDistributed systemsKubernetesServerless
Education:Bachelor's in Computer Science or Computer Engineering
Skills:CommunicationTeamwork
Tech Stack:KubernetesDockerServerlessRayPythonGoC++RustJavaUnix/LinuxAWSAzureGCPAWS SageMakerAzure MLGCP Vertex AIYarnMesosKubewharfAiBrix

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn