Senior Software Engineer - Compute Infrastructure (Cloud Native)

ByteDance
Seattle
Workplace: OnsiteFull timeFunction: Software EngineeringExperience: 3+ yearsEducation: mastersSkills: []

Build and evolve ByteDance’s large-scale, Kubernetes-based compute infrastructure for AI/LLM workloads. You’ll improve Kubernetes performance across control and data planes, create observability and SLO-driven tuning in production, and develop intelligent resource management and scheduling at node and cluster level. The role also drives standardization and optimization of container runtime environments, contributing to open-source infrastructure projects like key Kubernetes initiatives.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Senior Software Engineer - Compute Infrastructure (Cloud Native)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 31 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Build and evolve ByteDance’s large-scale, Kubernetes-based compute infrastructure for AI/LLM workloads. You’ll improve Kubernetes performance across control and data planes, create observability and SLO-driven tuning in production, and develop intelligent resource management and scheduling at node and cluster level. The role also drives standardization and optimization of container runtime environments, contributing to open-source infrastructure projects like key Kubernetes initiatives.
Location: Seattle
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Design and evolve the architecture of large-scale Kubernetes-based infrastructure platforms for microservices, big data, and AI/LLM applications.
  • •Improve Kubernetes performance across control and data planes by optimizing pod lifecycle, resource orchestration, and throughput under high load.
  • •Create observability and performance analysis frameworks, define system-level SLOs, and lead data-driven tuning and optimization in production.
  • •Develop intelligent, unified resource management and scheduling systems at node and cluster level for cloud-native compute environments.
  • •Standardize and optimize container runtime environments to improve workload isolation, reliability, and resource efficiency across heterogeneous compute systems.

Key Requirements

  • •B.S./M.S. in Computer Science, Computer Engineering, or related field with 3+ years of relevant industry experience (Ph.D. with strong publications may be an exception).
  • •Solid understanding of at least one area such as Unix/Linux environments, distributed and parallel systems, high-performance networking, or large-scale software systems.
  • •Familiarity with container and orchestration technologies such as Docker and Kubernetes.
  • •Proficiency in at least one major programming language such as Python, Go, C++, Rust, or Java.
  • •Preferred: knowledge of big data or machine learning workflows in a Kubernetes environment and experience with cloud-native open-source projects.
Experience:3+ yearsCloud-nativeKubernetesDistributed systemsAI/MLOpen source
Education:Master's in Computer Science, Computer Engineering, or related area
Tech Stack:KubernetesDockerServerlessRayKubewharfAiBrixUnix/LinuxPrometheusGrafanaDistributed tracingMicroservicesBig dataCPUGPUPower

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn