Software Engineer - Compute Infrastructure (Cloud Native)

ByteDance
San Jose
Workplace: OnsiteFull timeFunction: Software EngineeringEducation: bachelorsSkills: []

Build and evolve Kubernetes-based compute infrastructure that powers large-scale clusters for AI/LLM workloads. Improve Kubernetes control/data-plane performance, optimize pod lifecycle and orchestration, and establish observability with SLOs to drive production tuning. Design unified resource management and scheduling for node and cluster-level compute, and standardize container runtime environments for isolation, reliability, and cost efficiency. Contribute to cloud-native open-source initiatives across the K8s ecosystem.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Software Engineer - Compute Infrastructure (Cloud Native)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 31 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Build and evolve Kubernetes-based compute infrastructure that powers large-scale clusters for AI/LLM workloads. Improve Kubernetes control/data-plane performance, optimize pod lifecycle and orchestration, and establish observability with SLOs to drive production tuning. Design unified resource management and scheduling for node and cluster-level compute, and standardize container runtime environments for isolation, reliability, and cost efficiency. Contribute to cloud-native open-source initiatives across the K8s ecosystem.
Location: San Jose
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Design and evolve the architecture of large-scale Kubernetes-based infrastructure platforms for diverse workloads, including microservices, big data, and AI/LLM applications.
  • •Improve Kubernetes system performance across control and data planes by optimizing pod lifecycle, resource orchestration, and throughput under high load.
  • •Build observability and performance analysis frameworks, define Kubernetes system-level SLOs, and lead data-driven tuning and optimization in production.
  • •Develop intelligent unified resource management and scheduling systems at node and cluster level for large-scale cloud-native compute resources.
  • •Standardize and optimize container runtime environments to improve workload isolation, reliability, and resource efficiency across heterogeneous compute environments.

Key Requirements

  • •B.S./M.S. in Computer Science, Computer Engineering, or related field with 3+ years of relevant industry experience; Ph.D. with strong publication records may substitute.
  • •Solid understanding of at least one: Unix/Linux environments, distributed and parallel systems, high-performance networking systems, or developing large-scale software systems.
  • •Familiarity with container and orchestration technologies such as Docker and Kubernetes.
  • •Proficiency in at least one major programming language such as Python, Go, C++, Rust, or Java.
  • •Hands-on experience in cloud-native environments (e.g., big data or ML workflows in Kubernetes) and familiarity with observability tools is preferred.
Experience:AI/LLMCloud-nativeKubernetesOpen-sourceBig dataMachine learning
Education:Bachelor's in Computer Science, Computer Engineering or a related area
Tech Stack:KubernetesK8sServerlessDockerRayPrometheusGrafanaDistributed tracingUnix/LinuxPythonGoC++RustJavaKubewharfAiBrixMicroservicesCPUGPUPower

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn