Senior Software Engineer - Compute Infrastructure (Cloud Native)

ByteDance
San Jose
Workplace: OnsiteFull timeFunction: Software EngineeringEducation: bachelorsSkills: []

Build and evolve large-scale Kubernetes-based compute infrastructure for diverse workloads, including microservices, big data, and AI/LLM applications. Improve Kubernetes performance across control and data planes, optimize pod lifecycle and orchestration, and deliver data-driven tuning in production. Develop observability frameworks and define system-level SLOs. Create unified resource management and scheduling systems and standardize container runtime environments to improve isolation, reliability, and cost efficiency.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Senior Software Engineer - Compute Infrastructure (Cloud Native)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 31 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Build and evolve large-scale Kubernetes-based compute infrastructure for diverse workloads, including microservices, big data, and AI/LLM applications. Improve Kubernetes performance across control and data planes, optimize pod lifecycle and orchestration, and deliver data-driven tuning in production. Develop observability frameworks and define system-level SLOs. Create unified resource management and scheduling systems and standardize container runtime environments to improve isolation, reliability, and cost efficiency.
Location: San Jose
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Design and evolve the architecture of large-scale Kubernetes-based infrastructure platforms for performance, scalability, and resilience across workloads.
  • •Improve Kubernetes system performance across control and data planes by optimizing pod lifecycle, resource orchestration, and throughput under high load.
  • •Build observability and performance analysis frameworks, define Kubernetes system-level SLOs, and lead data-driven tuning and optimization in production.
  • •Develop intelligent, unified resource management and scheduling systems at node and cluster level for diverse compute resources in cloud-native environments.
  • •Standardize and optimize container runtime environments to improve workload isolation, reliability, and resource efficiency across heterogeneous compute environments.

Key Requirements

  • •B.S./M.S. in Computer Science, Computer Engineering, or related field with 3+ years of relevant experience; Ph.D. with strong publications may be an exception.
  • •Solid understanding of at least one area such as Unix/Linux environments, distributed and parallel systems, high-performance networking, or building large-scale software systems.
  • •Familiarity with container and orchestration technologies such as Docker and Kubernetes.
  • •Proficiency in at least one major programming language: Python, Go, C++, Rust, or Java.
  • •Knowledge of big data or machine learning workflows in a Kubernetes environment (preferred).
Education:Bachelor's in Computer Science, Computer Engineering or related area
Tech Stack:KubernetesDockerServerlessRayPrometheusGrafanaUnix/LinuxLinuxPythonGoC++RustJavaDistributed tracingMicroservicesAI/LLMAiBrixKubewharf

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn