Senior Backend Engineer - AML Engine Orchestration

ByteDance
Singapore
Workplace: OnsiteFull timeFunction: Software EngineeringExperience: 3+ yearsEducation: bachelorsSkills: ["Analytical abilities","Problem solving","Communication","Self-motivation","Documentation"]

Build and extend distributed orchestration and scheduling systems for AML workloads, optimizing Kubernetes/Godel-based frameworks for cluster utilization, load balancing, autoscaling, and preemption-aware rescheduling across multi-datacenter and multi-cloud environments. Design training and online inference runtimes for ultra-large recommendation models, develop distributed computing APIs for future ML paradigms, and improve MLops workflows in collaboration with platform teams.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Senior Backend Engineer - AML Engine Orchestration

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 31 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Build and extend distributed orchestration and scheduling systems for AML workloads, optimizing Kubernetes/Godel-based frameworks for cluster utilization, load balancing, autoscaling, and preemption-aware rescheduling across multi-datacenter and multi-cloud environments. Design training and online inference runtimes for ultra-large recommendation models, develop distributed computing APIs for future ML paradigms, and improve MLops workflows in collaboration with platform teams.
Location: Singapore
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Develop and extend distributed orchestration frameworks in the Kubernetes/Godel ecosystem, optimizing cluster utilization and load balancing for different business scenarios.
  • •Integrate autoscaling and automatic parallelization, using load modeling to optimize resource requests and improve resource efficiency at scale.
  • •Implement preemption and re-scheduling mechanisms by service priority, and manage automatic resource multiplexing across clusters and resource types.
  • •Build distributed training runtimes for ultra-large/ultra-deep recommendation models, focusing on hyper-scaled embeddings and large-scale GPU training.
  • •Design online orchestration for distributed model inference and optimize online recommendation/ads architectures and MLops workflows.

Key Requirements

  • •Bachelor’s degree or above in Computer Science, Engineering, or a related field.
  • •Strong programming/coding experience with modern languages such as Golang or Python.
  • •Experience contributing to large-scale distributed and multi-tenant systems (architecture, reliability, scaling).
  • •Strong analytical abilities and problem solving.
  • •At least 3 years of relevant experience.
  • •Familiar with large-scale distributed scheduling systems such as Kubernetes, Yarn, Flink and/or Spark.
  • •Familiar with open-source orchestration frameworks such as VeRL, vLLM, Ray or TFX.
Experience:3+ years
Education:Bachelor's in Computer Science, Engineering or related fields
Skills:Analytical abilitiesProblem solvingCommunicationSelf-motivationDocumentation
Tech Stack:GolangPythonKubernetesGodelAutoScalingYarnFlinkSparkVeRLVLLMRayTFXGPU trainingMLopsReinforcement learningFine-tuningDistillationOnline inference

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn