Backend Engineer , AML Engine Orchestration

ByteDance
Singapore
Workplace: OnsiteFull timeFunction: Software EngineeringExperience: 3+ yearsEducation: bachelorsSkills: ["Analytical abilities","Problem solving","Communication","Self-motivation","Engineering practice"]

Build and extend distributed orchestration and scheduling frameworks that optimize resource efficiency across multi-datacenter, multi-region, and multi-cloud environments. Integrate autoscaling and parallelization, and implement preemption and re-scheduling with intelligent load balancing. You’ll also design training and online inference architectures for ultra-large recommendation models, including distributed APIs/runtimes and MLops workflow improvements for diagnosability and usability.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Backend Engineer , AML Engine Orchestration

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 31 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Build and extend distributed orchestration and scheduling frameworks that optimize resource efficiency across multi-datacenter, multi-region, and multi-cloud environments. Integrate autoscaling and parallelization, and implement preemption and re-scheduling with intelligent load balancing. You’ll also design training and online inference architectures for ultra-large recommendation models, including distributed APIs/runtimes and MLops workflow improvements for diagnosability and usability.
Location: Singapore
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Develop and extend distributed orchestration frameworks in the Kubernetes/Godel ecosystem, optimizing cluster utilization and load balancing for different business scenarios.
  • •Integrate autoscaling and automatic parallelization, using load modeling/analytic methods to optimize resource requests and improve global resource efficiency.
  • •Implement preemption and re-scheduling mechanisms for services with different priorities, and manage automatic resource multiplexing across clusters and resource types.
  • •Build a flexible, robust distributed training runtime for ultra-large/ultra-deep recommendation models, focused on hyper-scaled embeddings and large-scale GPU training.
  • •Construct online orchestration architecture for next-generation recommendation systems, including distributed inference for online learning and improvements to MLops workflows.

Key Requirements

  • •Bachelor’s degree or above in Computer Science, Engineering, or a related field.
  • •Strong programming and coding experience with at least one modern language such as Golang or Python.
  • •Experience contributing to large-scale distributed systems and multi-tenant systems, including architecture, reliability, and scaling.
  • •Strong analytical abilities and problem solving.
  • •Good communication, self-motivation, engineering practice, and documentation, with at least 3 years of relevant experience.
Experience:3+ yearsDistributed systemsMachine learningRecommendation systems
Education:Bachelor's in Computer Science, Engineering or related fields
Skills:Analytical abilitiesProblem solvingCommunicationSelf-motivationEngineering practice
Tech Stack:GolangPythonKubernetesGodelAutoScalingMulti-tenant systemsGPUYarnFlinkSparkVeRLVLLMRayTFXMLops

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn