Backend Engineer , AML Engine Orchestration

TikTok
Singapore
Workplace: OnsiteFull timeFunction: Software EngineeringExperience: 3+ yearsEducation: bachelorsSkills: ["Analytical thinking","Problem-solving","Communication","Self-motivation","Engineering practice","Documentation"]

Build and extend distributed orchestration and scheduling systems for AML-driven recommendation, ads ranking, and search ranking. Optimize Kubernetes/Godel-based cluster utilization with autoscaling, parallelization, load modeling, preemption, and multi-datacenter/multi-region/multi-cloud scheduling. Design training runtimes for ultra-large GPU models, implement distributed computing APIs, and develop online inference orchestration while improving MLops workflows and diagnosability.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
TikTok
TikTok
1 month ago

Backend Engineer , AML Engine Orchestration

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 3 hours agoStatus: Live

Job Summary

Build and extend distributed orchestration and scheduling systems for AML-driven recommendation, ads ranking, and search ranking. Optimize Kubernetes/Godel-based cluster utilization with autoscaling, parallelization, load modeling, preemption, and multi-datacenter/multi-region/multi-cloud scheduling. Design training runtimes for ultra-large GPU models, implement distributed computing APIs, and develop online inference orchestration while improving MLops workflows and diagnosability.
Location: Singapore
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Develop and extend distributed orchestration frameworks within the Kubernetes/Godel ecosystem, optimizing cluster utilization and load balancing for different business scenarios.
  • •Integrate and expand autoscaling and automatic parallelization, using load modeling to optimize resource requests for improved efficiency.
  • •Implement preemption and re-scheduling mechanisms, and manage automatic resource multiplexing across clusters and resource types across multi-datacenter, multi-region, and multi-cloud environments.
  • •Design training system architecture and distributed training runtimes for ultra-large/ultra-deep recommendation models, including hyper-scaled embeddings and large-scale GPU training.
  • •Build online orchestration and distributed inference architecture for online learning, and optimize MLops workflows and usability of recommendation/ads model architectures.

Key Requirements

  • •Bachelor’s degree or above in Computer Science, Engineering, or related fields.
  • •Strong programming experience in at least one modern language such as Go (Golang) or Python.
  • •Experience contributing to large-scale distributed and multi-tenant systems (architecture, reliability, scalability).
  • •At least 3 years of relevant experience.
  • •Strong analytical/problem-solving skills and good communication, self-motivation, engineering practice, and documentation.
Experience:3+ yearsMachine learningDistributed systemsMulti-tenant systemsRecommendation systems
Education:Bachelor's in Computer Science, Engineering or related fields
Skills:Analytical thinkingProblem-solvingCommunicationSelf-motivationEngineering practiceDocumentation
Tech Stack:GolangPythonKubernetesGodelAutoScalingGPUYarnFlinkSparkVeRLVLLMRayTFXMLops

Company Brief

TikTok
Short-form video platform that lets users create, share, and discover entertainment content through algorithmic recommendations. It also offers advertising and creator tools for brands, influencers, and businesses.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Headquarters: Singapore, Singapore
Founded: 2016
WebsiteLinkedIn