Machine Learning Engineer Graduate (AML-Engine-Orchestration) - 2027 Start (PhD)

ByteDance
San Jose
Workplace: OnsiteInternshipFunction: Data Science & Machine LearningEducation: phdSkills: ["Ownership","Collaboration","Problem-solving","Quantitative analysis","Experimentation"]

Join the Data-AML-Engine Orchestration team to build large-scale ML infrastructure that powers online model serving across ByteDance products like TikTok. You’ll design orchestration and scheduling systems, including Kubernetes operators, multi-tenant resource/quota management, and lifecycle orchestration for deployment, upgrades, rollbacks, autoscaling, and disaster recovery. You’ll also develop serving orchestration and traffic management for disaggregated clusters, optimizing GPU utilization, latency, reliability, and MLE productivity.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
15 hours ago

Machine Learning Engineer Graduate (AML-Engine-Orchestration) - 2027 Start (PhD)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 15 hours agoStatus: Live

Job Summary

Join the Data-AML-Engine Orchestration team to build large-scale ML infrastructure that powers online model serving across ByteDance products like TikTok. You’ll design orchestration and scheduling systems, including Kubernetes operators, multi-tenant resource/quota management, and lifecycle orchestration for deployment, upgrades, rollbacks, autoscaling, and disaster recovery. You’ll also develop serving orchestration and traffic management for disaggregated clusters, optimizing GPU utilization, latency, reliability, and MLE productivity.
Location: San Jose
Workplace: Onsite
Employment Type: Internship
Job Function: Data Science & Machine Learning
Seniority: Graduate level

Key Responsibilities

  • •Design and build foundational ML platform orchestration capabilities, including Kubernetes operators, container runtimes, and job/service/lifecycle management.
  • •Build multi-tenant resource and quota systems with priorities, preemption, fair sharing, elasticity, and cross-cluster scheduling to improve GPU utilization and cost efficiency.
  • •Build lifecycle orchestration for online model serving, covering model/image distribution, deployment, upgrades, rollback, autoscaling, multi-cluster operations, and disaster recovery.
  • •Build serving orchestration and traffic management for disaggregated serving clusters, including topology-aware scheduling, KV cache affinity, request routing, and QoS/SLA management.

Key Requirements

  • •Currently completing or recently completed a PhD in Computer Science, Software Engineering, Artificial Intelligence, or a related technical field.
  • •Proficiency in at least one of Go, C++, or Python, with strong foundations in data structures, algorithms, and software engineering.
  • •Familiarity with Linux plus fundamentals of operating systems, computer networks, concurrent programming, and distributed systems.
  • •Hands-on exploratory skills using source code, metrics, logs, profiling, and experiments to investigate systems.
  • •A systematic, quantitative approach to define measurements, test hypotheses, and validate improvements.
Experience:Open-source infrastructureDistributed systemsMachine learningModel servingOnline services
Education:PhD / Doctorate in Computer Science, Software Engineering, Artificial Intelligence, or a related technical field
Skills:OwnershipCollaborationProblem-solvingQuantitative analysisExperimentation
Tech Stack:KubernetesKubernetes OperatorsContainer runtimesGoC++PythonLinuxFinOpsKV CacheAutoscalingDisaster recoveryDistributed systemsGPUNPUTraffic management

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn