Production Engineer - Applied Machine Learning

ByteDance
San Jose
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningSkills: ["Responsibility","Logical thinking","Cross-functional collaboration","Fault diagnosis","Performance analysis"]

Support and advance an applied machine learning (AML) platform by managing production training, inference, and storage systems. Own reliability practices including SLO/SLA, observability, alerting, on-call, fault diagnosis, auto-healing, disaster recovery, and incident reviews. Improve engineering excellence with CI/CD, canary releases, auto-rollback, pre-flight checks, capacity forecasting, and elastic scaling, while governing GPU/CPU/storage/network resources and costs.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Production Engineer - Applied Machine Learning

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 30 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Support and advance an applied machine learning (AML) platform by managing production training, inference, and storage systems. Own reliability practices including SLO/SLA, observability, alerting, on-call, fault diagnosis, auto-healing, disaster recovery, and incident reviews. Improve engineering excellence with CI/CD, canary releases, auto-rollback, pre-flight checks, capacity forecasting, and elastic scaling, while governing GPU/CPU/storage/network resources and costs.
Location: San Jose
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning

Key Responsibilities

  • •Manage production operations and stability for AML training, inference, and storage systems, including scheduling/orchestration and distributed components.
  • •Build and maintain reliability mechanisms such as SLO/SLA governance, observability/alerting, on-call processes, fault diagnosis, auto-healing, disaster recovery, and post-mortems.
  • •Drive engineering excellence with CI/CD, canary releases, auto-rollback, automated inspections, pre-flight checks, capacity forecasting, and elastic auto-scaling.
  • •Oversee resource governance across GPU/CPU/storage/network including quota management, cost attribution, and performance tuning.
  • •Support cross-functional issue resolution to maintain system availability and R&D efficiency.

Key Requirements

  • •Bachelor’s degree or above in Computer Science, Software Engineering, Artificial Intelligence, or related fields.
  • •Familiar with Linux and proficient in at least one of: Shell, Python, Go, or C++.
  • •Understanding of machine learning training/inference architectures, Kubernetes, GPU clusters, or distributed storage systems.
  • •Proven experience in online troubleshooting, performance analysis, and building automation platforms.
  • •Strong sense of responsibility, clear logical thinking, and ability to drive resolution across cross-functional teams.
Experience:Machine learningDistributed systemsReliability engineering
Education:
Skills:ResponsibilityLogical thinkingCross-functional collaborationFault diagnosisPerformance analysis
Tech Stack:LinuxShellPythonGoC++KubernetesGPU clustersCI/CDCanary releasesAuto-rollbackDistributed trainingOnline inference servingParameterServerNoSQLSLO/SLAObservabilityAlertingDisaster recoveryIncident reviewsCapacity forecasting

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn