Machine Learning Engineer - Inference

ByteDance
San Jose
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningExperience: 3+ yearsSkills: ["Collaboration","Problem solving","Performance optimization"]

Build and scale distributed inference infrastructure powering ByteDance’s feed and ads ranking, search ranking, and live/e-commerce recommendations. You’ll develop monitoring and management tools for reliable, scalable online inference servers, diagnose bottlenecks and instability, and implement performance improvements. Collaborate with product teams to deliver inference solutions that meet their requirements, while contributing to the AML team’s next-generation AI infrastructure and recommendation platform.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Machine Learning Engineer - Inference

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 30 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Build and scale distributed inference infrastructure powering ByteDance’s feed and ads ranking, search ranking, and live/e-commerce recommendations. You’ll develop monitoring and management tools for reliable, scalable online inference servers, diagnose bottlenecks and instability, and implement performance improvements. Collaborate with product teams to deliver inference solutions that meet their requirements, while contributing to the AML team’s next-generation AI infrastructure and recommendation platform.
Location: San Jose
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Mid level

Key Responsibilities

  • •Design and implement distributed inference infrastructure for feeds, ads, and search ranking models.
  • •Build monitoring and management tools to ensure reliability and scalability of online inference servers.
  • •Triaging system inefficiency and bottlenecks to improve system performance.
  • •Analyze bottlenecks and instability sources, then design and implement solutions.
  • •Collaborate with product teams to provide solutions that meet their requirements.

Key Requirements

  • •At least 3 years of experience developing and deploying large-scale systems.
  • •Contribute to an open source machine learning framework (TensorFlow, JAX, PyTorch, TorchScript, MXNet, or TensorRT).
  • •Strong background in Hardware-Software Co-Design, High Performance Computing, ML hardware acceleration (e.g., GPU/RDMA), or ML for Systems.
Experience:3+ yearsOpen sourceMachine learning
Skills:CollaborationProblem solvingPerformance optimization
Tech Stack:Machine learningTensorFlowJAXPyTorchTorchScriptMXNetTensorRTGPURDMADistributed inference

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn