Backend Inference Framework Engineer Graduate (AML Inference) - 2027 Start

ByteDance
Singapore
Workplace: OnsiteFull timeFunction: Software EngineeringSkills: ["Learning","Problem analysis","Cross-team collaboration","Communication","Stress tolerance"]

Design and implement scalable model inference services for high-parameter, high-complexity AI models. Develop and optimize core inference framework modules, including scheduling, monitoring/alerting, and canary releases, to improve performance, stability, and resource utilization under high concurrency. Stay current with inference technologies, select and innovate for business scenarios, and help standardize the team’s technical system across ByteDance’s ML mid-platform.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
3 days ago

Backend Inference Framework Engineer Graduate (AML Inference) - 2027 Start

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 12 hours agoStatus: Live

Job Summary

Design and implement scalable model inference services for high-parameter, high-complexity AI models. Develop and optimize core inference framework modules, including scheduling, monitoring/alerting, and canary releases, to improve performance, stability, and resource utilization under high concurrency. Stay current with inference technologies, select and innovate for business scenarios, and help standardize the team’s technical system across ByteDance’s ML mid-platform.
Location: Singapore
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Graduate level

Key Responsibilities

  • •Design and implement the overall architecture of model inference services for scalable, highly available enterprise-level inference systems.
  • •Build inference framework capabilities including inference engine scheduling, monitoring/alerting, and canary release.
  • •Continuously iterate on framework performance and resolve performance, resource, and stability bottlenecks under high-concurrency inference.
  • •Track the latest inference technologies and perform technology selection and innovation aligned to business scenarios.
  • •Accumulate distributed high-concurrency service architecture solutions and drive upgrade and standardization of the team’s technical system.

Key Requirements

  • •Completing or recently completed a Bachelor's or Master's degree in computing or a related discipline.
  • •Strong Linux command familiarity and solid C/C++ programming skills, including data structures and algorithms.
  • •Knowledge of multi-threaded concurrency (thread usage, synchronization locks, thread pools), with ability to identify concurrency issues and perform basic performance tuning.
  • •Experience in R&D projects for high-concurrency distributed services, including service latency and resource optimization.
  • •Ability to learn inference service architecture and translate problems into abstractions, with strong collaboration, communication, presentation, documentation, and stress tolerance.
Experience:Machine learningDistributed systemsHigh-concurrency services
Skills:LearningProblem analysisCross-team collaborationCommunicationStress tolerance
Tech Stack:LinuxCC++RedisRocksDBBRPCGRPCGPUThread poolsMulti-threading

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn