Backend Inference Framework Engineer Graduate (AML Inference) - 2027 Start

ByteDance
San Jose
Workplace: OnsiteFull timeFunction: Software EngineeringSkills: ["Learning ability","Execution ability","Problem analysis","Cross-team collaboration","Communication"]

Design and implement large-model inference services for high-concurrency recommendation/advertising use cases, delivering scalable, highly available systems. Build and optimize the inference framework core (scheduling, monitoring/alerting, canary releases), resolving performance and stability bottlenecks. Drive technology selection and innovation by researching the latest inference approaches and improving distributed service architecture and standardization across business scenarios.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
2 hours ago

Backend Inference Framework Engineer Graduate (AML Inference) - 2027 Start

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 2 hours agoStatus: Live

Job Summary

Design and implement large-model inference services for high-concurrency recommendation/advertising use cases, delivering scalable, highly available systems. Build and optimize the inference framework core (scheduling, monitoring/alerting, canary releases), resolving performance and stability bottlenecks. Drive technology selection and innovation by researching the latest inference approaches and improving distributed service architecture and standardization across business scenarios.
Location: San Jose
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Graduate level

Key Responsibilities

  • •Own the architecture design and implementation of model inference services for large-parameter, high-complexity AI models, enabling scalable enterprise-level inference.
  • •R&D and optimize core inference framework modules, including inference engine scheduling, monitoring/alerting, and canary release.
  • •Continuously improve framework performance by identifying and resolving performance bottlenecks, resource bottlenecks, and stability issues in large-model, high-concurrency inference.
  • •Track and evaluate the latest inference technologies, perform technology selection, and drive innovations tailored to business scenarios.
  • •Accumulate distributed high-concurrency service architecture solutions and promote technical system upgrades and standardization.

Key Requirements

  • •Completing or recently completed a Bachelor's/Master's in Software Development, Computer Science, Computer Engineering, or a related technical discipline.
  • •Proficiency with basic Linux commands, strong C/C++ skills, and knowledge of data structures and algorithms.
  • •Understanding of multi-threaded concurrency, including thread usage, synchronization, thread pools, and basic performance tuning.
  • •Experience with R&D projects involving high-concurrency distributed services and familiarity with latency/resource optimization.
  • •Strong learning, execution, problem analysis/abstraction, and cross-team collaboration skills.
Skills:Learning abilityExecution abilityProblem analysisCross-team collaborationCommunication
Tech Stack:LinuxCC++Multi-threadingThread poolsRedisRocksDBBRPCGRPCGPU

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn