Tech Lead, Research Scientist/Engineer - AI Infrastructure

ByteDance
San Jose
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningEducation: phdSkills: ["Communication","Rapid learning"]

Define and build next-generation AI infrastructure at ByteDance, translating evolving AI requirements into scalable architectures across compute, storage, networking, chips, power, and data/application layers. Own end-to-end AI factory architecture for large-scale training, RL, and inference, exploring emerging systems and hardware trends. Benchmark and optimize ML stack and AI memory systems for long-horizon agents, while aligning across research, engineering, hardware, and product teams.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Tech Lead, Research Scientist/Engineer - AI Infrastructure

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 29 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Define and build next-generation AI infrastructure at ByteDance, translating evolving AI requirements into scalable architectures across compute, storage, networking, chips, power, and data/application layers. Own end-to-end AI factory architecture for large-scale training, RL, and inference, exploring emerging systems and hardware trends. Benchmark and optimize ML stack and AI memory systems for long-horizon agents, while aligning across research, engineering, hardware, and product teams.
Location: San Jose
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Mid level

Key Responsibilities

  • •Design and evaluate scalable AI factory architectures across compute, storage, networking, chips, power, and data/application layers for training, RL, and inference.
  • •Develop technical proposals that balance supply-chain and energy constraints with silicon and software trade-offs.
  • •Track and explore emerging trends across AI systems, distributed training/RL, and hardware acceleration, including cognitive-science and psychology influences on AI memory and reasoning.
  • •Analyze and optimize ML stack performance (scheduling, networking, storage, training/RL frameworks, and long-horizon AI memory) via benchmarking and bottleneck analysis.
  • •Collaborate across research, engineering, hardware, data-center, and product teams to translate workload requirements into scalable solutions and drive cross-team initiatives.

Key Requirements

  • •PhD (or recently completed) in Computer Science, Computer Engineering, Electrical Engineering, or a related technical discipline.
  • •Experience in distributed systems, infrastructure engineering, or ML systems, including exposure to large-scale training or RL pipelines.
  • •Strong ability to integrate AI tools into knowledge discovery and research workflows.
  • •Ability to learn quickly and stay productive on a fast-evolving technical horizon.
  • •Excellent communication skills to collaborate across teams.
Experience:AI infrastructureDistributed systemsMachine learningRLHigh-performance computing
Education:PhD / Doctorate
Skills:CommunicationRapid learning
Tech Stack:AI infrastructureAI memory systemsDistributed trainingReinforcement learningRL pipelinesSchedulingHigh-performance networkingRDMANCCLGPUAcceleratorsKV cacheRetrieval-augmented architectures

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn