Research scientist - AI Infra - Global Frontier Tech Recruitment Program - 2027 Start (PhD)

ByteDance
San Jose
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningEducation: phdSkills: ["Communication","Collaboration","Learning agility","Technical research","Cross-team alignment","Ability to evaluate trade-offs"]

Join a lean research team defining next-generation AI infrastructure. You’ll design and evaluate scalable “AI factory” architectures spanning compute, storage, networking, chips, power, and data/application layers for large-scale training, RL, and inference. Drive research on AI-native infrastructure topics such as fault localization and observability, elastic storage and acceleration, power scheduling, vector retrieval, and agent intelligence—then collaborate cross-team to prototype, benchmark, and translate findings into production-ready innovations.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Research scientist - AI Infra - Global Frontier Tech Recruitment Program - 2027 Start (PhD)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 30 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Join a lean research team defining next-generation AI infrastructure. You’ll design and evaluate scalable “AI factory” architectures spanning compute, storage, networking, chips, power, and data/application layers for large-scale training, RL, and inference. Drive research on AI-native infrastructure topics such as fault localization and observability, elastic storage and acceleration, power scheduling, vector retrieval, and agent intelligence—then collaborate cross-team to prototype, benchmark, and translate findings into production-ready innovations.
Location: San Jose
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Graduate level

Key Responsibilities

  • •Design and evaluate scalable AI-factory architectures across compute, storage, networking, chips, power, and data/application layers for large-scale training, RL, and inference workloads.
  • •Develop technical proposals addressing supply-chain and energy constraints alongside silicon and software trade-offs.
  • •Track and explore emerging trends in AI systems, distributed training/RL, and hardware acceleration; build prototypes and share insights via technical reports.
  • •Analyze and optimize system performance across the ML stack—scheduling, networking, storage, training/RL frameworks, and AI memory systems—using benchmarking and bottleneck analysis.
  • •Collaborate across research, engineering, hardware, data-center, and product teams to translate workload requirements into scalable solutions and drive cross-team initiatives.

Key Requirements

  • •Completing or recently completed a PhD in Computer Science, Computer Engineering, Electrical Engineering, or a related technical discipline; cognitive science/computational neuroscience/psychology backgrounds welcome with strong systems fundamentals.
  • •Experience with distributed systems, infrastructure engineering, or ML systems, including exposure to large-scale training or RL pipelines, with comfort evaluating trade-offs across hardware, software, algorithms, energy, and supply-chain constraints.
  • •Strong proficiency integrating AI tools into knowledge discovery and research workflows.
  • •Demonstrated ability to learn quickly and remain productive on a fast-evolving technical horizon.
  • •Excellent communication skills to collaborate across teams.
Experience:Distributed systemsML systemsLarge-scale model trainingReinforcement learningHPC-style distributed workloadsAI infrastructureKnowledge discoveryOpen-source
Education:PhD / Doctorate in PhD
Skills:CommunicationCollaborationLearning agilityTechnical researchCross-team alignmentAbility to evaluate trade-offs
Tech Stack:AI infrastructureLLMsAI agentsPretrainingReinforcement learning (RL)Agentic workloadsLarge-scale trainingDistributed trainingInferenceDistributed systemsSchedulingNetworkingStorage systemsTime-series databasesFault localizationRoot cause analysisBenchmarkingBottleneck analysisRetrieval-augmented architecturesVector retrieval

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn