Research Scientist - Compute AI Infra - Global Frontier Tech Recruitment Program - 2027 Start (PhD)

ByteDance
San Jose
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningEducation: phdSkills: ["Communication","Teamwork"]

Research and develop next-generation AI-native infrastructure for large-scale LLM and AI agent workloads. Work on AI training and inference infrastructure, fault localization and observability for AI clusters, serverless elastic storage and acceleration for AI scenarios, heterogeneous GPU/CPU/MEM power and scheduling systems, and low-latency, low-cost vector retrieval engines. Collaborate with academia and open source communities and publish academic papers.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Research Scientist - Compute AI Infra - Global Frontier Tech Recruitment Program - 2027 Start (PhD)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 30 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Research and develop next-generation AI-native infrastructure for large-scale LLM and AI agent workloads. Work on AI training and inference infrastructure, fault localization and observability for AI clusters, serverless elastic storage and acceleration for AI scenarios, heterogeneous GPU/CPU/MEM power and scheduling systems, and low-latency, low-cost vector retrieval engines. Collaborate with academia and open source communities and publish academic papers.
Location: San Jose
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Graduate level

Key Responsibilities

  • •Develop key technologies to optimize the AI infrastructure stack, including training infrastructure, inference infrastructure, and AI agents.
  • •Conduct systematic research across the AI infrastructure stack, including network and observability, storage systems, and data center power scheduling.
  • •Research and build capabilities for low-latency, low-cost vector retrieval engines and distributed vector index systems for LLM-powered applications.
  • •Explore intelligent infrastructure optimization driven by AI agent workflows and enable full-stack intelligent optimization through AI for infrastructure.
  • •Collaborate with academia and open source communities, and present research and products in academic papers.

Key Requirements

  • •Completing or recently completed a PhD in Software Development, Computer Science, Computer Engineering, Artificial Intelligence, or a related technical field.
  • •Experience in at least one area of LLM training infrastructure (e.g., RL training, knowledge distillation), LLM inference infrastructure, or AI agents/agent infrastructure (e.g., coding agents, agent memory, agent sandbox).
  • •Strong ability to continuously learn and quickly grasp and apply new AI technologies.
  • •Good communication and teamwork skills.
  • •Ability to commit to an onboarding date by end of year 2027 and clearly state availability and graduation date in the resume.
Experience:AI infrastructureLLM trainingLLM inferenceAI agentsOpen source
Education:PhD / Doctorate
Skills:CommunicationTeamwork
Tech Stack:LLMsRL trainingKnowledge distillationPyTorchVLLMSGLangTime-series databasesGPU kernel optimizationsVector retrievalFault localizationRoot cause analysisServerlessDPU

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn