Research Scientist - AI Compute & DPU - Global Frontier Tech Recruitment Program - 2027 Start (PhD)

ByteDance
San Jose
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningEducation: phdSkills: ["System-building","Research","Ownership","Technical communication","Rigorous reasoning"]

Conduct PhD-level research for an AI infrastructure team that owns the AI Compute layer and DPU, including the open-source Kubernetes-native control plane AIBrix. Work on systematic research across the full stack—network observability, storage systems, data center power scheduling, vector retrieval, and AI-agent-driven intelligence—aiming to improve performance, cost efficiency, latency, and elastic scaling for large-scale LLM workloads.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Research Scientist - AI Compute & DPU - Global Frontier Tech Recruitment Program - 2027 Start (PhD)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 31 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Conduct PhD-level research for an AI infrastructure team that owns the AI Compute layer and DPU, including the open-source Kubernetes-native control plane AIBrix. Work on systematic research across the full stack—network observability, storage systems, data center power scheduling, vector retrieval, and AI-agent-driven intelligence—aiming to improve performance, cost efficiency, latency, and elastic scaling for large-scale LLM workloads.
Location: San Jose
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Graduate level

Key Responsibilities

  • •Perform systematic research across the AI infrastructure stack to address ultra-high performance and elasticity needs of LLM and AI-agent workloads.
  • •Research intelligent fault localization, root cause analysis, and time-series database tuning to improve large-scale cluster stability.
  • •Develop serverless high-performance elastic storage systems and storage-acceleration architectures for AI scenarios, including hardware-software co-optimization for DPU.
  • •Design heterogeneous collaborative GPU/CPU/MEM scheduling and power orchestration systems, addressing scheduling challenges with heterogeneous workloads and state dependencies.
  • •Optimize vector retrieval technologies for low-latency, low-cost ultra-large-scale vector retrieval using a cloud-native distributed vector index engine.

Key Requirements

  • •Completing or recently completed PhD in Software Development, Computer Science, Computer Engineering, or a related technical discipline, focused on distributed and ML systems.
  • •Proven first-author publications in top venues with clear technical contributions.
  • •Strong system-building ability with hands-on experience implementing or optimizing real systems beyond prototypes.
  • •Solid understanding of compute, network architecture, and operating systems.
  • •Deep expertise in at least one area: LLM inference/AI-ML systems, AI system optimization (scheduling, observability, resource management, high-performance networking), or software-hardware co-design.
Education:PhD / Doctorate
Skills:System-buildingResearchOwnershipTechnical communicationRigorous reasoning
Tech Stack:KubernetesAIBrixLLM inferenceKV cacheMultimodal servingKubernetes-native control planeNetworkingRDMADPDKNCCLOVSSR-IOVEBPFFPGAASICVector retrievalVector index engine

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn