AI Infrastructure Engineer Intern (Algorithm Infrastructure) - 2027 Start (PhD)

TikTok
San Jose
Workplace: OnsiteInternshipFunction: DevOps, Cloud & InfrastructureEducation: phdSkills: []

Build and evolve AI inference infrastructure for ultra-large-scale language and multimodal models. Work on global traffic orchestration, throughput and latency optimization, kernel efficiency, and production reliability, including distributed inference strategies (TP/EP/DP). Develop high-performance kernels for architectures like MoE and multimodal fusion layers using CUDA and Triton, and explore AI-driven infrastructure for optimization and deployment.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
TikTok
TikTok
1 hour ago

AI Infrastructure Engineer Intern (Algorithm Infrastructure) - 2027 Start (PhD)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 1 hour agoStatus: Live

Job Summary

Build and evolve AI inference infrastructure for ultra-large-scale language and multimodal models. Work on global traffic orchestration, throughput and latency optimization, kernel efficiency, and production reliability, including distributed inference strategies (TP/EP/DP). Develop high-performance kernels for architectures like MoE and multimodal fusion layers using CUDA and Triton, and explore AI-driven infrastructure for optimization and deployment.
Location: San Jose
Workplace: Onsite
Employment Type: Internship
Job Function: DevOps, Cloud & Infrastructure
Seniority: Intern level

Key Responsibilities

  • •Build and evolve next-generation inference systems for large-scale online traffic, including global scheduling across heterogeneous compute, high-concurrency load balancing, and efficient batch formation.
  • •Optimize distributed inference for 200B+ models and complex multimodal models using TP/EP/DP and related strategies to improve throughput and latency in production.
  • •Develop high-performance kernels for frontier model architectures such as MoE, emerging attention mechanisms, and multimodal fusion layers using CUDA, Triton, and related tools.
  • •Explore AI-driven infrastructure for inference systems, including AI Agents for kernel optimization, performance tuning, consistency validation, deployment pipelines, and intelligent operations.

Key Requirements

  • •Currently pursuing a PhD in Software Development, Computer Science, Computer Engineering, or a related technical field.
  • •Familiarity with large-model architectures and strong system design skills for complex, high-concurrency environments.
  • •Strong engineering skills in performance optimization and production system development.
  • •Strong understanding of asynchronous scheduling, resource pooling, and load balancing in distributed microservice systems.
Experience:AI infrastructureLarge modelsDistributed systemsMultimodal AIMicroservices
Education:PhD / Doctorate in Software Development, Computer Science, Computer Engineering, or related technical discipline
Tech Stack:CUDATritonTPEPDPMoEMicroserviceDistributed servingLow-latency inference

Company Brief

TikTok
Short-form video platform that lets users create, share, and discover entertainment content through algorithmic recommendations. It also offers advertising and creator tools for brands, influencers, and businesses.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Headquarters: Singapore, Singapore
Founded: 2016
WebsiteLinkedIn