AI Infrastructure and Frameworks Intern, Cosmos Lab - 2027

NVIDIA
Beijing, Shanghai, Shenzhen
Workplace: OnsiteInternshipFunction: Laboratory & Clinical OperationsSkills: ["Analytical skills","Communication","Curiosity","Willingness to learn"]

Join NVIDIA’s Cosmos Lab Infrastructure team to build training and post-training systems for advanced Physical AI models, including world models and robot policies. You’ll work on distributed training and RL workflows that connect simulation and real-robot experience collection to inference, reward computation, and evaluation. The role emphasizes systems improvements, performance analysis, profiling/benchmarking, and collaboration with researchers using NVIDIA GPU infrastructure.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
15 hours ago

AI Infrastructure and Frameworks Intern, Cosmos Lab - 2027

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 15 hours agoStatus: Live

Job Summary

Join NVIDIA’s Cosmos Lab Infrastructure team to build training and post-training systems for advanced Physical AI models, including world models and robot policies. You’ll work on distributed training and RL workflows that connect simulation and real-robot experience collection to inference, reward computation, and evaluation. The role emphasizes systems improvements, performance analysis, profiling/benchmarking, and collaboration with researchers using NVIDIA GPU infrastructure.
Location: Beijing, Shanghai, Shenzhen
Workplace: Onsite
Employment Type: Internship
Job Function: Laboratory & Clinical Operations
Seniority: Intern level

Key Responsibilities

  • •Develop and optimize training infrastructure for advanced Physical AI world models, including pre-training, SFT, and RL using distributed and low-precision techniques.
  • •Build post-training and RL infrastructure that connects simulation or real-robot interaction to experience collection, rollout inference, reward computation, training, and evaluation.
  • •Improve efficiency and scalability across training, inference, simulation, and evaluation via scheduling, placement, dynamic resource allocation, and load balancing with fault recovery.
  • •Analyze and optimize system performance by profiling, benchmarking, and performance modeling to identify bottlenecks and measure throughput, latency, GPU utilization, and policy freshness.
  • •Share results through tested code, documentation, technical presentations, and contribute to research publications where appropriate.

Key Requirements

  • •Pursuing a Bachelor’s, Master’s, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or a related field.
  • •Strong Python and debugging skills with systems fundamentals in concurrency, distributed execution, memory management, or data movement.
  • •Practical experience in at least one area such as training infrastructure, RL infrastructure, simulation or robotics integration, or inference infrastructure.
  • •Strong analytical and communication skills, curiosity, and a willingness to learn.
Education:
Skills:Analytical skillsCommunicationCuriosityWillingness to learn
Tech Stack:PythonC++CUDAGPUDistributed parallelismShardingLow-precision trainingProfilingBenchmarking

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor