Summer 2027 PhD ML Systems Research Engineering Intern

AMD
Santa Clara
Workplace: HybridInternshipUSD 91,520 - 137,280 annuallyFunction: Research & Scientific (R&D)Education: phdSkills: ["Analytical skills","Problem-solving","Communication"]

Develop and optimize infrastructure for distributed training, reinforcement learning, inference, and model optimization. Build scalable systems for experiment management, rollout generation, model serving, and evaluation, improving efficiency through parallelism, scheduling, checkpointing, caching, and performance tuning. Create tools integrating AI models with engineering systems, evaluation pipelines, and benchmarking frameworks, and support experiment tracking, logging, monitoring, and reproducibility. Analyze system performance and document designs and results.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
AMD
AMD
1 day ago

Summer 2027 PhD ML Systems Research Engineering Intern

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 20 hours agoStatus: Live

Job Summary

Develop and optimize infrastructure for distributed training, reinforcement learning, inference, and model optimization. Build scalable systems for experiment management, rollout generation, model serving, and evaluation, improving efficiency through parallelism, scheduling, checkpointing, caching, and performance tuning. Create tools integrating AI models with engineering systems, evaluation pipelines, and benchmarking frameworks, and support experiment tracking, logging, monitoring, and reproducibility. Analyze system performance and document designs and results.
Location: Santa Clara
Workplace: Hybrid
Employment Type: Internship
Job Function: Research & Scientific (R&D)
Seniority: Intern level

Key Responsibilities

  • •Develop and optimize infrastructure for distributed training, reinforcement learning, inference, and model optimization.
  • •Build scalable systems for experiment management, rollout generation, model serving, and evaluation.
  • •Improve training and inference efficiency via parallelism, scheduling, checkpointing, caching, and performance optimization.
  • •Develop tools integrating AI models with engineering systems, evaluation pipelines, and benchmarking frameworks.
  • •Support experiment tracking, logging, monitoring, and reproducibility, and analyze system performance for reliability, scalability, and resource utilization improvements.

Pay and Benefits

Salary: USD 91,520 - 137,280 annually

Key Requirements

  • •Must be currently enrolled in a PhD program in Computer Science, Computer Engineering, Electrical Engineering, Artificial Intelligence, or a related field.
  • •Programming experience in Python.
  • •Coursework or hands-on experience in artificial intelligence or computer systems.
  • •Familiarity with PyTorch or similar machine learning frameworks.
  • •Interest in distributed systems, reinforcement learning, LLMs, or AI infrastructure.
Experience:Machine learningDistributed systemsAI infrastructure
Education:PhD / Doctorate
Skills:Analytical skillsProblem-solvingCommunication
Languages:English
Tech Stack:PythonPyTorch

Eligibility

Work Authorization:Authorization required. Sponsorship not provided.

Company Brief

AMD
Designs and produces semiconductor products including CPUs, GPUs, and adaptive SoCs for consumer, enterprise, and embedded markets, competing across PCs, data centers, and gaming industries.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1969
Glassdoor
Glassdoor: 3.9
WebsiteLinkedIn