Post-Training Platform Infrastructure Engineer

AMD
San Jose
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Problem-solving","Communication","Collaboration"]

Engineer at the intersection of large-scale model inference, distributed systems, and performance optimization, focusing on post-training and inference infrastructure with KV cache lifecycle management and offloading across inference and RL systems. You will analyze bottlenecks, improve latency and throughput, and translate research insights into production-grade features, collaborating with ML teams to benchmark and guide system design.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
AMD
AMD
6 months ago

Post-Training Platform Infrastructure Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 12 hours agoStatus: Live

Job Summary

Engineer at the intersection of large-scale model inference, distributed systems, and performance optimization, focusing on post-training and inference infrastructure with KV cache lifecycle management and offloading across inference and RL systems. You will analyze bottlenecks, improve latency and throughput, and translate research insights into production-grade features, collaborating with ML teams to benchmark and guide system design.
Location: San Jose
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Research and deeply understand modern LLM inference frameworks, including P/D disaggregation and KV cache lifecycle across GPU, CPU, and storage backends.
  • •Analyze inference execution paths to identify performance bottlenecks and scheduling inefficiencies.
  • •Develop infrastructure-level features to improve inference latency, throughput, and memory efficiency; optimize KV cache management and offloading strategies.
  • •Enhance scalability across multi-GPU and multi-node deployments; apply approaches to RL frameworks and post-training pipelines.
  • •Collaborate with research and applied ML teams to translate model requirements into infrastructure capabilities and validate gains with benchmarks.

Key Requirements

  • •Strong background in systems engineering, distributed systems, or ML infrastructure.
  • •Hands-on experience with GPU-accelerated workloads and memory-constrained systems.
  • •Proficiency in Python and C++ (or similar systems languages).
  • •Solid understanding of LLM inference workflows (prefill vs decode) and KV cache behavior.
  • •Experience with distributed RL or post-training pipelines.
Experience:ML infrastructureDistributed systemsLLM inference
Skills:Problem-solvingCommunicationCollaboration
Languages:English
Tech Stack:PythonC++GPUHBMNUMANVMeMemoryKV cacheDistributed systems

Company Brief

AMD
Designs and produces semiconductor products including CPUs, GPUs, and adaptive SoCs for consumer, enterprise, and embedded markets, competing across PCs, data centers, and gaming industries.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1969
Glassdoor
Glassdoor: 3.9
WebsiteLinkedIn