AI Systems Engineer – AI Model (Training & Inference)

AMD
Markham
Workplace: HybridFull timeFunction: IT Operations (Systems/Network Admin)Skills: ["Debugging","Optimization","Observability","Collaboration","Reproducibility"]

Own the end-to-end model execution stack on AMD Instinct GPUs, spanning large-scale training infrastructure and high-performance inference serving. Enable and optimize LLM/VLM/MoE training on GPU clusters using Kubernetes, distributed checkpointing, and robust validation. Deliver inference acceleration by optimizing GPU kernels and serving frameworks for throughput/latency, quantization pipelines, and observability for large-scale reliability.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
AMD
AMD
21 hours ago

AI Systems Engineer – AI Model (Training & Inference)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 7 hours agoStatus: Live
Reposted: similar role first listed 4 months ago

Job Summary

Own the end-to-end model execution stack on AMD Instinct GPUs, spanning large-scale training infrastructure and high-performance inference serving. Enable and optimize LLM/VLM/MoE training on GPU clusters using Kubernetes, distributed checkpointing, and robust validation. Deliver inference acceleration by optimizing GPU kernels and serving frameworks for throughput/latency, quantization pipelines, and observability for large-scale reliability.
Location: Markham
Workplace: Hybrid
Employment Type: Full time
Job Function: IT Operations (Systems/Network Admin)
Seniority: Mid level

Key Responsibilities

  • •Enable and optimize large-scale LLM/VLM/MoE training on AMD Instinct GPU clusters, ensuring correctness, reproducibility, and throughput.
  • •Build and maintain training infrastructure including job orchestration, distributed checkpointing, data loading pipelines, and storage optimization on Kubernetes.
  • •Debug and resolve training issues such as gradient instability, non-determinism across GPU generations, and training communication/compute overlap.
  • •Optimize inference enablement and serving frameworks on AMD GPUs, including batching, KV-cache management, speculative decoding, and continuous batching.
  • •Collaborate with silicon and pre-silicon teams and build observability/validation tooling to keep training and inference platforms production-ready across Instinct generations.

Key Requirements

  • •Shipped LLMs on real hardware and built systems for training and inference across large-scale GPU setups.
  • •Hands-on experience writing GPU kernels (e.g., GEMM/attention/quantized matmul) targeting AMD Instinct architectures using HIP, Triton, and/or MLIR.
  • •Experience enabling large-scale distributed training using frameworks such as FSDP, DeepSpeed, and Megatron-LM, plus optimization of RCCL communication patterns.
  • •Background building observability and automated analysis tooling for large-scale distributed GPU clusters (log analysis, anomaly detection, performance baselining, regression detection).
  • •Bachelor’s, Master’s, or Ph.D. in Computer/Software Engineering, Computer Science, or a related technical discipline.
Experience:AI/ML infrastructureDistributed systemsHPCGPU computing
Education:
Skills:DebuggingOptimizationObservabilityCollaborationReproducibility
Tech Stack:LLMsVLMsMoEKubernetesDistributed checkpointingHIPTritonMLIRRCCLFSDPDeepSpeedMegatron-LMAll-reduceAll-gatherReduce-scatterVLLMSGLangTorchServeKV-cacheSpeculative decoding

Eligibility

Work Authorization:Authorization required. Sponsorship not provided.

Company Brief

AMD
Designs and produces semiconductor products including CPUs, GPUs, and adaptive SoCs for consumer, enterprise, and embedded markets, competing across PCs, data centers, and gaming industries.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1969
Glassdoor
Glassdoor: 3.9
WebsiteLinkedIn