Sr. AI/ML Platform Engineer

AMD
Santa Clara
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureEducation: bachelorsSkills: ["Reproducibility","Observability","Reliability","Performance","Developer experience"]

Build and operate a shared AI platform for agentic engineering workflows, enabling scalable job submission, scheduling, orchestration, retries, logging, artifact storage, and experiment tracking. Develop infrastructure for distributed training and inference across GPU clusters, automate benchmarking and evaluation, and improve reproducibility, observability, performance, and developer experience. Partner with ML research and applied AI engineers to productionize workflows and make Blueprint harnesses reusable across hardware/software tooling.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
AMD
AMD
2 months ago

Sr. AI/ML Platform Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 6 hours agoStatus: Live

Job Summary

Build and operate a shared AI platform for agentic engineering workflows, enabling scalable job submission, scheduling, orchestration, retries, logging, artifact storage, and experiment tracking. Develop infrastructure for distributed training and inference across GPU clusters, automate benchmarking and evaluation, and improve reproducibility, observability, performance, and developer experience. Partner with ML research and applied AI engineers to productionize workflows and make Blueprint harnesses reusable across hardware/software tooling.
Location: Santa Clara
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Build and operate the shared AI platform for agentic engineering workflows, including job submission, scheduling, orchestration, retries, logging, artifact storage, and experiment tracking.
  • •Develop reliable infrastructure for distributed training, distributed inference, batch evaluation, and large-scale agent rollout across GPU clusters.
  • •Build platform services for benchmark execution, correctness checking, profiling, regression tracking, and reproducible evaluation.
  • •Maintain artifact systems for kernels, RTL edits, traces, logs, profiler outputs, benchmark results, simulator outputs, and formal verification artifacts.
  • •Support integrations with ROCm/HIP tooling, profilers, simulators, EDA tools, vLLM, SGLang, and internal engineering systems, improving GPU cluster utilization and observability.

Key Requirements

  • •Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, Machine Learning, or related field, or equivalent practical experience.
  • •Strong programming skills in Python and one or more systems languages such as C++, Go, or Rust.
  • •Experience building ML platforms, AI infrastructure, distributed systems, workflow orchestration, experiment platforms, GPU cluster infrastructure, or developer platforms.
  • •Strong understanding of job scheduling, distributed workloads, logging, monitoring, reliability, storage systems, and production operations.
  • •Experience with Kubernetes, Ray, Slurm, workflow engines, containerization, CI/CD, data pipelines, or large-scale compute orchestration.
Experience:AI infrastructureDistributed systemsGPU cluster infrastructure
Education:Bachelor's in Computer Science, Computer Engineering, Electrical Engineering, Machine Learning, or related field
Skills:ReproducibilityObservabilityReliabilityPerformanceDeveloper experience
Languages:English
Tech Stack:PythonC++GoRustKubernetesRaySlurmCI/CDContainersPyTorchJAXTritonMLflowWeights & BiasesCUDAROCm/HIPVLLMSGLangCompilersProfilers

Company Brief

AMD
Designs and produces semiconductor products including CPUs, GPUs, and adaptive SoCs for consumer, enterprise, and embedded markets, competing across PCs, data centers, and gaming industries.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1969
Glassdoor
Glassdoor: 3.9
WebsiteLinkedIn