AI Infrastructure Engineer

AMD
San Jose
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsSkills: ["Kubernetes","DevOps","GitOps","Helm","Secret management","Configuration management","Deployment automation","CI pipelines","Self-service workflows","CSI","Networking"]

DevOps/Platform Engineer building and operating large-scale GPU compute infrastructure powering AI/ML workloads. You’ll design and extend a developer-focused platform with Kubernetes orchestration across on-prem and multi-cloud, implement secret/configuration management and deployment automation, and collaborate with development teams to enhance the GPU developer platform while managing service lifecycles with GitOps.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
AMD
AMD
3 months ago

AI Infrastructure Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 16 hours agoStatus: Live
Reposted: similar role first listed 10 months ago

Job Summary

DevOps/Platform Engineer building and operating large-scale GPU compute infrastructure powering AI/ML workloads. You’ll design and extend a developer-focused platform with Kubernetes orchestration across on-prem and multi-cloud, implement secret/configuration management and deployment automation, and collaborate with development teams to enhance the GPU developer platform while managing service lifecycles with GitOps.
Location: San Jose
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Build and extend platform capabilities to enable new classes of workloads (e.g., interactive development pods, CI pipelines, inference services, benchmarking jobs).
  • •Design and operate scalable orchestration systems using Kubernetes across both on-prem and multi-cloud environments.
  • •Develop platform features such as secret management, configuration management, and deployment automation for customers.
  • •Partner with development teams to extend the GPU developer platform with features, APIs, templates, and self-service workflows that streamline job orchestration and environment management.
  • •Manage service lifecycle within Kubernetes using Helm and GitOps workflows (e.g., ArgoCD or Flux).

Key Requirements

  • •5+ years of experience in DevOps, Platform, or Infrastructure Engineering.
  • •Deep hands-on experience with Kubernetes and container orchestration at scale.
  • •Proven ability to design and deliver platform features that serve internal customers or developer teams.
  • •Experience building developer-facing platforms or internal developer portals (e.g. custom workflow tooling).
Experience:5+ yearsAIGPUComputing
Skills:KubernetesDevOpsGitOpsHelmSecret managementConfiguration managementDeployment automationCI pipelinesSelf-service workflowsCSINetworking
Languages:English
Tech Stack:KubernetesHelmGitOpsArgoCDFluxCSIPrometheusGrafanaPyTorch

Company Brief

AMD
Designs and produces semiconductor products including CPUs, GPUs, and adaptive SoCs for consumer, enterprise, and embedded markets, competing across PCs, data centers, and gaming industries.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1969
Glassdoor
Glassdoor: 3.9
WebsiteLinkedIn