Software Development Engineer — GPU Fleet Management & AI Infrastructure

AMD
San Jose
Workplace: HybridFull timeFunction: Transportation & Fleet OperationsSkills: ["Technical judgment","Clear communication","Collaboration","Project leadership","Operability focus"]

Build Fleet Manager, a secure control plane for operating large-scale AMD GPU infrastructure. Design and implement distributed control-plane services, Kubernetes integrations, GPU scheduling, inference and developer-facing APIs/tools. Work across systems software, cloud infrastructure, and accelerated computing to improve orchestration, reliability, security, observability, and performance of AI training and inference across current and future AMD GPU platforms.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
AMD
AMD
2 days ago

Software Development Engineer — GPU Fleet Management & AI Infrastructure

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 2 hours agoStatus: Live

Job Summary

Build Fleet Manager, a secure control plane for operating large-scale AMD GPU infrastructure. Design and implement distributed control-plane services, Kubernetes integrations, GPU scheduling, inference and developer-facing APIs/tools. Work across systems software, cloud infrastructure, and accelerated computing to improve orchestration, reliability, security, observability, and performance of AI training and inference across current and future AMD GPU platforms.
Location: San Jose
Workplace: Hybrid
Employment Type: Full time
Job Function: Transportation & Fleet Operations
Seniority: Mid level

Key Responsibilities

  • •Design and develop Fleet Manager distributed control-plane services, APIs, schedulers, inference gateway, and command-line tools.
  • •Build reliable orchestration for GPU training, inference, custom jobs, and interactive development workloads.
  • •Develop scalable scheduling and admission-control capabilities, including priority, fairness, quotas, backfilling, and multi-node coordination.
  • •Implement durable reconciliation, lifecycle management, retries, idempotency, and recovery across PostgreSQL and external execution systems.
  • •Integrate Fleet Manager with Kubernetes and enable GPU health/diagnostics/quarantine and secure multi-tenant operations with auditing and least-privilege defaults.

Key Requirements

  • •Strong systems-software development experience in Rust, C++, Go, or a comparable language, with production Rust experience highly desirable.
  • •Experience designing and operating distributed systems/control planes, schedulers, or cloud infrastructure.
  • •Strong understanding of concurrency, asynchronous programming, state machines, and failure recovery.
  • •Experience building reliable services using REST, streaming, WebSocket, or gRPC APIs.
  • •Experience with Kubernetes internals/schedulers/operators and PostgreSQL-backed services, including schema evolution and transactional reliability.
Experience:Distributed systemsCloud infrastructureKubernetesAI infrastructureGPU infrastructure
Education:
Skills:Technical judgmentClear communicationCollaborationProject leadershipOperability focus
Languages:English
Tech Stack:RustC++GoKubernetesKueueJobSetSlurmPostgreSQLRESTWebSocketGRPCROCmHIPAmd-smiRCCLVLLMSGLangPyTorchContainer runtimesMetrics

Eligibility

Work Authorization:Authorization required. Sponsorship not provided.

Company Brief

AMD
Designs and produces semiconductor products including CPUs, GPUs, and adaptive SoCs for consumer, enterprise, and embedded markets, competing across PCs, data centers, and gaming industries.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1969
Glassdoor
Glassdoor: 3.9
WebsiteLinkedIn