Software Development Engineer — GPU Fleet Management & AI Infrastructure

AMD
San Jose
Workplace: HybridFull timeUSD 150,500 - 258,000 annuallyFunction: Transportation & Fleet OperationsEducation: bachelorsSkills: ["Technical judgment","Clear communication","Collaboration","Project leadership","Focus on correctness and security"]

Build Fleet Manager, a secure control plane for operating large-scale AMD GPU infrastructure. Design and develop distributed control-plane services, APIs, schedulers, an inference gateway, and developer tools. You’ll orchestrate GPU training and inference workloads, integrate with Kubernetes using scheduling technologies like Kueue and JobSet, and implement reliable lifecycle management backed by PostgreSQL. Work on health/diagnostics and multi-tenant security while improving AI inference performance and operability.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
AMD
AMD
2 days ago

Software Development Engineer — GPU Fleet Management & AI Infrastructure

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 14 hours agoStatus: Live

Job Summary

Build Fleet Manager, a secure control plane for operating large-scale AMD GPU infrastructure. Design and develop distributed control-plane services, APIs, schedulers, an inference gateway, and developer tools. You’ll orchestrate GPU training and inference workloads, integrate with Kubernetes using scheduling technologies like Kueue and JobSet, and implement reliable lifecycle management backed by PostgreSQL. Work on health/diagnostics and multi-tenant security while improving AI inference performance and operability.
Location: San Jose
Workplace: Hybrid
Employment Type: Full time
Job Function: Transportation & Fleet Operations

Key Responsibilities

  • •Design and develop Fleet Manager’s distributed control-plane services, APIs, schedulers, inference gateway, and command-line tools.
  • •Build reliable orchestration for GPU training, inference, custom jobs, and interactive development workloads.
  • •Implement scalable scheduling and admission control (priority/fairness/topology-aware placement/quotas/backfilling) with multi-node coordination.
  • •Integrate Fleet Manager with Kubernetes and related technologies (e.g., Kueue, JobSet, container runtimes, storage, observability).
  • •Improve reliability and performance of AI inference services (routing, streaming, load shedding, health detection, and usage metering).

Pay and Benefits

Salary: USD 150,500 - 258,000 annually

Key Requirements

  • •Experience designing and developing distributed systems, control planes, schedulers, or cloud infrastructure.
  • •Hands-on systems-software development experience in Rust, C++, Go, or a comparable language; production Rust experience is highly desirable.
  • •Strong understanding of concurrency, asynchronous programming, state machines, and failure recovery.
  • •Experience building reliable services with REST, streaming, WebSocket, or gRPC APIs.
  • •Experience with Kubernetes internals (controllers/operators/scheduling/resource management/custom resources) and scheduling technologies such as Kueue or JobSet or Slurm.
Experience:Distributed systemsCloud infrastructureKubernetesGPU infrastructureAI inference
Education:Bachelor's in Computer Science, Computer Engineering, Electrical Engineering, or a related field
Skills:Technical judgmentClear communicationCollaborationProject leadershipFocus on correctness and security
Languages:En-us
Tech Stack:KubernetesPostgreSQLRustC++GoRESTWebSocketGRPCKueueJobSetSlurmROCmHIPAmd-smiRCCLVLLMSGLangPyTorchOpenAI-compatible inference APIsContainer runtimes

Eligibility

Work Authorization:Authorization required. Sponsorship not provided.

Company Brief

AMD
Designs and produces semiconductor products including CPUs, GPUs, and adaptive SoCs for consumer, enterprise, and embedded markets, competing across PCs, data centers, and gaming industries.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1969
Glassdoor
Glassdoor: 3.9
WebsiteLinkedIn