Senior Network Engineer – GPU Cluster Networking

AMD
San Jose
Workplace: HybridFull timeUSD 151,900 - 260,400 annuallyFunction: IT Operations (Systems/Network Admin)Education: mastersSkills: ["Incident response","Root-cause analysis","Capacity planning","Continuous improvement","Documentation"]

Own the architecture, deployment, optimization, automation, and production operations of high-performance backend networks for large-scale AMD GPU clusters. Design and scale Ethernet and RoCEv2 fabrics (100/200/400 GbE), model bandwidth and oversubscription, and tune lossless/near-lossless settings for predictable low-latency, reliable collective communication. Partner across AI, storage, security, and platform teams, driving incident response, capacity planning, and observability using tools like Prometheus and Grafana.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
AMD
AMD
1 day ago

Senior Network Engineer – GPU Cluster Networking

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 3 hours agoStatus: Live

Job Summary

Own the architecture, deployment, optimization, automation, and production operations of high-performance backend networks for large-scale AMD GPU clusters. Design and scale Ethernet and RoCEv2 fabrics (100/200/400 GbE), model bandwidth and oversubscription, and tune lossless/near-lossless settings for predictable low-latency, reliable collective communication. Partner across AI, storage, security, and platform teams, driving incident response, capacity planning, and observability using tools like Prometheus and Grafana.
Location: San Jose
Workplace: Hybrid
Employment Type: Full time
Job Function: IT Operations (Systems/Network Admin)
Seniority: Mid level

Key Responsibilities

  • •Architect, deploy, operate, and continuously improve high-performance backend networks for large-scale AMD Instinct GPU clusters.
  • •Design network fabrics for AI/HPC environments ranging from individual GPU racks to clusters with ~10,000+ GPUs.
  • •Own the backend network architecture from GPU server and NIC through the leaf-spine switching fabric.
  • •Optimize Ethernet/RoCEv2 networking performance and troubleshoot lossless/near-lossless environments.
  • •Lead production incident response, root-cause analysis, corrective actions, and preventive engineering improvements.

Pay and Benefits

Salary: USD 151,900 - 260,400 annually

Key Requirements

  • •Architect, deploy, operate, and continuously improve high-performance backend networks for large-scale AMD Instinct GPU clusters.
  • •Design scalable network fabrics for AI/HPC environments from GPU racks to clusters with ~10,000+ GPUs.
  • •Configure, tune, validate, and troubleshoot RoCEv2 lossless/near-lossless networking (PFC, ECN, DCQCN, QoS, ECMP, DSCP mappings).
  • •Design and operate routing/switching using technologies including BGP, VLAN, VRF, EVPN, and VXLAN.
  • •Provide end-to-end performance optimization across GPUs, NICs, switches, CPUs, PCIe devices, storage systems, and Linux networking stack.
Experience:AIGPU clustersHPCDistributed computingCloud
Education:Master's in Computer Engineering
Skills:Incident responseRoot-cause analysisCapacity planningContinuous improvementDocumentation
Languages:English
Tech Stack:EthernetRoCEv2RDMA100/200/400 GbEPCIeNUMAROCmRCCLSLURMKubernetesBGPECMPVLANVRFEVPNVXLANPFCECNDCQCNQoS

Eligibility

Work Authorization:Authorization required. Sponsorship not provided.

Company Brief

AMD
Designs and produces semiconductor products including CPUs, GPUs, and adaptive SoCs for consumer, enterprise, and embedded markets, competing across PCs, data centers, and gaming industries.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1969
Glassdoor
Glassdoor: 3.9
WebsiteLinkedIn