Senior Network Engineer – GPU Cluster Networking

AMD
Hyderabad
Workplace: OnsiteFull timeINR 3,103,450 - 4,433,500 annuallyFunction: IT Operations (Systems/Network Admin)Education: mastersSkills: ["Hands-on execution","Technical leadership","Mentoring","Documentation","Data-driven decision-making"]

Own the end-to-end backend network for large-scale AMD GPU clusters, ensuring predictable bandwidth, low latency, and reliable collective communication. Architect, deploy, optimize, and operate high-performance high-speed Ethernet and RoCEv2 fabrics, including routing, switching, lossless tuning, capacity planning, and incident response. Partner with AI, platform, storage, security, and application teams while validating designs with telemetry and repeatable performance testing.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
AMD
AMD
8 hours ago

Senior Network Engineer – GPU Cluster Networking

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 49 minutes agoStatus: Live

Job Summary

Own the end-to-end backend network for large-scale AMD GPU clusters, ensuring predictable bandwidth, low latency, and reliable collective communication. Architect, deploy, optimize, and operate high-performance high-speed Ethernet and RoCEv2 fabrics, including routing, switching, lossless tuning, capacity planning, and incident response. Partner with AI, platform, storage, security, and application teams while validating designs with telemetry and repeatable performance testing.
Location: Hyderabad
Workplace: Onsite
Employment Type: Full time
Job Function: IT Operations (Systems/Network Admin)
Seniority: Mid level

Key Responsibilities

  • •Architect, deploy, operate, and continuously improve high-performance backend networks for large-scale AMD Instinct GPU clusters.
  • •Design network fabrics supporting AI and HPC environments from GPU racks to clusters with 10,000+ GPUs.
  • •Own the backend network architecture from GPU server and NIC through the leaf-spine switching fabric.
  • •Configure, tune, validate, and troubleshoot RoCEv2 environments, including PFC/ECN/DCQCN and QoS/buffering/queue management.
  • •Lead production incident response, root-cause analysis, corrective actions, and preventive engineering improvements for GPU cluster networks.

Pay and Benefits

Salary: INR 3,103,450 - 4,433,500 annually

Key Requirements

  • •Bachelor’s or Master’s degree in Computer Engineering or a related field, or equivalent practical experience.
  • •Significant experience designing, deploying, and operating production data center networks for AI, GPU, HPC, cloud, or similarly large distributed computing environments.
  • •Strong hands-on expertise in RDMA and RoCEv2 in production environments, including lossless/near-lossless tuning.
  • •Deep knowledge of data center networking fundamentals including routing/switching, VLAN/subnetting, BGP/ECMP, QoS, MTU configuration, and switch buffering.
  • •Experience with GPU-cluster topology and locality (GPU-to-GPU, GPU-to-NIC/CPU-to-NIC, PCIe, NUMA), plus monitoring/observability tooling.
Experience:AIGPUHPCCloudLarge-scale distributed computing
Education:Master's in Computer Engineering
Skills:Hands-on executionTechnical leadershipMentoringDocumentationData-driven decision-making
Languages:En-us
Tech Stack:RoCEv2Ethernet100/200/400 GbERDMAPFCECNDCQCNQoSECMPBGPVLANVRFEVPNVXLANDSCPLinux networkingPCIeNUMAOpticsSwitch buffer

Company Brief

AMD
Designs and produces semiconductor products including CPUs, GPUs, and adaptive SoCs for consumer, enterprise, and embedded markets, competing across PCs, data centers, and gaming industries.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1969
Glassdoor
Glassdoor: 3.9
WebsiteLinkedIn