Senior Network Engineer – GPU Cluster Networking
San Jose
Workplace: HybridFull timeUSD 151,900 - 260,400 annuallyFunction: IT Operations (Systems/Network Admin)Education: mastersSkills: ["Incident response","Root-cause analysis","Capacity planning","Continuous improvement","Documentation"]Own the architecture, deployment, optimization, automation, and production operations of high-performance backend networks for large-scale AMD GPU clusters. Design and scale Ethernet and RoCEv2 fabrics (100/200/400 GbE), model bandwidth and oversubscription, and tune lossless/near-lossless settings for predictable low-latency, reliable collective communication. Partner across AI, storage, security, and platform teams, driving incident response, capacity planning, and observability using tools like Prometheus and Grafana.
Loading
Loading job details...
Preparing the role view and application actions.

