GPU Fabric Engineer
Vultr
United States
Workplace: RemoteFull timeUSD 125,000 - 135,000Function: DevOps, Cloud & InfrastructureExperience: 3-7 yearsSkills: ["Troubleshooting","Analytical skills","Collaboration","Documentation"]Validate, troubleshoot, and optimize high-speed InfiniBand/RoCE networking fabrics for GPU clusters used in large-scale AI training and inference. Tune interconnect parameters for distributed workloads, monitor fabric health with NVIDIA UFM, diagnose issues like congestion and packet loss, and optimize RDMA, PFC/ECN, and lossless queue settings. Collaborate with GPU and networking teams, respond to production alerts, and document runbooks and tuning procedures.

