Member of Technical Staff - Datacenter Networking

Prime Intellect
San Francisco
Workplace: HybridFull timeUSD 150,000 - 300,000 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 3+ yearsSkills: ["Reliability","Performance troubleshooting","Incident response","Benchmarking","Automation"]

Design and operate datacenter networks that connect large GPU clusters for training, inference, storage, and management traffic. Own reliability and performance across Ethernet/RoCE and InfiniBand fabrics, automate provisioning and upgrades, and troubleshoot issues like packet loss, congestion, and link failures. Build monitoring and incident runbooks, benchmark network performance with infrastructure and ML teams, and partner with operators and hardware vendors to ensure successful deployments.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Prime Intellect
Prime Intellect
1 day ago

Member of Technical Staff - Datacenter Networking

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 15 minutes agoStatus: Live

Job Summary

Design and operate datacenter networks that connect large GPU clusters for training, inference, storage, and management traffic. Own reliability and performance across Ethernet/RoCE and InfiniBand fabrics, automate provisioning and upgrades, and troubleshoot issues like packet loss, congestion, and link failures. Build monitoring and incident runbooks, benchmark network performance with infrastructure and ML teams, and partner with operators and hardware vendors to ensure successful deployments.
Location: San Francisco
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Design and deploy scalable datacenter network topologies for GPU training, inference, storage, and management traffic.
  • •Configure and operate high-performance Ethernet/RoCE and InfiniBand fabrics with clear routing, redundancy, and capacity standards.
  • •Automate network provisioning, configuration validation, upgrades, and rollback procedures.
  • •Diagnose packet loss, congestion, link failures, and collective communication performance across hosts and switches.
  • •Build monitoring and improve incident response and runbooks based on network health and performance signals.

Pay and Benefits

Salary: USD 150,000 - 300,000 annually

Key Requirements

  • •3+ years of production datacenter networking experience.
  • •Strong understanding of Ethernet/TCP-IP, routing, switching, and redundant network design.
  • •Hands-on experience with high-performance GPU networking using InfiniBand or RoCE.
  • •Troubleshoot network problems across Linux hosts, NICs, switches, and physical links.
  • •Automate network operations with Python, Ansible, or comparable tools.
Experience:3+ yearsDatacenter networkingGPU networkingDistributed training
Skills:ReliabilityPerformance troubleshootingIncident responseBenchmarkingAutomation
Tech Stack:EthernetTCP/IPBGPECMPVLANsNetwork segmentationInfiniBandRoCEPythonAnsibleLinuxPacket captureRDMANCCLEVPN/VXLANVXLANSONiCTelemetryAlertingOptics

Company Brief

Prime Intellect
Builds a decentralized, open compute and training platform that enables distributed training and collective ownership of AI models, aggregating global GPU resources and offering tools for agentic RL and model evaluation.
Industry: AI & Machine Learning
Company Size: Small (11 to 50 employees)
Growth: Early Stage Startup
Funding: Seed
Headquarters: San Francisco, United States
Founded: 2024
WebsiteLinkedIn