Network Engineer (Supercomputer Infrastructure)

X AI
Tennessee, Memphis, Mississippi
Workplace: OnsiteFull timeFunction: IT Operations (Systems/Network Admin)Experience: 3+ yearsEducation: bachelorsSkills: ["Communication","Initiative","Operational excellence","Prioritization","Cross-functional collaboration"]

Design, build, and operate the high-availability, low-latency networks powering AI supercomputer campuses. Own network fabrics for GPU training and inference clusters, storage, and OT/site networks—balancing routing, congestion control, and redundancy. Evaluate and deploy data-center hardware, mature network automation with GitOps/IaC, and provide 24x7 on-call support. Troubleshoot cluster-impacting incidents, publish RCA, and maintain documentation and monitoring/telemetry.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
X AI
X AI
2 days ago

Network Engineer (Supercomputer Infrastructure)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 6 hours agoStatus: Live

Job Summary

Design, build, and operate the high-availability, low-latency networks powering AI supercomputer campuses. Own network fabrics for GPU training and inference clusters, storage, and OT/site networks—balancing routing, congestion control, and redundancy. Evaluate and deploy data-center hardware, mature network automation with GitOps/IaC, and provide 24x7 on-call support. Troubleshoot cluster-impacting incidents, publish RCA, and maintain documentation and monitoring/telemetry.
Location: Tennessee, Memphis, Mississippi
Workplace: Onsite
Employment Type: Full time
Job Function: IT Operations (Systems/Network Admin)
Seniority: Mid level

Key Responsibilities

  • •Design and implement highly available, low-latency, high-bandwidth networks for AI training fabrics, inference front-ends, storage, and site/OT networks.
  • •Design and maintain supercomputer data center and campus networks, coordinating with adjacent infrastructure, compute, storage, SiteOps, and enterprise teams.
  • •Evaluate, procure, and deploy network hardware supporting 400G/800G and beyond (switches, NICs, firewalls, optical multiplexers, and related appliances).
  • •Mature network automation tooling with configuration analysis, linting, validation, and scalable deployment frameworks using GitOps/IaC.
  • •Troubleshoot network issues impacting cluster health and job performance, document RCA, host retrospectives, and provide direct on-call networking support during operations and cluster bring-up/expansions.
Travel: Medium travel

Key Requirements

  • •Bachelor’s degree in CS/computer engineering or other STEM, with 3+ years of professional network engineering experience; or 5+ years of professional network engineering experience in lieu of a degree.
  • •Extensive hands-on experience designing, deploying, supporting, and troubleshooting Layer 2/Layer 3 networks in latency-sensitive and/or data-center/industrial environments.
  • •Functional experience with multiple network vendors in production or lab environments.
  • •Experience with GitOps and Infrastructure as Code frameworks as a user and contributor.
  • •Ability to pass applicable background checks for site access and provide 24x7 on-call support in emergency situations.
Experience:3+ yearsData centerIndustrialAerospaceDefenseHigh-reliability environmentsOT networksAI/HPC
Education:Bachelor's in computer science, computer engineering, or other STEM discipline
Skills:CommunicationInitiativeOperational excellencePrioritizationCross-functional collaboration
Licenses:Valid driver’s license
Certifications:CCNACCNP
Tech Stack:Layer 2Layer 3GitOpsInfrastructure as CodeGitOps/IaCNCCLNVIDIA Spectrum-XCiscoAristaJuniperRoCEv2InfiniBandWDMOTDRECMPAdaptive routingQoSMulticastTerraformAnsible

Company Brief

X AI
Develops advanced artificial intelligence models and research aimed at building safe, general AI and understanding the fundamental nature of the universe. Focuses on large-scale AI systems, research publications, and building foundational AI capabilities.
Industry: AI & Machine Learning
Company Size: Medium (51 to 250 employees)
Growth: Early Stage Startup
Headquarters: San Francisco, United States
Founded: 2023
Website