Network Engineer, Supercomputing

Thinking Machines Lab
San Francisco
Workplace: HybridFull timeUSD 350,000 - 475,000 annuallyFunction: IT Operations (Systems/Network Admin)Education: bachelorsSkills: ["Collaboration","Communication","Initiative","Ownership","Reliability reasoning"]

Own the low-level network stack that powers large-scale AI training and inference. You’ll ensure interconnect reliability across GPU fabrics, debugging RDMA/RoCEv2 and collective failures (including NCCL and congestion control tuning). Build host-level instrumentation with Linux tooling for dashboards and alerts, triage issues across NIC/driver/kernel/switch/workload boundaries, and drive escalations with cloud networking teams to resolution.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Thinking Machines Lab
Thinking Machines Lab
1 month ago

Network Engineer, Supercomputing

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 5 hours agoStatus: Live

Job Summary

Own the low-level network stack that powers large-scale AI training and inference. You’ll ensure interconnect reliability across GPU fabrics, debugging RDMA/RoCEv2 and collective failures (including NCCL and congestion control tuning). Build host-level instrumentation with Linux tooling for dashboards and alerts, triage issues across NIC/driver/kernel/switch/workload boundaries, and drive escalations with cloud networking teams to resolution.
Location: San Francisco
Workplace: Hybrid
Employment Type: Full time
Job Function: IT Operations (Systems/Network Admin)

Key Responsibilities

  • •Validate GPU network fabric design across deployments.
  • •Debug RDMA/RoCEv2 across NIC vendors and diagnose NCCL collective failures, PFC/ECN tuning, and congestion control behavior.
  • •Own NVLink/NVSwitch interconnect including fabric manager and IMEX health, and understand how the GPU fabric interacts with collectives.
  • •Build host-level network instrumentation using Linux tooling to create dashboards and alerts.
  • •Triage cross-cloud fabric issues across NIC, driver, kernel, switch, and workload boundaries, and drive escalations to resolution with networking teams.

Pay and Benefits

Salary: USD 350,000 - 475,000 annually
Perks:Health InsuranceDentalVisionPaid LeaveParental LeaveRelocation

Key Requirements

  • •Bachelor’s degree or equivalent experience in computer science, engineering, or similar.
  • •Proficiency in at least one backend language (Python or Rust).
  • •Experience operating large-scale clusters and container orchestration systems such as Kubernetes or Slurm.
  • •Comfort operating across the stack and owning projects end-to-end.
  • •Thrive in collaborative, cross-functional environments and take initiative to ship.
Education:Bachelor's
Skills:CollaborationCommunicationInitiativeOwnershipReliability reasoning
Tech Stack:PythonRustKubernetesSlurmLinuxRDMARoCEv2NCCLPFC/ECNCongestion controlNVLinkNVSwitchFabric managerIMEXKernelsCUDADeep learning frameworks

Eligibility

Work Authorization:Sponsorship available.

Company Brief

Thinking Machines Lab
Develops enterprise AI solutions, custom large language models, and ML platforms to help organizations deploy intelligent applications. Services include data engineering, model development, and AI consulting for scale and production readiness.
Industry: AI & Machine Learning
Company Size: Medium (51 to 250 employees)
Growth: Growth Stage Startup
Headquarters: Mumbai, India
Website