AI Cluster Architect

Vultr
United States
Workplace: RemoteFull timeUSD 165,000 - 185,000Function: Hardware, Embedded & Systems EngineeringExperience: 7+ yearsSkills: ["Documentation","Communication","Cross-functional collaboration","Power-aware design","Tradeoff analysis"]

Design and refine large-scale GPU cluster architectures within fixed power and facility limits. You’ll determine optimal GPU counts using power-aware modeling across the full bill of materials—GPUs, CPUs, NICs, switches, fabrics, storage, and cooling—while evaluating networking tradeoffs across InfiniBand, RoCE, and SpectrumX. Build scalable capacity-planning templates, document design choices, and collaborate with vendors on fabric innovations for 100k+ GPU deployments.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Vultr
Vultr
5 days ago

AI Cluster Architect

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live
Reposted: similar role first listed 6 months ago

Job Summary

Design and refine large-scale GPU cluster architectures within fixed power and facility limits. You’ll determine optimal GPU counts using power-aware modeling across the full bill of materials—GPUs, CPUs, NICs, switches, fabrics, storage, and cooling—while evaluating networking tradeoffs across InfiniBand, RoCE, and SpectrumX. Build scalable capacity-planning templates, document design choices, and collaborate with vendors on fabric innovations for 100k+ GPU deployments.
Location: United States
Workplace: Remote
Employment Type: Full time
Job Function: Hardware, Embedded & Systems Engineering
Seniority: Mid level

Key Responsibilities

  • •Architect large-scale GPU clusters within fixed site power budgets to maximize GPU density while reserving necessary headroom.
  • •Model and validate cluster power consumption across GPUs, CPUs, NICs, switches, fabric components, storage, and facility limits.
  • •Evaluate networking architecture tradeoffs across InfiniBand, RoCE, and SpectrumX, including multi-plane and multi-tier topologies.
  • •Determine network scale limits using switch radix, link speed, topology, and blocking requirements.
  • •Document architecture and design tradeoffs and guide future-proofing for next-gen GPUs, NICs, and fabrics, collaborating with vendors for novel architectures.

Pay and Benefits

Salary: USD 165,000 - 185,000
Perks:Health InsuranceDentalVision401kLearning BudgetPaid LeaveRemote OfficeGym Membership

Key Requirements

  • •7+ years designing or building large-scale HPC, AI, or hyperscale GPU clusters.
  • •Expert understanding of GPU/accelerator system design, including node topology and NIC-to-GPU affinity considerations.
  • •Strong familiarity with InfiniBand, RoCE, and SpectrumX networking, including multi-tier, multi-plane, and large-radix switch design.
  • •Demonstrated ability to model power draw and thermal characteristics across GPUs, servers, NICs, switches, optics, and storage.
  • •Ability to gather and analyze vendor SKU-level specifications and incorporate them into scalable cluster architectures.
Experience:7+ yearsHPCAIHyperscaleGPU clustersCloud infrastructure
Skills:DocumentationCommunicationCross-functional collaborationPower-aware designTradeoff analysis
Tech Stack:InfiniBandRoCESpectrumXPCIeNVLinkNVSwitchROCmClosDragonflyMulti-planeRail-optimized

Company Brief

Vultr
Provides cloud infrastructure services including VPS, dedicated instances, block storage, and bare metal across global data centers. Targets developers and businesses with simple, high-performance, and cost-effective cloud compute and networking solutions.
Industry: Cloud Computing
Company Size: Large (251 to 1,000 employees)
Growth: Established Company
Headquarters: West Palm Beach, United States
Founded: 2014
WebsiteLinkedIn