HPC Solutions Engineer

Hydra Host
United States
Workplace: RemoteFull timeFunction: Solutions Engineering & Sales EngineeringExperience: 7+ yearsSkills: ["Problem-solving","Communication","Teamwork","Ownership","Self-motivation","Organization"]

Own bespoke HPC solutions for premium clients by scoping, deploying, and operating large-scale GPU clusters. Work with customers to translate technical needs into requirements, select tools, and create reusable reference implementations. Manage NVIDIA GPU environments and high-speed networking (InfiniBand), oversee SLURM-based HPC scheduling and distributed ML/HPC workflows, and automate infrastructure with Ansible and Terraform. Collaborate with data scientists and engineers on production integration, tuning, and documentation.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Hydra Host
Hydra Host
2 months ago

HPC Solutions Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 8 hours agoStatus: Live

Job Summary

Own bespoke HPC solutions for premium clients by scoping, deploying, and operating large-scale GPU clusters. Work with customers to translate technical needs into requirements, select tools, and create reusable reference implementations. Manage NVIDIA GPU environments and high-speed networking (InfiniBand), oversee SLURM-based HPC scheduling and distributed ML/HPC workflows, and automate infrastructure with Ansible and Terraform. Collaborate with data scientists and engineers on production integration, tuning, and documentation.
Location: United States
Workplace: Remote
Employment Type: Full time
Job Function: Solutions Engineering & Sales Engineering
Seniority: Mid level

Key Responsibilities

  • •Partner with customers on technical discovery to define requirements and deliverables for distributed GPU computing use cases.
  • •Identify and recommend tools for each customer, building reusable boilerplate and reference implementations.
  • •Manage GPU clusters and coordinate with IT to ensure efficient operation of NVIDIA GPUs using InfiniBand.
  • •Oversee deployment and maintenance of ML environments using virtual storage solutions and distributed HPC tools like SLURM.
  • •Automate provisioning and manage performance tuning and optimization to maximize throughput and reduce latency.

Pay and Benefits

Equity and Bonus:Equity
Perks:EquityPaid Leave

Key Requirements

  • •7+ years of experience in high performance compute, distributed machine learning, GPU computing, and/or system architecture.
  • •Proficiency managing NVIDIA GPU environments and familiarity with GPU computing frameworks and libraries.
  • •Strong experience with high-speed networking, specifically InfiniBand.
  • •Experience with HPC job schedulers, preferably SLURM.
  • •Expertise automating environment setup and maintenance using Ansible and Terraform, with strong Python coding skills.
Experience:7+ years
Skills:Problem-solvingCommunicationTeamworkOwnershipSelf-motivationOrganization
Tech Stack:PythonNVIDIA GPUsInfiniBandSLURMAnsibleTerraformHPCDistributed computingMachine learning librariesVirtual storage

Company Brief

Hydra Host
Hydra Host provides bare metal GPU infrastructure and an operating system for AI factories. It helps customers and data center operators deploy, manage, and monetize GPU compute across distributed data centers, with a focus on AI workloads and sovereign infrastructure.
Industry: Cloud Computing
Company Size: Small (11 to 50 employees)
Growth: Growth Stage Startup
Funding: Series A
Headquarters: Miami, United States
WebsiteLinkedIn