HPC Infrastructure & Cluster Engineer

General Dynamics
Virginia
Workplace: OnsiteFull timeUSD 119,850 - 162,150 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsSkills: ["Troubleshooting","Security/compliance focus"]

Manage the administration, health, and performance of a dedicated customer compute cluster supporting UDS. Own day-to-day Linux and bare-metal operations, hardware monitoring, patching, upgrades, and workload orchestration using Run:AI and SLURM. Optimize compute, networking, and storage performance (SAN, InfiniBand GPU-to-GPU). Configure environments with Red Hat OpenShift and ensure compliance with federal security standards and access controls for secure, high-availability operations.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
General Dynamics
General Dynamics
2 days ago

HPC Infrastructure & Cluster Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 47 minutes agoStatus: Live

Job Summary

Manage the administration, health, and performance of a dedicated customer compute cluster supporting UDS. Own day-to-day Linux and bare-metal operations, hardware monitoring, patching, upgrades, and workload orchestration using Run:AI and SLURM. Optimize compute, networking, and storage performance (SAN, InfiniBand GPU-to-GPU). Configure environments with Red Hat OpenShift and ensure compliance with federal security standards and access controls for secure, high-availability operations.
Location: Virginia
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Administer the customer compute cluster, including Linux administration, hardware monitoring, patching, and system upgrades.
  • •Configure and optimize workload management and orchestration platforms to efficiently distribute intensive AI/ML workloads across the cluster.
  • •Tune cluster performance across hardware, OS, and network layers to maximize compute efficiency and data throughput.
  • •Administer storage and high-speed networking fabrics, including ongoing management of an InfiniBand GPU-to-GPU network infrastructure.
  • •Collaborate with integration teams to provision environments and container platforms, leveraging Red Hat OpenShift for model deployment.

Pay and Benefits

Salary: USD 119,850 - 162,150 annually
Perks:401kHealth InsuranceDentalVisionPaid LeaveLearning Budget

Key Requirements

  • •Active TS/SCI clearance with ability to obtain CI Poly.
  • •5+ years managing Linux systems administration and infrastructure in high-performance computing environments.
  • •Expertise administering bare-metal servers, enterprise storage arrays, and advanced network configurations (InfiniBand).
  • •Strong knowledge of workload managers/job schedulers and AI orchestration tools (Run:AI, SLURM).
  • •Hands-on experience with enterprise container orchestration (OpenShift or Kubernetes) and automation scripting (Bash, Python).
Experience:5+ yearsHigh-performance computingAI/MLLinux systems administration
Skills:TroubleshootingSecurity/compliance focus
Tech Stack:LinuxRun:AISLURMInfiniBandSANRed Hat OpenShiftKubernetesOpenShiftBashPython

Eligibility

Work Authorization:Authorization required. Sponsorship not provided.
Security Clearance:Top Secret/SCITop Secret SCI + Polygraph

Company Brief

General Dynamics
Provides IT, systems engineering, cybersecurity, cloud, and mission support services to U.S. federal civilian and defense agencies, delivering large-scale technology solutions and managed services for national security and government operations.
Industry: Defense Technology
Company Size: Enterprise (1,001+ employees)
Growth: Established Company
Headquarters: Tysons / Fairfax, United States
WebsiteLinkedIn