Senior HPC Cluster Engineer

NVIDIA
Santa Clara, Austin, Redmond
Workplace: OnsiteFull timeUSD 152,000 - 287,500 annuallyFunction: Solutions Engineering & Sales EngineeringExperience: 5+ yearsEducation: bachelorsSkills: ["Communication","Collaboration","Problem-solving","Teamwork"]

Design, deploy, and operate GPU compute clusters for EDA and high-performance computing workloads, leading large-scale HPC infrastructure, automating provisioning, and collaborating with researchers and infrastructure teams to optimize performance and reliability.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
4 months ago

Senior HPC Cluster Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

Design, deploy, and operate GPU compute clusters for EDA and high-performance computing workloads, leading large-scale HPC infrastructure, automating provisioning, and collaborating with researchers and infrastructure teams to optimize performance and reliability.
Location: Santa Clara, Austin, Redmond
Workplace: Onsite
Employment Type: Full time
Job Function: Solutions Engineering & Sales Engineering

Key Responsibilities

  • •Design, deploy, and operate GPU compute clusters for EDA and HPC workloads across multiple teams.
  • •Lead automation efforts to improve infrastructure provisioning, management, observability, and day-to-day operations.
  • •Provide technical leadership for managing large-scale HPC systems, including compute, networking, and storage deployment.
  • •Foster partnerships with researchers and cross-functional teams to ensure responsive cluster support and alignment with user needs.
  • •Support researchers with performance analysis and optimizations for EDA workloads.

Pay and Benefits

Salary: USD 152,000 - 287,500 annually
Equity and Bonus:Equity
Perks:Equity

Key Requirements

  • •Bachelor’s degree in Computer Science, Electrical Engineering or related field or equivalent experience.
  • •Minimum of 5 years of proven experience crafting and operating large scale compute infrastructure, including cluster configuration management tools such as BCM or Ansible.
  • •Experience with AI/HPC job schedulers and orchestrators, such as Slurm, LSF, PBS or K8s; MPI and NCCL workflows.
  • •Proficient in using Linux distributions (Rocky/Centos/RHEL and/or Ubuntu); understanding of container technologies such as Enroot and Docker.
  • •Proficiency in Python and Bash
  • •Note: only top 5 were selected; if more required, please refer to full description.
Experience:5+ yearsHPCGPUAIEDA
Education:Bachelor's
Skills:CommunicationCollaborationProblem-solvingTeamwork
Tech Stack:LinuxAnsibleSlurmLSFPBSK8sMPINCCLEnrootDockerPythonBashRockyCentOSRHELUbuntuPrometheusGrafana

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor