Senior HPC AI Cluster Engineer

NVIDIA
Santa Clara
Workplace: HybridFull timeUSD 176,000 - 333,500 annuallyFunction: Data Science & Machine LearningExperience: 8+ yearsSkills: ["Troubleshooting","Documentation","Automation","Collaboration","Systems thinking"]

Build and operate large-scale HPC/AI clusters for GPU-accelerated supercomputing and deep learning workflows. Own monitoring, logging, alerting, and server/network/storage observability, while managing Linux workload scheduling and orchestration. Develop CI/CD pipelines, automation for deployment and infrastructure management, and self-service resource consumption. Troubleshoot end-to-end from bare metal to application layers, and support R&D POCs/POVs for next-generation performance platforms.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
3 days ago

Senior HPC AI Cluster Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 1 hour agoStatus: Live

Job Summary

Build and operate large-scale HPC/AI clusters for GPU-accelerated supercomputing and deep learning workflows. Own monitoring, logging, alerting, and server/network/storage observability, while managing Linux workload scheduling and orchestration. Develop CI/CD pipelines, automation for deployment and infrastructure management, and self-service resource consumption. Troubleshoot end-to-end from bare metal to application layers, and support R&D POCs/POVs for next-generation performance platforms.
Location: Santa Clara
Workplace: Hybrid
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Mid level

Key Responsibilities

  • •Design, implement, and maintain large-scale HPC/AI clusters with monitoring, logging, and alerting.
  • •Manage Linux job/workload schedules and orchestration tools.
  • •Develop and maintain CI/CD pipelines for infrastructure and software delivery.
  • •Create tooling to automate deployment and management of large-scale infrastructure environments, including operational monitoring and self-service resource consumption.
  • •Perform end-to-end troubleshooting from bare metal and OS/software stack through the application level, and support R&D POCs/POVs.

Pay and Benefits

Salary: USD 176,000 - 333,500 annually
Equity and Bonus:Equity

Key Requirements

  • •8+ years of experience and a degree in Computer Science, Engineering, or a related field (or equivalent experience).
  • •Knowledge of HPC and AI solution technologies spanning CPU/GPU systems through high-speed interconnects and supporting software.
  • •Experience with job scheduling and orchestration tools such as Slurm and Kubernetes.
  • •Excellent knowledge of Windows and Linux networking and internals, including firewalld/iptables and protocols such as TCP, DHCP, and DNS.
  • •Experience with Linux tooling and automation (Python, bash, CI/CD) plus configuration management tools like Jenkins and Ansible, and storage solutions such as Lustre, GPFS, and Weka.io.
Experience:8+ yearsHPCAIGPU computingCloud computingDeep learning
Skills:TroubleshootingDocumentationAutomationCollaborationSystems thinking
Tech Stack:LinuxWindowsRedhat/CentOSUbuntuPythonBashSlurmKubernetesJenkinsAnsiblePuppetChefVMwareHyper-VKVMCitrixInfiniBandEthernetInfiniBand or RoCERDMA

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor