Senior HPC Cluster Engineer - AI, ML

NVIDIA
Bengaluru, Pune
Workplace: HybridFull timeFunction: Data Science & Machine LearningExperience: 5+ yearsSkills: ["Leadership","Incident response","Collaboration","Analytical mindset","Continual learning"]

Lead systems administration and service delivery for production AI/HPC GPU clusters on the MARS team. Own day-to-day cluster operations, coordinating upgrades, incident response, and reliability improvements while ensuring system health and efficient resource utilization. Collaborate with global engineers to enhance researcher user experience, build automation and a scalable GPU-accelerated computing ecosystem, and support performance analysis and optimization using MPI-based AI/HPC workloads.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
16 hours ago

Senior HPC Cluster Engineer - AI, ML

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 16 hours agoStatus: Live

Job Summary

Lead systems administration and service delivery for production AI/HPC GPU clusters on the MARS team. Own day-to-day cluster operations, coordinating upgrades, incident response, and reliability improvements while ensuring system health and efficient resource utilization. Collaborate with global engineers to enhance researcher user experience, build automation and a scalable GPU-accelerated computing ecosystem, and support performance analysis and optimization using MPI-based AI/HPC workloads.
Location: Bengaluru, Pune
Workplace: Hybrid
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Mid level

Key Responsibilities

  • •Provide leadership for systems administration and service delivery on the AI/HPC fleet by coordinating upgrades, responding to incidents, and improving reliability.
  • •Own day-to-day operations of production AI/HPC clusters, ensuring system health, user satisfaction, and efficient resource utilization.
  • •Develop and improve a GPU-accelerated computing ecosystem by building scalable automation solutions.
  • •Build and maintain heterogeneous AI/ML clusters on-premises and in the cloud.
  • •Support researchers with workload performance analysis and optimizations, including root cause analysis, SEV triage, postmortems, and participating in on-call incident response.

Key Requirements

  • •Bachelor’s degree in Computer Science, Electrical Engineering (or related) or equivalent experience.
  • •Minimum 5 years of experience designing and operating large-scale compute infrastructure.
  • •Experience with AI/HPC advanced job schedulers such as Slurm, K8s, PBS, RTDA, BCM, or LSF.
  • •Proficiency administering CentOS/RHEL and/or Ubuntu Linux distributions.
  • •Strong configuration management and container ecosystem experience (e.g., Terraform/Ansible/Puppet/Salt; Docker/Singularity/Podman/Shifter/Charliecloud), plus Python and bash; applied MPI experience.
Experience:5+ yearsAI/HPCInfrastructureDistributed clusters
Education:
Skills:LeadershipIncident responseCollaborationAnalytical mindsetContinual learning
Tech Stack:SlurmKubernetesPBSRTDABCMLSFCentOSRHELUbuntuTerraformAnsiblePuppetSaltDockerSingularityPodmanShifterCharliecloudPythonBash

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor