Senior HPC AI Cluster Engineer

NVIDIA
Switzerland, Zurich
Workplace: RemoteFull timeFunction: Data Science & Machine LearningExperience: 8+ yearsEducation: bachelorsSkills: ["Troubleshooting","Documentation"]

Design, implement, and run large-scale HPC/AI clusters, ensuring reliable monitoring, logging, alerting, and performance. Own Linux workload scheduling and orchestration, build CI/CD pipelines, and develop automation for deploying and managing infrastructure environments. Deploy monitoring across servers, network, and storage, troubleshoot from bare metal through the software stack, and document standard methodologies while supporting R&D and future improvements.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
1 hour ago

Senior HPC AI Cluster Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 1 hour agoStatus: Live

Job Summary

Design, implement, and run large-scale HPC/AI clusters, ensuring reliable monitoring, logging, alerting, and performance. Own Linux workload scheduling and orchestration, build CI/CD pipelines, and develop automation for deploying and managing infrastructure environments. Deploy monitoring across servers, network, and storage, troubleshoot from bare metal through the software stack, and document standard methodologies while supporting R&D and future improvements.
Location: Switzerland, Zurich
Workplace: Remote
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Mid level

Key Responsibilities

  • •Design, implement, and maintain large-scale HPC/AI clusters with monitoring, logging, and alerting.
  • •Manage Linux job/workload schedules and orchestration tools.
  • •Develop and maintain continuous integration and delivery pipelines.
  • •Create tooling to automate deployment and management of large-scale infrastructure environments, including operational monitoring and self-service resource consumption.
  • •Troubleshoot end-to-end from bare metal and OS/software stack to application level; develop and document standard methodologies and support R&D POCs/POVs.

Key Requirements

  • •8+ years of experience with a degree in Computer Science, Engineering, or related field.
  • •Knowledge of HPC and AI solution technologies across CPU/GPU and high-speed interconnects and supporting software.
  • •Experience with job scheduling and orchestration tools such as Slurm and Kubernetes.
  • •Strong Linux and Windows networking knowledge (e.g., Redhat/CentOS and Ubuntu), including protocols like TCP/DHCP/DNS and tools such as firewalld/iptables/wireshark.
  • •Python and bash scripting experience, plus automation/configuration management with tools like Jenkins, Ansible, and Puppet/Chef.
Experience:8+ yearsHPCAIGPU computingCloud computing
Education:Bachelor's
Skills:TroubleshootingDocumentation
Tech Stack:HPCAIGPU computingCPU architectureSlurmKubernetesK8sLinuxRedhatCentOSUbuntuWindowsFirewalldIptablesWiresharkTCPDHCPDNSLustreGPFS

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor