HPC Engineer – IFM

MBZUAI
Anywhere
Workplace: RemoteFull timeFunction: DevOps, Cloud & InfrastructureEducation: bachelorsSkills: []

Join the IFM infrastructure team to support and maintain large-scale GPU computing clusters powering frontier AI research. You’ll assist researchers with job submission and troubleshooting, monitor cluster health and performance, and handle issues across Linux, hardware, storage, networking, and software. The role includes Slurm administration, cluster deployment and upgrades, building automation scripts, documenting operations, and collaborating with researchers and vendors.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
MBZUAI
MBZUAI
2 months ago

HPC Engineer – IFM

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 10 hours agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Join the IFM infrastructure team to support and maintain large-scale GPU computing clusters powering frontier AI research. You’ll assist researchers with job submission and troubleshooting, monitor cluster health and performance, and handle issues across Linux, hardware, storage, networking, and software. The role includes Slurm administration, cluster deployment and upgrades, building automation scripts, documenting operations, and collaborating with researchers and vendors.
Location: Anywhere
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Entry level

Key Responsibilities

  • •Support operation and maintenance of large-scale GPU computing clusters.
  • •Assist researchers with job submission, troubleshooting, and resource utilization.
  • •Monitor cluster health, performance, and availability.
  • •Troubleshoot issues across Linux, hardware, storage, networking, and software.
  • •Support Slurm administration, cluster deployment/upgrades/validation, and develop automation scripts and maintain technical documentation.

Key Requirements

  • •Bachelor’s degree in Computer Science, Computer Engineering, Electrical Engineering, Software Engineering, Information Technology, Mathematics, Physics, or related disciplines.
  • •Linux administration experience.
  • •Programming in Python, Bash, Go, or C/C++.
  • •Networking fundamentals and cloud platforms exposure (Azure, AWS, GCP).
  • •Experience with containers (Docker, Apptainer, Enroot) and Git-based software development workflows.
Experience:HPCDistributed systemsResearch computingAI/ML infrastructureCloudContainers
Education:Bachelor's
Tech Stack:LinuxPythonBashGoC/C++AzureAWSGCPDockerApptainerEnrootGitSlurm

Company Brief

MBZUAI
Mohamed bin Zayed University of Artificial Intelligence is a graduate-level research university in Abu Dhabi focused on AI education, research, and industry collaboration. It offers advanced programs in machine learning, computer vision, natural language processing, and related fields.
Industry: Higher Education
Company Size: Large (251 to 1,000 employees)
Growth: Government & Public Sector
Funding: Government Funded
Headquarters: Abu Dhabi, United Arab Emirates
Founded: 2019
WebsiteLinkedIn