Senior HPC Engineer – IFM

MBZUAI
Anywhere
Workplace: RemoteFull timeFunction: Solutions Engineering & Sales EngineeringSkills: ["Reliability engineering","Troubleshooting","Root cause analysis","Collaboration","Mentorship"]

Provide technical leadership for large-scale GPU infrastructure supporting frontier AI research. Operate and optimize GPU clusters, improving reliability, scalability, and performance through monitoring, capacity planning, and incident response. Lead troubleshooting and root-cause analysis, design and validate cluster deployments and upgrades, and collaborate with researchers to optimize distributed AI training. Engage vendors, define operational standards, and mentor junior engineers.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
MBZUAI
MBZUAI
2 months ago

Senior HPC Engineer – IFM

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 10 hours agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Provide technical leadership for large-scale GPU infrastructure supporting frontier AI research. Operate and optimize GPU clusters, improving reliability, scalability, and performance through monitoring, capacity planning, and incident response. Lead troubleshooting and root-cause analysis, design and validate cluster deployments and upgrades, and collaborate with researchers to optimize distributed AI training. Engage vendors, define operational standards, and mentor junior engineers.
Location: Anywhere
Workplace: Remote
Employment Type: Full time
Job Function: Solutions Engineering & Sales Engineering
Seniority: Mid level

Key Responsibilities

  • •Lead operation and optimization of large-scale GPU clusters to improve reliability, scalability, and performance.
  • •Drive troubleshooting and root cause analysis for complex compute, storage, and networking issues.
  • •Design, validate, and roll out new cluster deployments and upgrades.
  • •Collaborate with researchers to optimize distributed AI training at scale.
  • •Lead vendor engagement and technical reviews, define monitoring/operational standards and capacity planning, and manage major incidents.

Key Requirements

  • •5+ years in HPC, Linux infrastructure, cloud infrastructure, distributed systems, or large-scale production environments.
  • •Experience with Slurm and Linux administration.
  • •Strong troubleshooting experience for compute, storage, and networking systems.
  • •Bachelor’s degree in computer science, computer engineering, electrical engineering, software engineering, information technology, applied mathematics, physics, or a related field.
  • •Knowledge of GPU cluster operations and NVIDIA technologies such as CUDA, NCCL, NVLink, and GPUDirect.
Skills:Reliability engineeringTroubleshootingRoot cause analysisCollaborationMentorship
Tech Stack:HPCLinuxSlurmGPU clustersCUDANCCLNVLinkGPUDirectInfiniBandWekaLustreBeeGFSAzureAWSGCPTerraformAnsibleInfrastructure-as-CodePyTorch DistributedMegatron-LM

Company Brief

MBZUAI
Mohamed bin Zayed University of Artificial Intelligence is a graduate-level research university in Abu Dhabi focused on AI education, research, and industry collaboration. It offers advanced programs in machine learning, computer vision, natural language processing, and related fields.
Industry: Higher Education
Company Size: Large (251 to 1,000 employees)
Growth: Government & Public Sector
Funding: Government Funded
Headquarters: Abu Dhabi, United Arab Emirates
Founded: 2019
WebsiteLinkedIn