Senior HPC Engineer – IFM
MBZUAI
Anywhere
Workplace: RemoteFull timeFunction: Solutions Engineering & Sales EngineeringSkills: ["Reliability engineering","Troubleshooting","Root cause analysis","Collaboration","Mentorship"]Provide technical leadership for large-scale GPU infrastructure supporting frontier AI research. Operate and optimize GPU clusters, improving reliability, scalability, and performance through monitoring, capacity planning, and incident response. Lead troubleshooting and root-cause analysis, design and validate cluster deployments and upgrades, and collaborate with researchers to optimize distributed AI training. Engage vendors, define operational standards, and mentor junior engineers.

