Senior HPC Engineer – IFM
Anywhere
Workplace: RemoteFull timeFunction: Solutions Engineering & Sales EngineeringSkills: ["Reliability engineering","Troubleshooting","Root cause analysis","Collaboration","Mentorship"]Provide technical leadership for large-scale GPU infrastructure supporting frontier AI research. Operate and optimize GPU clusters, improving reliability, scalability, and performance through monitoring, capacity planning, and incident response. Lead troubleshooting and root-cause analysis, design and validate cluster deployments and upgrades, and collaborate with researchers to optimize distributed AI training. Engage vendors, define operational standards, and mentor junior engineers.
Loading
Loading job details...
Preparing the role view and application actions.

