Senior HPC Cluster Engineer - AI, ML
Bengaluru, Pune
Workplace: HybridFull timeFunction: Data Science & Machine LearningExperience: 5+ yearsSkills: ["Leadership","Incident response","Collaboration","Analytical mindset","Continual learning"]Lead systems administration and service delivery for production AI/HPC GPU clusters on the MARS team. Own day-to-day cluster operations, coordinating upgrades, incident response, and reliability improvements while ensuring system health and efficient resource utilization. Collaborate with global engineers to enhance researcher user experience, build automation and a scalable GPU-accelerated computing ecosystem, and support performance analysis and optimization using MPI-based AI/HPC workloads.
Loading
Loading job details...
Preparing the role view and application actions.

