Senior HPC AI Cluster Engineer
Santa Clara
Workplace: HybridFull timeUSD 176,000 - 333,500 annuallyFunction: Data Science & Machine LearningExperience: 8+ yearsSkills: ["Troubleshooting","Documentation","Automation","Collaboration","Systems thinking"]Build and operate large-scale HPC/AI clusters for GPU-accelerated supercomputing and deep learning workflows. Own monitoring, logging, alerting, and server/network/storage observability, while managing Linux workload scheduling and orchestration. Develop CI/CD pipelines, automation for deployment and infrastructure management, and self-service resource consumption. Troubleshoot end-to-end from bare metal to application layers, and support R&D POCs/POVs for next-generation performance platforms.
Loading
Loading job details...
Preparing the role view and application actions.

