Senior Platform and EngOps Engineer - Cluster Operations
Santa Clara
Workplace: OnsiteFull timeUSD 176,000 - 276,000 annuallyFunction: Solutions Engineering & Sales EngineeringExperience: 8+ yearsEducation: bachelorsSkills: ["Ansible","Python","Shell","Linux","NVLink","InfiniBand","Slurm"]Lead the design, deployment, and ongoing operations of large GPU clusters (NVLink/InfiniBand) in a fast-paced AI/ HPC environment. You’ll automate provisioning, maintain software/firmware updates, troubleshoot failures, and coordinate with cross-timezone engineering teams to ensure high-availability cluster performance.
Loading
Loading job details...
Preparing the role view and application actions.

