Site Reliability Engineer (High Performance Computing)
United States
Workplace: OnsiteFull timeUSD 125,000 - 195,000 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 2+ yearsEducation: bachelorsSkills: ["Ownership","Self-critical","Fairness","Clear communication","Incident readiness"]Own the end-to-end operating model for a shared HPC compute platform, building reliability through Linux administration, infrastructure-as-code, storage/resource management, and production automation. Drive observability across clusters, nodes, and storage while reducing toil via software tooling. Partner with HPC systems engineers to improve cluster reliability, incidents, and capacity planning so engineers can run mission-critical work with minimal friction.
Loading
Loading job details...
Preparing the role view and application actions.

