Principal Site Reliability Engineer
Santa Clara
Workplace: HybridFull timeUSD 248,000 - 396,750 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 15+ yearsEducation: bachelorsSkills: ["Technical judgment","Communication","Collaboration","Mentoring","Influencing stakeholders"]Shape the technical vision for reliability across NVIDIA’s AI Platform Runtime, leading architecture and roadmap efforts across multiple organizations. Design and deliver highly available, resilient, secure distributed platforms, and build AI-driven automation to accelerate incident response, troubleshooting, and remediation. Set enterprise reliability standards (SLOs, error budgets, capacity models) and advance observability using OpenTelemetry and anomaly detection. Provide technical leadership during critical incidents and mentor engineering leaders to raise the technical bar.
Loading
Loading job details...
Preparing the role view and application actions.

