Site Reliability Engineer
RunPod
United States
Workplace: RemoteFull timeUSD 150,000 - 200,000 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsSkills: ["Ownership","Proactive problem solving","Automation-first thinking","Cross-functional collaboration","Excellent written communication"]Own reliability for a distributed AI infrastructure platform by defining and operationalizing SLIs/SLOs, improving incident response, and driving production readiness. Build observability and reliability tooling (monitoring, alerting, dashboards) using Prometheus/Grafana, and reduce operational toil through automation and safe CI/CD practices. Partner cross-functionally to strengthen resilience, improve MTTR, and prevent incidents before they impact developers running critical workloads.

