Site Reliability Engineer

RunPod
United States
Workplace: RemoteFull timeUSD 150,000 - 200,000 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsSkills: ["Ownership","Proactive problem solving","Automation-first thinking","Cross-functional collaboration","Excellent written communication"]

Own reliability for a distributed AI infrastructure platform by defining and operationalizing SLIs/SLOs, improving incident response, and driving production readiness. Build observability and reliability tooling (monitoring, alerting, dashboards) using Prometheus/Grafana, and reduce operational toil through automation and safe CI/CD practices. Partner cross-functionally to strengthen resilience, improve MTTR, and prevent incidents before they impact developers running critical workloads.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
RunPod
RunPod
1 month ago

Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 8 hours agoStatus: Live
Reposted: similar role first listed 2 years ago

Job Summary

Own reliability for a distributed AI infrastructure platform by defining and operationalizing SLIs/SLOs, improving incident response, and driving production readiness. Build observability and reliability tooling (monitoring, alerting, dashboards) using Prometheus/Grafana, and reduce operational toil through automation and safe CI/CD practices. Partner cross-functionally to strengthen resilience, improve MTTR, and prevent incidents before they impact developers running critical workloads.
Location: United States
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Define and implement SLIs/SLOs for critical services and enforce reliability standards
  • •Lead incident response, coordinate cross-team mitigation, and run blameless postmortems to ensure corrective actions
  • •Design and improve monitoring, alerting, and dashboards to reduce alert fatigue and improve reliability signal
  • •Automate recurring operational workflows and build tools/scripts to reduce manual toil and improve deployment safety
  • •Drive long-term reliability improvements via production readiness reviews and cross-functional reliability advocacy

Pay and Benefits

Salary: USD 150,000 - 200,000 annually
Equity and Bonus:Equity
Perks:Health InsuranceDentalVisionEquityRemote WorkPaid Leave

Key Requirements

  • •5+ years of experience in SRE, Reliability Engineering, or Production Engineering
  • •Strong Linux systems and networking expertise
  • •Experience managing containerized production systems
  • •Experience defining and managing SLIs/SLOs and understanding distributed systems failure modes
  • •Proven incident response/postmortem leadership and strong scripting or programming skills
Experience:5+ yearsAI infrastructureDistributed systemsProduction operations
Skills:OwnershipProactive problem solvingAutomation-first thinkingCross-functional collaborationExcellent written communication
Tech Stack:LinuxNetworkingContainerized systemsSLIsSLOsPrometheusGrafanaGPU performanceDistributed systemsPythonGoBashCI/CDSlackInfrastructure as CodeObservability tooling

Eligibility

Work Authorization:Authorization required. Sponsorship not provided.

Company Brief

RunPod
Provides on-demand GPU cloud and marketplace services for machine learning workloads, offering rentable GPU instances, scalable compute for training and inference, and tools to run ML jobs cost-effectively without long-term commitments.
Industry: Cloud Computing
Website