Sr. Staff Site Reliability Engineer

Obsidian Security
Palo Alto
Workplace: OnsiteFull timeUSD 232,000 - 263,000 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsSkills: ["Leadership","Problem-solving","Communication","Collaboration","Teamwork"]

Define and drive a company-wide reliability vision for a complex, multi-tenant SaaS platform serving enterprise and financial customers. Partner with DevOps and Platform Engineering to shape a unified reliability strategy, build end-to-end visibility, and lead initiatives that detect, diagnose, and communicate issues before customers are impacted, while hands-on contributing to system design and observability.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Obsidian Security
Obsidian Security
3 months ago

Sr. Staff Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 8 hours agoStatus: Live

Job Summary

Define and drive a company-wide reliability vision for a complex, multi-tenant SaaS platform serving enterprise and financial customers. Partner with DevOps and Platform Engineering to shape a unified reliability strategy, build end-to-end visibility, and lead initiatives that detect, diagnose, and communicate issues before customers are impacted, while hands-on contributing to system design and observability.
Location: Palo Alto
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Sr. Manager level

Key Responsibilities

  • •Reliability Strategy & Architecture - Define and lead long-term reliability strategy across services. Establish end-to-end system visibility frameworks and guide architecture for observability, detection, and resilience.
  • •Cross-Org Leadership - Partner across teams to embed reliability, standardize SLI/SLOs, and serve as a technical escalation expert.
  • •Detection & Observability - Build intelligent detection systems (anomaly detection, connector health models) and enable self-service observability.
  • •Incident Management - Define and evolve a tiered incident communication strategy, improve response practices, and lead postmortems to strengthen reliability and customer trust.
  • •Execution - Contribute hands-on to system design, monitoring, and debugging across distributed systems and data pipelines.

Pay and Benefits

Salary: USD 232,000 - 263,000 annually
Equity and Bonus:Equity
Perks:Health InsuranceDental401kEquity

Key Requirements

  • •5+ years in SRE, Production Engineering, or related roles.
  • •3+ years operating at a senior or technical leadership level (Staff or equivalent scope).
  • •Deep expertise in AWS and/or GCP; Kubernetes and Helm; Observability stacks (Prometheus, Grafana, or equivalent); and CI/CD systems (GitLab CI/CD, ArgoCD, etc.).
  • •Proven experience designing and scaling reliability systems for multi-tenant SaaS platforms.
  • •Strong debugging and systems thinking across distributed microservices and legacy systems.
Experience:5+ yearsSaaSEnterpriseSecurityCloudObservability
Skills:LeadershipProblem-solvingCommunicationCollaborationTeamwork
Languages:English
Tech Stack:AWSGCPKubernetesHelmPrometheusGrafanaGitLab CI/CDArgoCD

Company Brief

Obsidian Security
Provides cloud-native security solutions that detect and respond to identity- and configuration-based threats across SaaS, IaaS, and cloud identity platforms. Focuses on discovering risky exposures, prioritizing alerts, and enabling remediation for enterprise cloud environments.
Industry: Cybersecurity
Headquarters: San Mateo, United States
WebsiteLinkedIn