Staff Site Reliability Engineer - Observability

Okta
San Francisco
Workplace: HybridFull timeUSD 194,000 - 267,000 annuallyFunction: DevOps, Cloud & InfrastructureSkills: ["Problem-solving","Data-driven","On-call","Collaboration"]

Senior, highly technical SRE focused on observability and Google Cloud. You will build and scale an observability platform in GCP using Terraform and languages like Python/Go/Ruby, maintain IaC pipelines, and drive incident response and observability-driven development across distributed systems.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Okta
Okta
4 months ago

Staff Site Reliability Engineer - Observability

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 7 hours agoStatus: Live

Job Summary

Senior, highly technical SRE focused on observability and Google Cloud. You will build and scale an observability platform in GCP using Terraform and languages like Python/Go/Ruby, maintain IaC pipelines, and drive incident response and observability-driven development across distributed systems.
Location: San Francisco
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Automated Infrastructure: Design, build, and maintain scalable observability infrastructure using tools like Terraform.
  • •GCP Observability Engineering: Optimize collection, processing, and storage of Observability data to ensure high reliability and low latency of Splunk and Grafana services.
  • •Incident Response: Participate in on-call rotations and lead post-incident reviews to drive systemic improvements and observability-driven development.
  • •Automation: Eliminate toil by automating the deployment and scaling of observability agents and collectors.

Pay and Benefits

Salary: USD 194,000 - 267,000 annually
Perks:Health InsuranceDentalVision401kPaid Leave

Key Requirements

  • •GKE: Minimum 5+ years scaling and managing observability in Google Cloud Platform
  • •Visualization: Expertise in creating Splunk or Grafana dashboards that correlate data across multiple sources
  • •SRE Mindset: Minimum 3+ years in an SRE, DevOps, or Systems Engineering role focused on high-availability systems
  • •Programming Proficiency: Strong Python and Go (also Ruby) for building internal tools and automating workflows
  • •Distributed Systems: Linux internals, networking (TCP/IP, DNS, Load Balancing), and Kubernetes/GKE
Experience:CloudObservabilitySREGKEKubernetes
Skills:Problem-solvingData-drivenOn-callCollaboration
Languages:English
Tech Stack:TerraformPythonGoRubyLinuxKubernetesGKEGrafanaSplunkOpenTelemetryOTel

Eligibility

Nationality:US National
Work Authorization:Authorization required. Sponsorship not provided.

Company Brief

Okta
Provides identity and access management cloud solutions that help organizations secure and manage user authentication, single sign-on, multi-factor authentication, and lifecycle management across applications and devices.
Industry: Cybersecurity
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: San Francisco, United States
Founded: 2009
WebsiteLinkedIn