Staff Site Reliability Engineer

Okta
Bengaluru
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Technical leadership","Incident response","Collaboration","Communication","Continuous improvement"]

Join Okta’s Emerging Products Group as a technical leader in Site Reliability Engineering. You’ll design and operate large-scale cloud infrastructure, run on-call, lead incident response and post-incident improvements, and define reliability metrics like SLIs/SLOs and error budgets. Build automation and operational guardrails using Go, Python, Terraform, CI/CD, and GitOps. Drive observability improvements with Datadog and Splunk while mentoring engineers and influencing architecture.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Okta
Okta
2 weeks ago

Staff Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live
Reposted: similar role first listed 10 months ago

Job Summary

Join Okta’s Emerging Products Group as a technical leader in Site Reliability Engineering. You’ll design and operate large-scale cloud infrastructure, run on-call, lead incident response and post-incident improvements, and define reliability metrics like SLIs/SLOs and error budgets. Build automation and operational guardrails using Go, Python, Terraform, CI/CD, and GitOps. Drive observability improvements with Datadog and Splunk while mentoring engineers and influencing architecture.
Location: Bengaluru
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Design, build, and operate large-scale cloud infrastructure and production services, including participating in on-call for highly available customer-facing systems.
  • •Lead incident response and drive post-incident reviews focused on systemic improvements.
  • •Define, measure, and improve SLIs, SLOs, and error budgets to enhance reliability and resilience.
  • •Develop automation and infrastructure using Go, Python, Terraform, and related technologies to eliminate operational toil.
  • •Improve observability and operational workflows using metrics, logging, tracing, dashboards, alerting, CI/CD, and GitOps practices while mentoring engineers.

Key Requirements

  • •Strong experience operating large-scale production services in AWS and/or GCP.
  • •Deep expertise with Kubernetes in production, including troubleshooting networking, storage, scheduling, scaling, and workload lifecycle issues.
  • •Extensive Infrastructure as Code experience with Terraform and Helm.
  • •Strong software engineering skills in Golang and/or Python, including building automation and internal engineering platforms.
  • •Experience operating and troubleshooting distributed data platforms such as PostgreSQL, Redis, OpenSearch, MySQL, or Cassandra.
Experience:SaaSKubernetesCloud
Skills:Technical leadershipIncident responseCollaborationCommunicationContinuous improvement
Languages:English
Tech Stack:GoGolangPythonTerraformHelmGitArgoCDGitOpsKubernetesEKSGKEDatadogSplunkPostgreSQLRedisOpenSearchMySQLCassandraCI/CDTLS

Company Brief

Okta
Provides identity and access management cloud solutions that help organizations secure and manage user authentication, single sign-on, multi-factor authentication, and lifecycle management across applications and devices.
Industry: Cybersecurity
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: San Francisco, United States
Founded: 2009
WebsiteLinkedIn