Staff Site Reliability Engineer

Okta
Bengaluru
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Continuous improvement","Operational excellence","Technical leadership","Collaboration","Communication"]

Build and operate highly reliable, scalable, and secure cloud services for Okta’s Emerging Products Group. You’ll lead incident response and drive post-incident improvements, define SLIs/SLOs and error budgets, and partner with engineering teams to improve availability, performance, and resilience. The role focuses on automation-first reliability—developing software and infrastructure with Go, Python, Terraform, Kubernetes, and GitOps practices—plus technical leadership and mentoring across multiple teams.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Okta
Okta
2 days ago

Staff Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 3 hours agoStatus: Live

Job Summary

Build and operate highly reliable, scalable, and secure cloud services for Okta’s Emerging Products Group. You’ll lead incident response and drive post-incident improvements, define SLIs/SLOs and error budgets, and partner with engineering teams to improve availability, performance, and resilience. The role focuses on automation-first reliability—developing software and infrastructure with Go, Python, Terraform, Kubernetes, and GitOps practices—plus technical leadership and mentoring across multiple teams.
Location: Bengaluru
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Sr. Manager level

Key Responsibilities

  • •Design, build, and operate large-scale cloud infrastructure and production services.
  • •Participate in an on-call rotation for highly available customer-facing systems.
  • •Lead incident response and drive post-incident reviews focused on systemic improvements.
  • •Define, measure, and improve SLIs, SLOs, and error budgets while partnering to improve availability, scalability, performance, and resilience.
  • •Develop automation and infrastructure using Go, Python, Terraform, and related tools; improve observability and deployment safety using CI/CD and GitOps.

Key Requirements

  • •Strong experience operating large-scale production services in AWS and/or GCP.
  • •Deep expertise with Kubernetes in production, including troubleshooting networking, storage, scheduling, scaling, and workload lifecycle issues.
  • •Extensive infrastructure-as-code experience with Terraform and Helm.
  • •Strong software engineering skills in Golang and/or Python, including building automation and internal engineering platforms.
  • •Experience with observability/monitoring and reliability engineering concepts such as SLIs, SLOs, error budgets, and capacity planning.
Experience:SaaSCloudMicroservicesDistributed systems
Skills:Continuous improvementOperational excellenceTechnical leadershipCollaborationCommunication
Languages:English
Tech Stack:KubernetesEKSGKETerraformHelmGitArgoCDGitOpsGolangGoPythonDatadogSplunkPostgreSQLRedisOpenSearchMySQLCassandraCI/CD

Company Brief

Okta
Provides identity and access management cloud solutions that help organizations secure and manage user authentication, single sign-on, multi-factor authentication, and lifecycle management across applications and devices.
Industry: Cybersecurity
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: San Francisco, United States
Founded: 2009
WebsiteLinkedIn