Director, Site Reliability Engineering

Okta
Bengaluru
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureExperience: 16+ yearsSkills: ["Leadership","Communication","Cross-cultural collaboration","Mentoring","Process improvement"]

Lead Okta’s India-based Site Reliability Engineering organization to scale a highly available, AWS-hosted identity platform. Oversee reliability across platform, databases, edge networking, Kubernetes, CI/CD, observability, FinOps, and automation tooling. Partner globally to deliver resilient services, drive incident response and root-cause analysis, and implement automation and infrastructure-as-code to reduce toil and improve operational efficiency while hiring and developing SRE talent.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Okta
Okta
1 day ago

Director, Site Reliability Engineering

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 37 minutes agoStatus: Live

Job Summary

Lead Okta’s India-based Site Reliability Engineering organization to scale a highly available, AWS-hosted identity platform. Oversee reliability across platform, databases, edge networking, Kubernetes, CI/CD, observability, FinOps, and automation tooling. Partner globally to deliver resilient services, drive incident response and root-cause analysis, and implement automation and infrastructure-as-code to reduce toil and improve operational efficiency while hiring and developing SRE talent.
Location: Bengaluru
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Director level

Key Responsibilities

  • •Build and lead a high-caliber India-based SRE organization supporting Okta’s production fleet.
  • •Define and execute the India SRE strategy aligned with global reliability goals, partnering with global engineering, product, and infrastructure leaders.
  • •Lead post-incident reviews with root-cause analysis and long-term corrective actions; participate in incident management and on-call rotations.
  • •Implement automation and observability to reduce manual toil; drive adoption of infrastructure as code, Kubernetes, and infrastructure automation tooling.
  • •Hire, mentor, and develop SRE talent across India; improve SDLC processes for cloud infrastructure as code and CI/CD pipeline maturity.

Key Requirements

  • •16+ years of experience in site reliability, infrastructure, or production engineering roles.
  • •8+ years in technical leadership and people management, including managing managers, with expertise in automation, observability, performance optimization, and incident response.
  • •Experience building or scaling offshore SRE teams that partner with global counterparts.
  • •Strong expertise in cloud-native architectures, Kubernetes, Terraform (IaC), and CI/CD pipelines.
  • •Experience running an SRE org supporting a SaaS/Cloud service on a public cloud, preferably AWS.
Experience:16+ yearsSREInfrastructureProduction engineeringSaaSCloud
Education:
Skills:LeadershipCommunicationCross-cultural collaborationMentoringProcess improvement
Tech Stack:AWSTerraformKubernetesCI/CDObservabilityFinOpsInfrastructure as codeOn-call rotationsRoot-cause analysis

Company Brief

Okta
Provides identity and access management cloud solutions that help organizations secure and manage user authentication, single sign-on, multi-factor authentication, and lifecycle management across applications and devices.
Industry: Cybersecurity
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: San Francisco, United States
Founded: 2009
WebsiteLinkedIn