Senior Site Reliability Engineer

Okta
Bengaluru
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Continuous improvement","Operational excellence","Incident response","Collaboration","Communication"]

Build and operate highly reliable, scalable cloud services for Okta’s Emerging Products Group, with an automation-first SRE mindset. You’ll lead incident response and drive post-incident learning, define SLIs/SLOs and error budgets, and improve observability and operational workflows. Collaborate with software engineers and product teams to enhance availability, performance, and resilience, while developing automation and infrastructure using Go, Python, Terraform, and GitOps.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Okta
Okta
1 day ago

Senior Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 6 hours agoStatus: Live
Reposted: similar role first listed 7 months ago

Job Summary

Build and operate highly reliable, scalable cloud services for Okta’s Emerging Products Group, with an automation-first SRE mindset. You’ll lead incident response and drive post-incident learning, define SLIs/SLOs and error budgets, and improve observability and operational workflows. Collaborate with software engineers and product teams to enhance availability, performance, and resilience, while developing automation and infrastructure using Go, Python, Terraform, and GitOps.
Location: Bengaluru
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Design, build, and operate large-scale cloud infrastructure and production services.
  • •Participate in an on-call rotation and lead incident response with post-incident reviews to drive systemic improvements.
  • •Define, measure, and improve SLIs, SLOs, and error budgets while partnering with engineering teams to improve availability, scalability, performance, and resilience.
  • •Continuously improve observability via metrics, logging, tracing, dashboards, and alerting.
  • •Develop automation and infrastructure using Go, Python, Terraform and related technologies, reducing operational toil and improving CI/CD and GitOps workflows.

Key Requirements

  • •Strong experience operating large-scale production services in AWS and/or GCP.
  • •Deep Kubernetes expertise in production, including troubleshooting networking, storage, scheduling, scaling, and workload lifecycle.
  • •Extensive Infrastructure as Code experience with Terraform and Helm.
  • •Strong software engineering skills in Golang and/or Python, plus experience building automation and internal engineering platforms.
  • •Hands-on observability and reliability knowledge (SLIs, SLOs, error budgets) and incident response leadership.
Experience:SaaSCloud infrastructureKubernetesPlatform engineeringDistributed systemsObservability
Skills:Continuous improvementOperational excellenceIncident responseCollaborationCommunication
Languages:English
Tech Stack:KubernetesEKSGKETerraformHelmGitArgoCDGitOpsGolangGoPythonDatadogSplunkPostgreSQLRedisOpenSearchMySQLCassandraAI-assisted engineeringCI/CD

Company Brief

Okta
Provides identity and access management cloud solutions that help organizations secure and manage user authentication, single sign-on, multi-factor authentication, and lifecycle management across applications and devices.
Industry: Cybersecurity
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: San Francisco, United States
Founded: 2009
WebsiteLinkedIn