Staff Site Reliability Engineer - Ecosystem

Okta
Bengaluru
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureExperience: 8+ yearsEducation: bachelorsSkills: ["Communication","Problem-solving","Mentoring","Cross-team collaboration"]

Build and operate highly scalable, reliable infrastructure across AWS and GCP, leading reliability and modernization work including container platform migrations (ECS to EKS/GKE) and microservice enablement. Serve as a Kubernetes and CI/CD authority, implement infrastructure as code with Terraform/Ansible, and improve observability, performance, and cost via monitoring, logging, and alerting. Lead initiatives end-to-end, participate in on-call, and mentor engineers while partnering with security and compliance teams.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Okta
Okta
1 month ago

Staff Site Reliability Engineer - Ecosystem

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 8 hours agoStatus: Live

Job Summary

Build and operate highly scalable, reliable infrastructure across AWS and GCP, leading reliability and modernization work including container platform migrations (ECS to EKS/GKE) and microservice enablement. Serve as a Kubernetes and CI/CD authority, implement infrastructure as code with Terraform/Ansible, and improve observability, performance, and cost via monitoring, logging, and alerting. Lead initiatives end-to-end, participate in on-call, and mentor engineers while partnering with security and compliance teams.
Location: Bengaluru
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Design, build, and operate scalable, reliable, secure production infrastructure across AWS and GCP.
  • •Lead reliability and modernization initiatives, including container platform migrations (ECS to EKS/GKE) and microservice enablement across multi-cloud environments.
  • •Act as a technical authority in Kubernetes (EKS and GKE), cloud infrastructure (AWS and GCP), and modern CI/CD practices (GitOps and automation pipelines).
  • •Implement and manage infrastructure as code using Terraform and Ansible for provisioning, scaling, and configuration across cloud providers.
  • •Improve observability, performance, and cost with monitoring, logging, and alerting systems; define SLOs/SLIs and run blameless postmortems while participating in on-call and mentoring engineers.

Key Requirements

  • •8+ years in SRE, DevOps, or Infrastructure Engineering roles.
  • •3–5 years with Kubernetes (EKS/GKE) in production, including related ecosystem tools like Helm and Karpenter.
  • •3–5 years with AWS and GCP, plus 3–5 years using Terraform for multi-cloud infrastructure.
  • •3+ years of coding experience in Python, Go, or similar languages focused on automation and operational excellence.
  • •Hands-on experience with CI/CD pipelines, observability tools, and cloud databases/caching (e.g., Prometheus/Grafana, ELK/Loki, RDS/Cloud SQL, Redis/Memorystore).
Experience:8+ yearsSaaSCloud-nativeMicroservicesMulti-cloudDistributed systems
Education:Bachelor's in Computer Science
Skills:CommunicationProblem-solvingMentoringCross-team collaboration
Languages:English
Tech Stack:AWSGCPKubernetesEKSGKEECSEKS/GKEEKS to EKS/GKETerraformAnsibleCloudFormationPythonGoShellGitOpsCI/CDArgoCDGitLab CISpinnakerLinux

Company Brief

Okta
Provides identity and access management cloud solutions that help organizations secure and manage user authentication, single sign-on, multi-factor authentication, and lifecycle management across applications and devices.
Industry: Cybersecurity
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: San Francisco, United States
Founded: 2009
WebsiteLinkedIn