Staff Site Reliability Engineer (Linux/Network troubleshooting/Scripting)

Zscaler
Bengaluru
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureExperience: 4+ yearsSkills: ["Problem-solving","Data-driven","Growth mindset","Cross-functional collaboration","Innovation"]

Design, automate, and operate next-generation cloud infrastructure for large-scale distributed systems, ensuring high availability, security, and resilience. Lead cloud operations and incident management while tuning Linux/BSD systems. Oversee container orchestration on EKS/GKE, build scalable observability with monitoring and alerting (including Grafana, SLIs/SLOs), and collaborate across engineering teams to deliver secure, scalable solutions that improve production reliability.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Zscaler
Zscaler
5 days ago

Staff Site Reliability Engineer (Linux/Network troubleshooting/Scripting)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 2 hours agoStatus: Live

Job Summary

Design, automate, and operate next-generation cloud infrastructure for large-scale distributed systems, ensuring high availability, security, and resilience. Lead cloud operations and incident management while tuning Linux/BSD systems. Oversee container orchestration on EKS/GKE, build scalable observability with monitoring and alerting (including Grafana, SLIs/SLOs), and collaborate across engineering teams to deliver secure, scalable solutions that improve production reliability.
Location: Bengaluru
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Sr. Manager level

Key Responsibilities

  • •Design and implement cloud management automations to reduce toil and accelerate delivery.
  • •Optimize containerized architectures using EKS and GKE for robust production performance.
  • •Lead building and tuning of scalable monitoring and alerting systems for observability.
  • •Own cloud operations, deployments, on-call support, and incident management while continuously tuning Linux and BSD-based systems.
  • •Collaborate with cross-functional teams to consult on feasibility and deliver technology solutions.

Pay and Benefits

Perks:Health InsurancePaid LeaveParental LeavePensionLearning Budget

Key Requirements

  • •4+ years designing, analyzing, and troubleshooting large-scale distributed systems.
  • •Deep hands-on SRE best practices with strong technical experience in Python automation, Terraform, Ansible, networking, Kubernetes, and AWS cloud.
  • •Comprehensive understanding of web security and core protocols including HTTP, SSL/TLS, DNS, and networking fundamentals.
  • •Proven observability experience building dashboards and managing Grafana, including SLIs, SLOs, and error budgets.
  • •Strong DevOps skills across CI/CD, source control management, builds/releases, and continuous integration tools.
Experience:4+ yearsDistributed systemsSREDevOpsCloud infrastructure
Skills:Problem-solvingData-drivenGrowth mindsetCross-functional collaborationInnovation
Tech Stack:PythonTerraformAnsibleAWSEKSGKELinuxBSDGrafanaKubernetesHTTPSSL/TLSDNSSQLCI/CDCICDSource Control ManagementSCMSLIs

Company Brief

Zscaler
Provides cloud-native security platform delivering secure access, threat protection, and zero trust services to organizations, enabling secure internet and private application access without traditional network appliances.
Industry: Cybersecurity
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: San Jose, United States
Founded: 2007
WebsiteLinkedIn