Sr. Staff Site Reliability Engineer

Zscaler
Bengaluru
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureExperience: 7+ yearsSkills: ["Ownership","Problem-solving","Collaboration","Integrity","Growth mindset"]

Own the reliability and operational integrity of a global infrastructure fleet spanning GCP, AWS, and bare-metal. Build self-healing automation and fault-tolerant provisioning, embed observability across telemetry pipelines, and partner with SWE teams to define production-readiness standards such as SLOs, health checks, and rollback requirements. Collaborate cross-functionally in an Agile/Scrum environment while strengthening infrastructure automation with durable, idempotent workflows.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Zscaler
Zscaler
1 day ago

Sr. Staff Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 5 hours agoStatus: Live

Job Summary

Own the reliability and operational integrity of a global infrastructure fleet spanning GCP, AWS, and bare-metal. Build self-healing automation and fault-tolerant provisioning, embed observability across telemetry pipelines, and partner with SWE teams to define production-readiness standards such as SLOs, health checks, and rollback requirements. Collaborate cross-functionally in an Agile/Scrum environment while strengthening infrastructure automation with durable, idempotent workflows.
Location: Bengaluru
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Build production-grade automation for bare-metal provisioning and IPMI/Redfish hardware lifecycle operations, including autonomous remediation
  • •Design fault-tolerant provisioning systems with self-healing and auto-remediation capabilities
  • •Develop telemetry pipelines integrating metrics, logs, and distributed tracing for hardware and system monitoring
  • •Partner with SWE teams to define and enforce production-readiness criteria, including SLOs, health-check standards, and rollback requirements
  • •Collaborate with cross-functional teams and contribute to Agile/Scrum processes such as sprint planning and retrospectives

Pay and Benefits

Perks:Health InsurancePaid LeaveParental LeavePensionLearning Budget

Key Requirements

  • •7+ years in production engineering, platform engineering, or infrastructure engineering in a cloud or hybrid environment
  • •Deep Linux/Unix systems mastery, including OS internals, kernel networking, boot pipelines, and low-level debugging
  • •Hands-on operational experience across GCP and/or AWS and on-prem datacenter operations
  • •Working command of infrastructure-as-code tooling (Terraform, Ansible) for both environments
  • •Expertise building durable, idempotent, repeatable automation using Python and/or Go
Experience:7+ yearsCloudHybrid environmentProduction engineeringPlatform engineeringInfrastructure engineering
Skills:OwnershipProblem-solvingCollaborationIntegrityGrowth mindset
Tech Stack:AnsiblePythonGoGCPAWSBare-metalIPMIRedfishPrometheusGrafanaVictoria MetricsTelegrafLinuxUnixTerraformObservabilitySLOsTemporal.ioCadenceIPMI/Redfish

Company Brief

Zscaler
Provides cloud-native security platform delivering secure access, threat protection, and zero trust services to organizations, enabling secure internet and private application access without traditional network appliances.
Industry: Cybersecurity
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: San Jose, United States
Founded: 2007
WebsiteLinkedIn