Sr. Staff Production Engineer

Zscaler
Bellevue, San Jose
Workplace: HybridFull timeUSD 143,500 - 205,000 annuallyFunction: Manufacturing & Production OperationsExperience: 8+ yearsSkills: ["Problem-solving","Collaboration","Ownership","Integrity","High-trust communication"]

Own reliability for a global, multi-cloud platform by designing and implementing highly available infrastructure across AWS, Azure, GCP, and bare-metal. Drive an automation-first culture using Python/Go to build self-healing systems, and mature observability with Prometheus, Grafana, and OpenTelemetry to define SLIs/SLOs and error budgets. Lead incident response as Incident Commander on TDO on-call, run operability reviews, and analyze post-incident root causes.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Zscaler
Zscaler
2 days ago

Sr. Staff Production Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 17 hours agoStatus: Live

Job Summary

Own reliability for a global, multi-cloud platform by designing and implementing highly available infrastructure across AWS, Azure, GCP, and bare-metal. Drive an automation-first culture using Python/Go to build self-healing systems, and mature observability with Prometheus, Grafana, and OpenTelemetry to define SLIs/SLOs and error budgets. Lead incident response as Incident Commander on TDO on-call, run operability reviews, and analyze post-incident root causes.
Location: Bellevue, San Jose
Workplace: Hybrid
Employment Type: Full time
Job Function: Manufacturing & Production Operations
Seniority: Mid level

Key Responsibilities

  • •Design and implement highly available, scalable infrastructure across AWS, Azure, GCP, and bare-metal environments.
  • •Drive an automation-first culture by writing code (Python/Go) to eliminate manual toil and build self-healing systems.
  • •Implement and maintain observability (Prometheus, Grafana, OpenTelemetry), define SLIs/SLOs, and establish error budgets.
  • •Lead incident response as Incident Commander (TDO on-call), develop response playbooks, and conduct post-incident deep dives.
  • •Partner with engineering and partner teams to conduct operability reviews and improve service reliability.

Pay and Benefits

Salary: USD 143,500 - 205,000 annually
Perks:Health InsurancePaid LeaveParental LeaveRetirementLearning Budget

Key Requirements

  • •8+ years managing reliability, scalability, and availability for large-scale production services with deep programming expertise (e.g., Python, Go, or C/C++).
  • •Strong background in networking protocols, Linux/FreeBSD systems, and distributed architecture.
  • •Experience in high-stakes incident management, including participation in a 24/7 on-call rotation.
  • •Proficiency using ITIL frameworks and incident data to drive service maturity through problem management and technical operability reviews.
  • •Foundational understanding of AI/ML technologies and experience leveraging AI-driven solutions to optimize outcomes.
Experience:8+ yearsCloudCybersecurityAI/ML
Skills:Problem-solvingCollaborationOwnershipIntegrityHigh-trust communication
Tech Stack:AWSAzureGCPBare-metalPythonGoPrometheusGrafanaOpenTelemetrySLIsSLOsITILAnsibleTerraformChaos engineeringDisaster recoveryBGPGREIPSecHAProxy

Company Brief

Zscaler
Provides cloud-native security platform delivering secure access, threat protection, and zero trust services to organizations, enabling secure internet and private application access without traditional network appliances.
Industry: Cybersecurity
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: San Jose, United States
Founded: 2007
WebsiteLinkedIn