Sr. Production Engineer

Zscaler
Bellevue, San Jose, United States
Workplace: HybridFull timeUSD 103,600 - 148,000 annuallyFunction: Manufacturing & Production OperationsExperience: 4+ yearsSkills: ["Ownership","Problem-solving","Collaboration","Accountability","Growth mindset"]

Design and implement highly available, scalable infrastructure for a global platform across AWS, Azure, GCP, and bare-metal. Lead an automation-first approach by writing Python/Go to eliminate toil and build self-healing systems. Own observability using Prometheus, Grafana, and OpenTelemetry, define SLIs/SLOs and error budgets, and drive incident command with playbooks and post-incident analysis.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Zscaler
Zscaler
1 day ago

Sr. Production Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 10 hours agoStatus: Live

Job Summary

Design and implement highly available, scalable infrastructure for a global platform across AWS, Azure, GCP, and bare-metal. Lead an automation-first approach by writing Python/Go to eliminate toil and build self-healing systems. Own observability using Prometheus, Grafana, and OpenTelemetry, define SLIs/SLOs and error budgets, and drive incident command with playbooks and post-incident analysis.
Location: Bellevue, San Jose, United States
Workplace: Hybrid
Employment Type: Full time
Job Function: Manufacturing & Production Operations
Seniority: Mid level

Key Responsibilities

  • •Design and implement highly available, scalable infrastructure across AWS, Azure, GCP, and bare-metal environments.
  • •Drive an automation-first culture by writing Python/Go to eliminate manual toil and build self-healing systems.
  • •Implement and maintain observability (Prometheus, Grafana, OpenTelemetry), defining SLIs/SLOs and error budgets.
  • •Serve as lead Incident Commander (TDO on-call), develop response playbooks, and run deep-dive post-incident analyses.
  • •Partner with Engineering and other teams to conduct operability reviews.

Pay and Benefits

Salary: USD 103,600 - 148,000 annually
Perks:Health InsurancePaid LeaveParental LeaveRetirementLearning Budget

Key Requirements

  • •4 years of experience managing reliability, scalability, and availability for large-scale production services with deep programming expertise (e.g., Python, Go, or C/C++).
  • •Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes.
  • •Strong background in networking protocols, Linux/FreeBSD systems, and distributed architecture.
  • •Experience in high-stakes incident management and participation in a 24/7 on-call rotation.
  • •Proficiency using ITIL frameworks and incident data for service maturity via systematic problem management and technical operability reviews.
Experience:4+ yearsProduction servicesDistributed architectureIncident managementMulti-cloudInfrastructure reliability
Skills:OwnershipProblem-solvingCollaborationAccountabilityGrowth mindset
Languages:English
Tech Stack:AWSAzureGCPBare-metalPythonGoPrometheusGrafanaOpenTelemetrySLIsSLOsAnsibleTerraformAIOpsMachine learningAnomaly detectionPredictive capacity planningChaos engineeringDisaster recoveryBGP

Company Brief

Zscaler
Provides cloud-native security platform delivering secure access, threat protection, and zero trust services to organizations, enabling secure internet and private application access without traditional network appliances.
Industry: Cybersecurity
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: San Jose, United States
Founded: 2007
WebsiteLinkedIn