Sr. Production Engineer

Zscaler
San Jose
Workplace: HybridFull timeUSD 118,400 - 148,000 annuallyFunction: Manufacturing & Production OperationsExperience: 3-5 yearsSkills: ["Ownership","Bias for action","Collaboration","Integrity","Data-driven"]

Drive an automation-first culture to improve the reliability of a global, multi-cloud platform processing 200+ billion transactions daily. Provide hands-on execution and technical vision by maturing observability and architectural standards, implementing highly available infrastructure across AWS, GCP, and bare metal, defining SLIs/SLOs and error budgets, and leading incident response as Incident Commander. Partner on operability reviews to reduce MTTM and strengthen scalability.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Zscaler
Zscaler
2 months ago

Sr. Production Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 7 hours agoStatus: Live

Job Summary

Drive an automation-first culture to improve the reliability of a global, multi-cloud platform processing 200+ billion transactions daily. Provide hands-on execution and technical vision by maturing observability and architectural standards, implementing highly available infrastructure across AWS, GCP, and bare metal, defining SLIs/SLOs and error budgets, and leading incident response as Incident Commander. Partner on operability reviews to reduce MTTM and strengthen scalability.
Location: San Jose
Workplace: Hybrid
Employment Type: Full time
Job Function: Manufacturing & Production Operations
Seniority: Mid level

Key Responsibilities

  • •Implement highly available, scalable infrastructure across AWS, GCP, and bare-metal environments.
  • •Write automation code (Python/Go) to eliminate manual toil and build self-healing systems.
  • •Implement and maintain observability (Prometheus, Grafana, OpenTelemetry), define SLIs/SLOs, and establish error budgets.
  • •Serve as lead Incident Commander on TDO on-call, build response playbooks, and run deep-dive post-incident analyses.
  • •Partner with Engineering teams to conduct operability reviews and improve reliability practices.

Pay and Benefits

Salary: USD 118,400 - 148,000 annually
Perks:Health InsurancePaid LeaveParental LeaveRetirement OptionsLearning Budget

Key Requirements

  • •3-5+ years managing reliability, scalability, and availability for large-scale production services.
  • •Deep expertise in programming (Python, Go, or C/C++).
  • •Strong background in networking protocols, Linux/RHEL systems, and distributed architecture.
  • •Experience with high-stakes incident management and participating in a 24/7 on-call rotation.
  • •Proficiency using ITIL frameworks and incident data to improve service maturity through systematic problem management and operability reviews.
Experience:3-5 yearsCloud infrastructureMulti-cloudProduction services
Skills:OwnershipBias for actionCollaborationIntegrityData-driven
Languages:English
Tech Stack:AWSGCPPythonGoC/C++PrometheusGrafanaOpenTelemetryLinuxRHELITILAnsibleTerraformHelmTemporalChaos engineeringDisaster recoveryBGPGREIPSec

Company Brief

Zscaler
Provides cloud-native security platform delivering secure access, threat protection, and zero trust services to organizations, enabling secure internet and private application access without traditional network appliances.
Industry: Cybersecurity
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: San Jose, United States
Founded: 2007
WebsiteLinkedIn