Production Engineer

Zscaler
San Jose, United States
Workplace: HybridFull timeUSD 102,400 - 128,000Function: Manufacturing & Production OperationsExperience: 1-3 yearsSkills: ["Bias for action","Integrity","Pragmatic execution","Simplicity","Data-driven decision-making"]

Own automation-first reliability for a global, multi-cloud platform. Implement highly available infrastructure across AWS, GCP, and bare metal, and reduce toil by writing Python/Go for self-healing systems. Improve observability with Prometheus, Grafana, and OpenTelemetry, define SLIs/SLOs and error budgets, and lead incidents as the TDO on-call. Partner on operability reviews to shape scalable service practices.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Zscaler
Zscaler
2 months ago

Production Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 7 hours agoStatus: Live

Job Summary

Own automation-first reliability for a global, multi-cloud platform. Implement highly available infrastructure across AWS, GCP, and bare metal, and reduce toil by writing Python/Go for self-healing systems. Improve observability with Prometheus, Grafana, and OpenTelemetry, define SLIs/SLOs and error budgets, and lead incidents as the TDO on-call. Partner on operability reviews to shape scalable service practices.
Location: San Jose, United States
Workplace: Hybrid
Employment Type: Full time
Job Function: Manufacturing & Production Operations
Seniority: Entry level

Key Responsibilities

  • •Implement highly available, scalable infrastructure across AWS, GCP, and bare-metal environments.
  • •Drive an automation-first culture by writing Python/Go to eliminate manual toil and build self-healing systems.
  • •Implement and maintain observability with Prometheus, Grafana, and OpenTelemetry; define SLIs/SLOs and establish error budgets.
  • •Serve as lead Incident Commander (TDO on-call), develop response playbooks, and conduct deep-dive post-incident analyses.
  • •Partner with engineering teams to conduct operability reviews.

Pay and Benefits

Salary: USD 102,400 - 128,000
Perks:Health InsurancePaid LeaveParental LeaveRetirementLearning Budget

Key Requirements

  • •Demonstrated curiosity and active exploration of AI tools, including integrating new technologies to improve workflows.
  • •1-3 years of experience managing reliability, scalability, and availability for large-scale production services.
  • •Deep expertise in programming (e.g., Python, Go, or C/C++).
  • •Strong background in networking protocols, Linux/RHEL systems, and distributed architecture.
  • •Experience in high-stakes incident management and participating in a 24/7 on-call rotation.
Experience:1-3 yearsCloud infrastructureDistributed systemsIncident managementCybersecurity
Skills:Bias for actionIntegrityPragmatic executionSimplicityData-driven decision-making
Languages:English
Tech Stack:AWSGCPPythonGoC/C++PrometheusGrafanaOpenTelemetrySLIs/SLOsITILAnsibleTerraformHelmTemporalLinuxRHELBGPGREIPSecHAProxy

Company Brief

Zscaler
Provides cloud-native security platform delivering secure access, threat protection, and zero trust services to organizations, enabling secure internet and private application access without traditional network appliances.
Industry: Cybersecurity
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: San Jose, United States
Founded: 2007
WebsiteLinkedIn