Principal Production Engineer

Zscaler
San Jose
Workplace: HybridFull timeUSD 164,500 - 235,000 annuallyFunction: Manufacturing & Production OperationsExperience: 10+ yearsSkills: ["Python","Go","C/C++","Prometheus","Grafana","OpenTelemetry","Ansible","Terraform","Helm","Temporal","AWS","GCP","Networking","Linux","SRE","On-call","Incident management"]

Lead the design and implementation of highly available, scalable multi-cloud infrastructure. Drive an automation-first culture, improve observability, and define SLIs/SLOs while acting as Incident Commander to reduce MTTM. Partner with engineering teams to drive operability and reliability for a global platform processing billions of transactions daily.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Zscaler
Zscaler
2 months ago

Principal Production Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 9 hours agoStatus: Live

Job Summary

Lead the design and implementation of highly available, scalable multi-cloud infrastructure. Drive an automation-first culture, improve observability, and define SLIs/SLOs while acting as Incident Commander to reduce MTTM. Partner with engineering teams to drive operability and reliability for a global platform processing billions of transactions daily.
Location: San Jose
Workplace: Hybrid
Employment Type: Full time
Job Function: Manufacturing & Production Operations

Key Responsibilities

  • •Design and implement highly available, scalable infrastructure across AWS, GCP, and bare-metal environments.
  • •Drive an 'automation-first' culture by writing code (Python/Go) to eliminate manual toil and build self-healing systems.
  • •Implement and maintain sophisticated observability (Prometheus, Grafana, OpenTelemetry), define SLIs/SLOs, and establish error budgets.
  • •Act as a lead Incident Commander (TDO on-call), develop response playbooks, and conduct deep-dive post-incident analyses.
  • •Partner with Engineering and partner teams to conduct operability reviews.

Pay and Benefits

Salary: USD 164,500 - 235,000 annually
Perks:Health InsuranceTime OffParental LeaveRetirement Options

Key Requirements

  • •10+ years of experience managing reliability, scalability, and availability for large-scale production services.
  • •Deep expertise in programming (Python, Go, or C/C++).
  • •Strong background in networking protocols, Linux/RHEL systems, and distributed architecture.
  • •Experience in high-stakes incident management and participation in a 24/7 on-call rotation.
  • •Proficiency in leveraging ITIL frameworks and incident data to drive service maturity through systematic problem management and technical operability reviews.
Experience:10+ years
Skills:PythonGoC/C++PrometheusGrafanaOpenTelemetryAnsibleTerraformHelmTemporalAWSGCPNetworkingLinuxSREOn-callIncident management
Languages:English
Tech Stack:PythonGoC/C++PrometheusGrafanaOpenTelemetryAnsibleTerraformHelmTemporalAWSGCPAzureLinuxRHELKubernetesHAProxyBGPDNSGRE

Company Brief

Zscaler
Provides cloud-native security platform delivering secure access, threat protection, and zero trust services to organizations, enabling secure internet and private application access without traditional network appliances.
Industry: Cybersecurity
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: San Jose, United States
Founded: 2007
WebsiteLinkedIn