Sr. Staff Software Development Engineer - AI Platform

Zscaler
Bengaluru, Pune
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningExperience: 8+ yearsSkills: ["Ownership","Collaboration","Communication","Innovation","Data-driven decision-making"]

Design, scale, and maintain AWS-based infrastructure and platform capabilities for production AI/ML workloads. Build infrastructure automation, observability, and platform governance to support reliable, secure releases. Own AWS architecture, Terraform-based infrastructure, GitLab CI/CD pipelines, Prometheus/Grafana alerting and incident response, and DevOps/DORA-driven reliability improvements on Kubernetes (EKS). Mentor engineers through design and code reviews while partnering with AI/ML and data teams on AI/RAG deployments.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Zscaler
Zscaler
1 day ago

Sr. Staff Software Development Engineer - AI Platform

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 6 hours agoStatus: Live

Job Summary

Design, scale, and maintain AWS-based infrastructure and platform capabilities for production AI/ML workloads. Build infrastructure automation, observability, and platform governance to support reliable, secure releases. Own AWS architecture, Terraform-based infrastructure, GitLab CI/CD pipelines, Prometheus/Grafana alerting and incident response, and DevOps/DORA-driven reliability improvements on Kubernetes (EKS). Mentor engineers through design and code reviews while partnering with AI/ML and data teams on AI/RAG deployments.
Location: Bengaluru, Pune
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning

Key Responsibilities

  • •Design, build, and maintain scalable, secure, highly available AWS infrastructure for AI/ML workloads using Terraform and IaC best practices.
  • •Own and evolve GitLab CI/CD pipelines to automate build, test, security scanning, and multi-environment deployments.
  • •Architect centralized observability with Prometheus and Grafana, design alerting strategies, and lead incident response, root-cause analysis, and postmortems.
  • •Define and track DORA metrics and build self-healing, auto-scaling, and cost-optimized infrastructure for AI/ML services on Kubernetes (EKS).
  • •Implement infrastructure security best practices, establish platform governance standards, support production AI/RAG deployments, and mentor engineers through design/code reviews.

Pay and Benefits

Perks:Health InsurancePaid LeaveParental LeaveLearning Budget

Key Requirements

  • •8+ years as a Platform/Site Reliability/DevOps Engineer, including 3+ years supporting AI/ML or data platform infrastructure with deep AWS hands-on experience (EKS, Lambda, ECS, VPC, IAM, S3).
  • •Strong Infrastructure as Code experience with Terraform (modules, workspaces, remote state) and CI/CD pipeline design/maintenance with GitLab CI/CD (runners and deployment automation).
  • •Expertise in centralized observability with Prometheus and Grafana, including alerting/on-call systems (Alertmanager, PagerDuty, OpsGenie) and owning incident response.
  • •Solid hands-on understanding of DevOps/DORA metrics (deployment frequency, lead time for changes, change failure rate, MTTR) to drive engineering improvement.
  • •Experience deploying and operating containerized apps on Kubernetes (EKS), including Helm and autoscaling (HPA, Cluster Autoscaler, Karpenter), plus strong scripting in Python (and/or Go/Bash).
Experience:8+ yearsAI/MLData platforms
Skills:OwnershipCollaborationCommunicationInnovationData-driven decision-making
Languages:English
Tech Stack:AWSEKSLambdaECSVPCIAMS3TerraformGitLabGitLab CI/CDPrometheusGrafanaAlertmanagerPagerDutyOpsGenieKubernetesHelmHPACluster AutoscalerKarpenter

Company Brief

Zscaler
Provides cloud-native security platform delivering secure access, threat protection, and zero trust services to organizations, enabling secure internet and private application access without traditional network appliances.
Industry: Cybersecurity
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: San Jose, United States
Founded: 2007
WebsiteLinkedIn