Senior Lead Site Reliability Engineer

Zoom Video Communications
San Jose
Full timeUSD 146,700 - 339,300 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 10+ yearsSkills: ["Communication","Mentoring","Technical leadership","Cross-team collaboration","Incident management"]

Lead cross-team SRE efforts to design, operate, and improve Zoom’s hybrid infrastructure across global data centers. You’ll install, configure, patch, and monitor large-scale physical and cloud systems, develop automation to reduce toil, and tackle performance bottlenecks. Serve as a technical lead for major incidents, mentor engineers on best practices, and partner with Security, Networking, and Platform teams to build self-healing, resilient platforms while optimizing Linux systems at scale.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Zoom Video Communications
Zoom Video Communications
4 days ago

Senior Lead Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 18 hours agoStatus: Live

Job Summary

Lead cross-team SRE efforts to design, operate, and improve Zoom’s hybrid infrastructure across global data centers. You’ll install, configure, patch, and monitor large-scale physical and cloud systems, develop automation to reduce toil, and tackle performance bottlenecks. Serve as a technical lead for major incidents, mentor engineers on best practices, and partner with Security, Networking, and Platform teams to build self-healing, resilient platforms while optimizing Linux systems at scale.
Location: San Jose
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Sr. Manager level

Key Responsibilities

  • •Provide technical direction for cross-team initiatives and major incidents, participating in on-call shifts and incident management.
  • •Mentor SREs and developers by defining best practices and design patterns.
  • •Partner with Security, Networking, and Platform teams on architecture roadmaps and influence vendor and hardware strategy for on-prem and cloud workloads.
  • •Design self-healing platforms using automation, chaos engineering, and fault-tolerant patterns.
  • •Optimize Linux systems at scale (performance tuning, kernel parameters, networking, storage, and security hardening) and maintain system firewalls while troubleshooting connectivity and access permissions.

Pay and Benefits

Salary: USD 146,700 - 339,300 annually
Equity and Bonus:Equity

Key Requirements

  • •10+ years in SRE, production engineering, or large-scale systems administration.
  • •Strong Linux system administration experience (systemd, cgroups, networking, filesystems, performance analysis).
  • •Demonstrated coding ability in at least one programming language (e.g., Python).
  • •Experience with configuration management and IaC (Ansible, Terraform, Packer), plus CI/CD (Jenkins, GitLab) and container orchestration (k8s, Docker).
  • •Security-first and networking expertise, including incident response in mission-critical environments, plus resilience practices such as chaos engineering.
Experience:10+ years
Skills:CommunicationMentoringTechnical leadershipCross-team collaborationIncident management
Tech Stack:LinuxSystemdCgroupsPythonAnsibleTerraformPackerJenkinsGitLabKubernetesK8sDockerTPMSecure bootSecrets managementBGPDNSTLSTraffic engineeringChaos engineering

Company Brief

Zoom Video Communications
Provides video conferencing, chat, phone, webinar, and collaboration software for businesses, schools, and individuals. Its platform is widely used for remote meetings, virtual events, and hybrid work communication.
Industry: SaaS
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: San Jose, United States
Founded: 2011
WebsiteLinkedIn