Site Reliability Engineer - SRE (Platform Software Team)

Arista Networks
Bengaluru
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Problem-solving","Software troubleshooting","Collaboration","Communication"]

Build and operate the software infrastructure that powers HW Labs and Manufacturing sites, enabling scalable, reliable, observable, and secure production systems. You’ll monitor and improve deployment and alerting, automate away toil, maintain incident response runbooks, and write postmortems. Partner with infrastructure and tools engineers and other Arista teams to diagnose bottlenecks, deploy staged system updates, and implement fault-tolerance, performance, and automation using modern SRE practices.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Arista Networks
Arista Networks
3 days ago

Site Reliability Engineer - SRE (Platform Software Team)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 16 hours agoStatus: Live

Job Summary

Build and operate the software infrastructure that powers HW Labs and Manufacturing sites, enabling scalable, reliable, observable, and secure production systems. You’ll monitor and improve deployment and alerting, automate away toil, maintain incident response runbooks, and write postmortems. Partner with infrastructure and tools engineers and other Arista teams to diagnose bottlenecks, deploy staged system updates, and implement fault-tolerance, performance, and automation using modern SRE practices.
Location: Bengaluru
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Build, deploy safely, incrementally, and operate critical production systems with a focus on scalability, reliability, observability, performance, and security.
  • •Monitor, support, and enhance the product deployment experience across services; build automation to reduce toil.
  • •Proactively monitor and improve alerts, set up automated alert handling, and create and maintain incident response runbooks.
  • •Write postmortems and implement solutions to prevent recurring incidents; deploy new systems in staged manners.
  • •Triage platform and infrastructural issues, assist software engineers, plan maintenance windows, and implement infrastructure solutions to address bottlenecks.

Key Requirements

  • •At least a Bachelor’s in Computer Science or Engineering with 5+ years of experience, or an MS with 5+ years, or equivalent work experience.
  • •Knowledge of Go, Python, or bash shell scripting to implement medium-complexity automation workflows.
  • •Linux (or UNIX) knowledge for administration and debugging.
  • •Hands-on experience deploying and managing Kubernetes on bare-metal infrastructure.
  • •Experience with server provisioning using Ansible and Ansible Tower/AWX.
Experience:Infrastructure automationKubernetesSRE
Skills:Problem-solvingSoftware troubleshootingCollaborationCommunication
Languages:English
Tech Stack:GoPythonBashLinuxUNIXKubernetesAnsibleAnsible TowerAWXDockerVirtualizationAir-gapped systemsInfrastructure-as-codeMySQLPrometheusGrafanaArtifactoryGitLab CI/CDJenkinsZuul

Company Brief

Arista Networks
Designs and sells high-performance cloud networking hardware and software for large-scale data centers and enterprise environments, including switches, routers, and network operating systems focused on programmability, telemetry, and low-latency networking.
Industry: Networking Equipment
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 2004
Glassdoor
Glassdoor: 4.1
WebsiteLinkedIn