Site Reliability Engineer (SRE) - Engineering Productivity

Arista Networks
Bengaluru
Workplace: RemoteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsEducation: bachelorsSkills: ["Problem-solving","Software troubleshooting","Incident response","Automation"]

Build, deploy, and operate critical production systems for engineering productivity, with a focus on scalability, reliability, observability, performance, and security. Improve developer experience by monitoring and enhancing services, automating workflows, and proactively handling alerts. Create incident response runbooks and postmortems, triage platform/infrastructure issues, and work with product teams to remove workflow bottlenecks using tools and platforms such as Kubernetes, Jenkins, Grafana, Spinnaker, and Google Cloud.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Arista Networks
Arista Networks
2 days ago

Site Reliability Engineer (SRE) - Engineering Productivity

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 36 minutes agoStatus: Live

Job Summary

Build, deploy, and operate critical production systems for engineering productivity, with a focus on scalability, reliability, observability, performance, and security. Improve developer experience by monitoring and enhancing services, automating workflows, and proactively handling alerts. Create incident response runbooks and postmortems, triage platform/infrastructure issues, and work with product teams to remove workflow bottlenecks using tools and platforms such as Kubernetes, Jenkins, Grafana, Spinnaker, and Google Cloud.
Location: Bengaluru
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Build, deploy, and operate critical production systems, emphasizing scalability, reliability, observability, performance, and security.
  • •Monitor, support, and enhance developer experience; build automation to remove toil and improve production operations.
  • •Proactively monitor and respond to alerts with automated alert handling; create and maintain incident response runbooks.
  • •Triage platform/infrastructure issues and partner with software engineers to resolve bottlenecks and limitations in workflows.
  • •Plan maintenance windows, deploy systems in staged manner, and write postmortem documents with follow-up solutions to prevent repeat incidents.

Key Requirements

  • •BSc Computer Science/Engineering with 5+ years experience, or MS Computer Science/Engineering with 5+ years experience, or equivalent work experience.
  • •Knowledge of one or more of Go, Python, or shell scripting to implement automation workflows.
  • •Linux/UNIX knowledge for administration and debugging.
  • •Hands-on experience operating software systems at scale and server provisioning (storage and networking perspective).
  • •Strong problem-solving and software troubleshooting skills with experience in infrastructure-as-code.
Experience:5+ years
Education:Bachelor's in Computer Science or Engineering
Skills:Problem-solvingSoftware troubleshootingIncident responseAutomation
Languages:English
Tech Stack:GoPythonShell scriptingLinuxAnsibleArtifactoryGerritJenkinsKubernetesGrafanaSpinnakerMySQLElasticSearchGoogle CloudVarnishPerforceDockerPrometheusLokiTempo

Company Brief

Arista Networks
Designs and sells high-performance cloud networking hardware and software for large-scale data centers and enterprise environments, including switches, routers, and network operating systems focused on programmability, telemetry, and low-latency networking.
Industry: Networking Equipment
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 2004
Glassdoor
Glassdoor: 4.1
WebsiteLinkedIn