Senior Network Site Reliability Engineer

NVIDIA
India
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 10+ yearsEducation: bachelorsSkills: ["Problem-solving","Critical thinking","Ownership","Interpersonal communication","Drive"]

Own the operational reliability of enterprise network infrastructure, driving high availability through incident and service-request management, monitoring, and proactive risk mitigation. Partner with architecture and deployment teams to ensure new implementations meet production standards. Lead automation initiatives to reduce toil and improve SLOs, using observability and debugging to implement refinements. Perform blameless postmortems/RCAs and publish automation/bot knowledge-base content.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
1 week ago

Senior Network Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 1 hour agoStatus: Live

Job Summary

Own the operational reliability of enterprise network infrastructure, driving high availability through incident and service-request management, monitoring, and proactive risk mitigation. Partner with architecture and deployment teams to ensure new implementations meet production standards. Lead automation initiatives to reduce toil and improve SLOs, using observability and debugging to implement refinements. Perform blameless postmortems/RCAs and publish automation/bot knowledge-base content.
Location: India
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Own network infrastructure operations to ensure high availability and reliability, handling incidents and service requests.
  • •Monitor network performance, identify improvement areas, and collaborate to implement refinements and continuous improvements.
  • •Partner with architecture and deployment teams to ensure new implementations are supportable and meet production standards.
  • •Drive automation initiatives to reduce toil, improve operational efficiency, and maintain Service Level Objectives (SLOs).
  • •Run blameless postmortems and follow through on Root Cause Analyses (RCAs), and publish knowledge-base articles for automation and bots.

Key Requirements

  • •10+ years in network operations or related support, automation, and site reliability engineering, including enterprise and data center networks.
  • •Strong network fundamentals and ability to troubleshoot complex issues using TCP/UDP, IPv4/IPv6, BGP, OSPF/ISIS, VPN, L2 switching, firewalls, and load balancers.
  • •Hands-on experience with observability/monitoring tools including Prometheus, Grafana, Alertmanager, Nautobot/NetBox, and BigPanda.
  • •Automation experience using Salt, Ansible, Python (or similar) plus operational signals such as SNMP, Syslog, and streaming telemetry.
  • •Linux system fundamentals and proficiency in scripting/programming such as Python or Go, with knowledge of ServiceNow/Jira and ITIL fundamentals.
Experience:10+ yearsNetwork operationsSite reliability engineeringData center networksEnterprise networks
Education:Bachelor's
Skills:Problem-solvingCritical thinkingOwnershipInterpersonal communicationDrive
Tech Stack:PrometheusGrafanaAlertmanagerNautobotNetBoxBigPandaSaltAnsiblePythonGoServiceNowJiraITILSNMPSyslogLinuxTCPUDPIPv4IPv6

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor