Site Reliability Engineer - Hardware Infrastructure

NVIDIA
Santa Clara
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 8+ yearsEducation: bachelorsSkills: []

Site Reliability Engineer role focused on defining and supporting large-scale production systems, blending software and systems engineering to ensure high availability. You will lead incident response, postmortems, reliability metrics, and on-call practices; drive monitoring, automation, and AI-assisted tooling; and guide teams toward fault-tolerant, scalable infrastructure.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
6 months ago

Site Reliability Engineer - Hardware Infrastructure

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

Site Reliability Engineer role focused on defining and supporting large-scale production systems, blending software and systems engineering to ensure high availability. You will lead incident response, postmortems, reliability metrics, and on-call practices; drive monitoring, automation, and AI-assisted tooling; and guide teams toward fault-tolerant, scalable infrastructure.
Location: Santa Clara
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Sr. Manager level

Key Responsibilities

  • •Develop and support guidelines for incident management, planned maintenance, and blameless postmortems.
  • •Assist teams in responding to high severity incidents, driving root cause analysis, crafting high-quality postmortems, and developing post-incident corrective actions.
  • •Define reliability and supportability metrics, Service Level Objectives, and error budgets.
  • •Develop and drive the adoption of actionable, customer-centric monitoring and alerting.
  • •Apply automation and Generative AI/Agentic solutions to minimize manual and tedious activities and boost customer support.

Key Requirements

  • •Degree in Computer Science or a related technical field involving coding, or equivalent experience.
  • •8+ years of experience in SRE, DevOps, or Production Engineering.
  • •Strong understanding of SRE principles, including incident management, error budgets, SLOs, and SLAs.
  • •Experience running critical services in production and with infrastructure automation.
  • •Experience with Python, Go, Perl, or Ruby and hands-on experience with observability platforms (Prometheus, Grafana).
Experience:8+ yearsSREDevOpsProduction Engineering
Education:Bachelor's
Languages:English
Tech Stack:PythonGoPerlRubyPrometheusGrafana

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor