Senior Site Reliability Engineering - Storage

NVIDIA
Israel
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 12+ yearsEducation: bachelorsSkills: ["Communication","Collaboration","Mentoring","Troubleshooting","Root cause analysis"]

Own the reliability, performance, and scalability of global NAS, SAN, and Object Storage platforms powering critical internal and external services. Lead storage platform design, deployment, and operations; define SLOs/SLIs and error budgets; and drive automation for provisioning, monitoring, incident response, and lifecycle management. Partner with SRE, infrastructure, networking, and application teams in a follow-the-sun model, and mentor engineers while improving reliability through data-driven root cause analysis.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
2 days ago

Senior Site Reliability Engineering - Storage

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 12 hours agoStatus: Live

Job Summary

Own the reliability, performance, and scalability of global NAS, SAN, and Object Storage platforms powering critical internal and external services. Lead storage platform design, deployment, and operations; define SLOs/SLIs and error budgets; and drive automation for provisioning, monitoring, incident response, and lifecycle management. Partner with SRE, infrastructure, networking, and application teams in a follow-the-sun model, and mentor engineers while improving reliability through data-driven root cause analysis.
Location: Israel
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Lead design, deployment, and operations of production NAS, SAN, and Object Storage platforms to ensure reliability, performance, and security.
  • •Gather requirements from partner teams, architect storage solutions, and drive end-to-end implementation for new and existing services.
  • •Develop and continuously improve automation for provisioning, configuration, monitoring, incident response, and lifecycle management.
  • •Participate in on-call and incident response; troubleshoot complex storage/performance issues; drive root cause analysis and preventive actions.
  • •Define and track SLOs/SLIs and error budgets using observability and analytics, and build runbooks and documentation for storage services and automation.

Key Requirements

  • •12+ years of experience in Site Reliability, DevOps, or Infrastructure Engineering with significant focus on storage systems.
  • •Bachelor’s degree in Computer Science/Engineering (or equivalent practical experience).
  • •Hands-on experience designing, deploying, and operating enterprise-grade NAS, SAN, and/or Object Storage platforms.
  • •Strong SRE fundamentals including SLOs/SLIs, error budgets, incident management, observability, and postmortems.
  • •Proficiency with Infrastructure as Code and configuration management tools (e.g., Terraform, Ansible, Puppet, SaltStack) plus source control.
Experience:12+ yearsStorage systemsInfrastructure engineering
Education:Bachelor's in Computer Science, Computer Engineering, or related technical field
Skills:CommunicationCollaborationMentoringTroubleshootingRoot cause analysis
Tech Stack:TerraformAnsiblePuppetSaltStackSource controlPythonGoShellDockerKubernetesHypervisorsCI/CDObservabilityAnalyticsSLOs/SLIsIncident responsePostmortems

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor