Senior Site Reliability Engineering, Storage

NVIDIA
Bengaluru
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureEducation: bachelorsSkills: ["Communication","Collaboration","Mentoring","Troubleshooting","Root cause analysis"]

Own the reliability, performance, and scalability of global NAS, SAN, and object storage platforms. Lead design, deployment, and production operations while gathering requirements, architecting storage solutions, and driving end-to-end implementation. Improve automation for provisioning, monitoring, incident response, and lifecycle management, including on-call troubleshooting and root-cause analysis. Define SLOs/SLIs and error budgets using observability and analytics, maintain runbooks and documentation, forecast capacity, and collaborate across SRE, infrastructure, networking, and applications while mentoring engineers.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
1 week ago

Senior Site Reliability Engineering, Storage

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 11 minutes agoStatus: Live
Reposted: similar role first listed 3 months ago

Job Summary

Own the reliability, performance, and scalability of global NAS, SAN, and object storage platforms. Lead design, deployment, and production operations while gathering requirements, architecting storage solutions, and driving end-to-end implementation. Improve automation for provisioning, monitoring, incident response, and lifecycle management, including on-call troubleshooting and root-cause analysis. Define SLOs/SLIs and error budgets using observability and analytics, maintain runbooks and documentation, forecast capacity, and collaborate across SRE, infrastructure, networking, and applications while mentoring engineers.
Location: Bengaluru
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Sr. Manager level

Key Responsibilities

  • •Lead design, deployment, and operations of production NAS, SAN, and object storage platforms to ensure reliability, performance, and security.
  • •Capture requirements, architect storage solutions, and drive end-to-end implementation for new and existing services.
  • •Develop, maintain, and improve automation for provisioning, configuration, monitoring, incident response, and lifecycle management.
  • •Participate in on-call and incident response; lead troubleshooting and drive root cause analysis and preventive actions.
  • •Define and track SLOs/SLIs and error budgets using observability and analytics; build runbooks, documentation, and capacity forecasting while collaborating across teams and mentoring engineers.

Key Requirements

  • •12+ years of experience in Site Reliability, DevOps, or Infrastructure Engineering with significant focus on storage systems.
  • •Bachelor’s degree in Computer Science/Engineering or equivalent practical experience.
  • •Hands-on experience designing, deploying, and operating enterprise-grade NAS, SAN, and/or object storage platforms.
  • •Strong SRE fundamentals including SLOs/SLIs, error budgets, incident management, observability, and postmortems.
  • •Proficiency with Infrastructure as Code and configuration management tools (e.g., Terraform, Ansible, Puppet, SaltStack) plus source control and scripting/programming skills (Python/Go/Shell).
Experience:Storage systemsDevOpsInfrastructureSREAutomationNASSAN
Education:Bachelor's in Computer Science, Computer Engineering, or related technical field
Skills:CommunicationCollaborationMentoringTroubleshootingRoot cause analysis
Tech Stack:TerraformAnsiblePuppetSaltStackPythonGoShellDockerKubernetesHypervisorsCI/CDSource control

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor