Senior Site Reliability Engineer - Storage

NVIDIA
Santa Clara
Full timeUSD 168,000 - 333,500 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 8+ yearsEducation: bachelorsSkills: ["Communication","Collaboration","Technology evaluation","Documentation","Performance optimization"]

Design, implement, and optimize on-prem HPC storage solutions supplemented with cloud computing. Own scalable storage architectures for data-intensive applications, develop automation tools for deployment and operational monitoring/alerting, and enable self-service resource consumption. Evaluate distributed file systems and perform technology assessments, while collaborating with engineering teams to translate infrastructure requirements into reliable, well-documented best practices for high-performance environments.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
5 days ago

Senior Site Reliability Engineer - Storage

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 19 hours agoStatus: Live
Reposted: similar role first listed 7 months ago

Job Summary

Design, implement, and optimize on-prem HPC storage solutions supplemented with cloud computing. Own scalable storage architectures for data-intensive applications, develop automation tools for deployment and operational monitoring/alerting, and enable self-service resource consumption. Evaluate distributed file systems and perform technology assessments, while collaborating with engineering teams to translate infrastructure requirements into reliable, well-documented best practices for high-performance environments.
Location: Santa Clara
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Design and implement on-prem HPC infrastructure supplemented with cloud computing to support NVIDIA’s growing IT needs.
  • •Design and implement scalable, efficient storage solutions optimized for data-intensive applications and performance/cost-effectiveness.
  • •Develop automation tooling for deploying and managing large-scale infrastructure, including operational monitoring and alerting.
  • •Document distributed file system procedures and practices, and perform related technology evaluations.
  • •Collaborate across teams to understand developers’ workflows, gather infrastructure requirements, and guide methodologies for building, testing, and deploying performant applications.

Pay and Benefits

Salary: USD 168,000 - 333,500 annually
Equity and Bonus:Equity

Key Requirements

  • •BS in Computer Science (or equivalent) with 8+ years relevant experience; MS with 5+ years or Ph.D. with 3 years.
  • •8+ years solving performance bottlenecks for HPC applications.
  • •Experience designing, deploying, and managing enterprise NAS/storage solutions such as NetApp and Pure Storage, and S3-based storage such as Cloudian MinIO.
  • •Experience with one or more parallel/distributed filesystems such as Lustre and GPFS.
  • •Python/Bash/Golang scripting experience, plus strong cloud operations experience in AWS, Azure, or GCP and monitoring tools such as Prometheus+Grafana or Elasticsearch+Kibana.
  • •Experience with RDMA (InfiniBand or RoCE) and HPC cluster management tools like Slurm, PBS, or LSF is a plus.
  • •Containerization experience with Docker, Mesosphere DCOS, or Kubernetes (k8s) is a plus.
Experience:8+ yearsHPCHigh-performance computingDistributed storageCloud computingInfrastructure automation
Education:Bachelor's in Computer Science
Skills:CommunicationCollaborationTechnology evaluationDocumentationPerformance optimization
Tech Stack:PythonBashGolangAWSAzureGCPNetAppPure StorageS3Cloudian MinIOLustreGPFSPrometheusGrafanaElasticsearchKibanaSplunkZabbixDockerMesosphere DCOS

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor