Senior Storage Production Engineer - DGX Cloud

NVIDIA
United States
Workplace: HybridFull timeUSD 176,000 - 333,500 annuallyFunction: Manufacturing & Production OperationsExperience: 8+ yearsEducation: bachelorsSkills: ["Communication","Teamwork","Work ethic","Quality","Problem-solving"]

Design, implement, and run large-scale storage clusters that power GPU cloud services, focusing on reliability, uptime, latency, and performance. Build storage monitoring, alerting, and automation to proactively detect and remediate issues, and optimize storage efficiency via compression, deduplication, tiering, and intelligent data placement. Partner with teams supporting AI/ML workloads and participate in incident response/on-call to continuously improve storage service lifecycles.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
1 month ago

Senior Storage Production Engineer - DGX Cloud

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 3 hours agoStatus: Live

Job Summary

Design, implement, and run large-scale storage clusters that power GPU cloud services, focusing on reliability, uptime, latency, and performance. Build storage monitoring, alerting, and automation to proactively detect and remediate issues, and optimize storage efficiency via compression, deduplication, tiering, and intelligent data placement. Partner with teams supporting AI/ML workloads and participate in incident response/on-call to continuously improve storage service lifecycles.
Location: United States
Workplace: Hybrid
Employment Type: Full time
Job Function: Manufacturing & Production Operations
Seniority: Mid level

Key Responsibilities

  • •Design, implement, and support large-scale storage clusters with scalability, high availability, and data integrity.
  • •Develop and maintain storage monitoring, logging, and alerting systems for proactive performance issue detection and resolution.
  • •Improve storage architectures for low-latency access and high-throughput performance for AI/ML workloads.
  • •Manage the storage service lifecycle from design through deployment and continuous optimization, including capacity management and launch reviews.
  • •Operate and optimize production storage infrastructure (availability, latency, system health) and support incident response/on-call with blameless root-cause analysis.

Pay and Benefits

Salary: USD 176,000 - 333,500 annually
Equity and Bonus:Equity

Key Requirements

  • •BS degree or equivalent experience with 8+ years of practical experience in Computer Science, Storage Systems, or a related technical field.
  • •Experience with distributed, high-performance storage solutions including clustered/parallel file systems and distributed object or enterprise-grade storage.
  • •Strong understanding of block, file, and object storage technologies with scalability, reliability, and performance knowledge.
  • •Experience with storage networking protocols such as NFS, SMB, iSCSI, S3, Fibre Channel, RDMA, and NVMe over Fabrics.
  • •Hands-on Linux-based storage systems engineering with strong programming and automation experience (e.g., C/C++, Java, Python, Go, NodeJS, Bash) plus config management/observability tools (e.g., Ansible/Chef/Puppet/Terraform; InfluxDB/Prometheus/Grafana/Elastic).
Experience:8+ yearsAI/MLHPCCloudDistributed systems
Education:Bachelor's
Skills:CommunicationTeamworkWork ethicQualityProblem-solving
Tech Stack:Storage systemsDistributed storageClustered file systemsParallel file systemsDistributed object storageEnterprise storageBlock storageFile storageObject storageNFSSMBISCSIS3Fibre ChannelRDMANVMe over FabricsLinuxCC++Java

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor