Senior Storage Production Engineer - DGX Cloud

NVIDIA
Australia
Workplace: RemoteFull timeFunction: Manufacturing & Production OperationsExperience: 8+ yearsEducation: bachelorsSkills: ["Excellent written and oral communication","Teamwork","Strong work ethic","Commitment to quality"]

Design, implement, and support large-scale storage clusters for NVIDIA GPU cloud services, ensuring scalability, reliability, uptime, and data integrity. Build storage monitoring, logging, and alerting for proactive issue detection, improve architectures for low-latency AI/ML workloads, and optimize storage efficiency and data placement. Manage lifecycle from inception to continuous optimization using automation, capacity planning, predictive analytics, and on-call support for production availability.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
5 days ago

Senior Storage Production Engineer - DGX Cloud

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 15 hours agoStatus: Live

Job Summary

Design, implement, and support large-scale storage clusters for NVIDIA GPU cloud services, ensuring scalability, reliability, uptime, and data integrity. Build storage monitoring, logging, and alerting for proactive issue detection, improve architectures for low-latency AI/ML workloads, and optimize storage efficiency and data placement. Manage lifecycle from inception to continuous optimization using automation, capacity planning, predictive analytics, and on-call support for production availability.
Location: Australia
Workplace: Remote
Employment Type: Full time
Job Function: Manufacturing & Production Operations
Seniority: Mid level

Key Responsibilities

  • •Design, implement, and support large-scale storage clusters to ensure scalability, high availability, and data integrity.
  • •Develop and maintain storage monitoring, logging, and alerting systems for proactive detection and resolution of performance issues.
  • •Improve storage architectures for AI/ML workloads to enable low-latency access, efficient caching, and high-throughput performance.
  • •Manage the lifecycle of storage services from inception and design through deployment, operation, and continuous optimization, including launch reviews and automation frameworks.
  • •Maintain and optimize production storage infrastructure by supervising availability, latency, and system health, using predictive analytics/automation and participating in on-call rotation.

Key Requirements

  • •BS degree or equivalent experience in Computer Science, Storage Systems, or a related field with 8+ years of practical experience.
  • •Experience with distributed and high-performance storage solutions, including clustered/parallel file systems and enterprise-grade storage.
  • •Knowledge of block, file, and object storage technologies with an understanding of scalability, reliability, and performance.
  • •Experience with storage networking protocols including NFS, SMB, iSCSI, S3, Fibre Channel, RDMA, and NVMe over Fabrics.
  • •Programming and automation experience using one or more of C/C++, Java, Python, Go, NodeJS, or Bash plus infrastructure configuration management (e.g., Ansible, Chef, Puppet, Terraform) and observability tools (InfluxDB, Prometheus, Grafana, Elastic).
Experience:8+ years
Education:Bachelor's in Computer Science, Storage Systems, or a related technical field
Skills:Excellent written and oral communicationTeamworkStrong work ethicCommitment to quality
Tech Stack:KubernetesContainersVirtualizationStorage architectureHigh-performance distributed storageData managementLinuxNFSSMBISCSIS3Fibre ChannelRDMANVMe over FabricsBlock storageFile storageObject storageCC++Java

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor