HPC Storage Engineer - West Coast

RunPod
United States
Workplace: RemoteFull timeUSD 180,000 - 260,000 annuallyFunction: DevOps, Cloud & InfrastructureSkills: ["Ownership","Self-starting","Continuous improvement","Collaboration","Performance analysis"]

Own and scale Runpod’s multi-region storage ecosystem—network volumes, local NVMe, and S3-compatible object storage. You’ll implement and operate distributed storage deployments, write production control-plane code, and automate runbooks. Partner with SRE, networking, and hardware teams to tune storage I/O paths and high-speed fabrics (RDMA/RoCE, MTU/jumbo frames, multipath). Drive capacity expansions, migrations, SLOs, and on-call reliability improvements.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
RunPod
RunPod
3 days ago

HPC Storage Engineer - West Coast

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 15 hours agoStatus: Live

Job Summary

Own and scale Runpod’s multi-region storage ecosystem—network volumes, local NVMe, and S3-compatible object storage. You’ll implement and operate distributed storage deployments, write production control-plane code, and automate runbooks. Partner with SRE, networking, and hardware teams to tune storage I/O paths and high-speed fabrics (RDMA/RoCE, MTU/jumbo frames, multipath). Drive capacity expansions, migrations, SLOs, and on-call reliability improvements.
Location: United States
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Own capacity, durability, availability, and performance for network volumes, local NVMe, and S3-compatible object storage.
  • •Tune end-to-end I/O paths, including device/filesystem configuration, caching/read-ahead, replication/erasure coding, and client mount behavior.
  • •Diagnose complex performance issues end to end and optimize storage traffic across high-throughput networks (including MTU/jumbo frames, multipath, and NIC/offload).
  • •Lead capacity expansions, hardware refreshes, migrations, and rebalances without customer-visible disruption.
  • •Write production code and automation for storage control-plane services, provisioning workflows, data movement, monitoring, and runbook replacement; drive observability (metrics, SLOs, alerts) and on-call follow-through.

Pay and Benefits

Salary: USD 180,000 - 260,000 annually
Equity and Bonus:Equity
Perks:Health InsuranceDentalVisionPaid LeaveEquityHome Office

Key Requirements

  • •8+ years in infrastructure, storage, or systems engineering with substantial ownership of production storage at scale.
  • •Deep practical experience with a distributed storage system (e.g., Ceph, MinIO, Lustre, GPFS/Spectrum Scale, ZFS-based systems, or comparable).
  • •Strong Linux internals and storage-stack knowledge (block layer, filesystems, NVMe, page cache, I/O schedulers, NFS/SMB, iSCSI/NVMe-oF).
  • •Experience building and/or operating S3-compatible object storage services.
  • •Proficiency writing and shipping production code (Go, Python, Rust, or similar) plus observability and performance debugging experience.
Experience:Distributed storageAI infrastructureHigh-speed networkingObservabilityS3-compatible storage
Skills:OwnershipSelf-startingContinuous improvementCollaborationPerformance analysis
Tech Stack:GoPythonRustCephMinIOLustreGPFS/Spectrum ScaleMooseFSWekaFSVASTZFSDistributed storageLinuxBlock layerFilesystemsNVMePage cacheI/O schedulersNFSSMB

Eligibility

Work Authorization:Authorization required. Sponsorship not provided.

Company Brief

RunPod
Provides on-demand GPU cloud and marketplace services for machine learning workloads, offering rentable GPU instances, scalable compute for training and inference, and tools to run ML jobs cost-effectively without long-term commitments.
Industry: Cloud Computing
Website