Senior Platform and EngOps Engineer - Cluster Operations

NVIDIA
Santa Clara
Workplace: OnsiteFull timeUSD 176,000 - 276,000 annuallyFunction: Solutions Engineering & Sales EngineeringExperience: 8+ yearsEducation: bachelorsSkills: ["Ansible","Python","Shell","Linux","NVLink","InfiniBand","Slurm"]

Lead the design, deployment, and ongoing operations of large GPU clusters (NVLink/InfiniBand) in a fast-paced AI/ HPC environment. You’ll automate provisioning, maintain software/firmware updates, troubleshoot failures, and coordinate with cross-timezone engineering teams to ensure high-availability cluster performance.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
4 months ago

Senior Platform and EngOps Engineer - Cluster Operations

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 23 hours agoStatus: Live
Reposted: similar role first listed 7 months ago

Job Summary

Lead the design, deployment, and ongoing operations of large GPU clusters (NVLink/InfiniBand) in a fast-paced AI/ HPC environment. You’ll automate provisioning, maintain software/firmware updates, troubleshoot failures, and coordinate with cross-timezone engineering teams to ensure high-availability cluster performance.
Location: Santa Clara
Workplace: Onsite
Employment Type: Full time
Job Function: Solutions Engineering & Sales Engineering

Key Responsibilities

  • •Develop automated tools to deploy, provision, and maintain extensive GPU clusters interconnected via NVLink and InfiniBand.
  • •Implement modern DevOps tools to automate software updates, perform maintenance tasks, and monitor cluster availability, ensuring seamless operations.
  • •Take ownership of daily cluster failures and issues, troubleshooting them promptly to maintain optimal cluster availability and performance.
  • •Manage the rollout and rollback of cluster software and firmware updates, ensuring smooth transitions and minimal disruptions.
  • •Collaborate effectively with dynamic Engineering and Product Teams across multiple time zones to align cluster operations with evolving project requirements.

Pay and Benefits

Salary: USD 176,000 - 276,000 annually
Equity and Bonus:Equity

Key Requirements

  • •BS or MS in Computer Science, Computer Engineering, Electrical Engineering, or equivalent experience.
  • •8+ years of hands-on experience deploying and administering clusters, servers, switches, and related infrastructure.
  • •Automation expertise with hands-on skills in Ansible, Python and Shell scripting.
  • •Deep understanding of operating systems, computer networks, and high-performance applications.
  • •Proven ability to work effectively with developers and test engineers across different teams and time zones.
Experience:8+ yearsHigh-performance computingGpuArtificial intelligence
Education:Bachelor's
Skills:AnsiblePythonShellLinuxNVLinkInfiniBandSlurm
Tech Stack:AnsiblePythonShellLinuxNVLinkInfiniBandSlurm

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor