Senior Cloud Infrastructure and DevOps Solutions Architect

NVIDIA
Germany, Spain, United Kingdom
Workplace: OnsiteFull timeFunction: Solutions Engineering & Sales EngineeringExperience: 8+ yearsSkills: ["Consulting","Communication","Troubleshooting","Knowledge transfer","Executive stakeholder presentations"]

Build and advise on large-scale AI/HPC infrastructure for NVIDIA’s Cloud Partner operating model. Own end-to-end validation and Day 2 production stability for partner GPU clusters, covering monitoring, fault detection/remediation, preventive maintenance, and fleet reliability. Work across bare metal, OS, container platforms, networking, and storage to enable heterogeneous open Kubernetes-based stacks and support R&D POCs with hands-on troubleshooting and runbooks.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
2 hours ago

Senior Cloud Infrastructure and DevOps Solutions Architect

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 2 hours agoStatus: Live

Job Summary

Build and advise on large-scale AI/HPC infrastructure for NVIDIA’s Cloud Partner operating model. Own end-to-end validation and Day 2 production stability for partner GPU clusters, covering monitoring, fault detection/remediation, preventive maintenance, and fleet reliability. Work across bare metal, OS, container platforms, networking, and storage to enable heterogeneous open Kubernetes-based stacks and support R&D POCs with hands-on troubleshooting and runbooks.
Location: Germany, Spain, United Kingdom
Workplace: Onsite
Employment Type: Full time
Job Function: Solutions Engineering & Sales Engineering
Seniority: Mid level

Key Responsibilities

  • •Own full-solution validation on the partner software stack, including cluster-wide stability testing and multi-day burn-in against MTBI and goodput targets.
  • •Minimise time from cluster handover to first production workload by reducing duplicated validation and handover friction across bring-up and managed-service intake.
  • •Own Day 2 production stability at fleet scale, including monitoring/logging, workload orchestration, fault detection/remediation, preventive maintenance, and rollout campaigns.
  • •Assess customer environments and operate heterogeneous open platforms combining Kubernetes/KubeVirt/Slurm with enterprise networking/storage to enable third-party workloads.
  • •Provide consultative guidance and hands-on troubleshooting across bare metal, OS, software stack, container platform, networking, and storage; support POCs/POVs and create runbooks and best-practice guides.

Key Requirements

  • •BS/MS/PhD (or equivalent experience) in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or related fields.
  • •8+ years managing scalable cloud environments and automation engineering roles.
  • •Proven cloud/HPC/GPU expertise, including networking fundamentals and data center architectures, with hands-on experience operating HPC/AI clusters and NVIDIA GPU-accelerated infrastructure.
  • •Extensive Kubernetes experience in GPU-accelerated and HPC environments, including scheduler internals, Slurm, and mixed bare-metal/virtualized estates (e.g., KubeVirt).
  • •Deep Linux and storage systems knowledge (RedHat, Ubuntu; Lustre, GPFS, ZFS, XFS) plus automation/GitOps/observability using Python/Bash and tools like Ansible and Terraform.
Experience:8+ yearsHPCAIGPU infrastructureDevOpsCloud
Education:
Skills:ConsultingCommunicationTroubleshootingKnowledge transferExecutive stakeholder presentations
Tech Stack:KubernetesKubeVirtSlurmPrometheusGrafanaCumulusSONiCPythonBashAnsibleTerraformGitOpsLokiLinuxRedHatUbuntuLustreGPFSZFSXFS

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor