Senior Solutions Architect, DevOps

NVIDIA
Tel Aviv
Workplace: OnsiteFull timeFunction: Solutions Engineering & Sales EngineeringExperience: 5+ yearsSkills: ["Consulting","Communication","Troubleshooting","Documentation","Technical leadership"]

Advise and provide hands-on technical leadership for large-scale AI/HPC infrastructure deployments, spanning bare metal through the software and container stack. Partner with customers and internal teams to troubleshoot, validate architectures, and recommend production-ready Kubernetes-based platforms integrated with networking and storage. Lead account technical strategy, develop standard methodologies and runbooks, and engage in POCs/POVs to support DevOps and platform architecture decisions.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
1 day ago

Senior Solutions Architect, DevOps

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 15 hours agoStatus: Live

Job Summary

Advise and provide hands-on technical leadership for large-scale AI/HPC infrastructure deployments, spanning bare metal through the software and container stack. Partner with customers and internal teams to troubleshoot, validate architectures, and recommend production-ready Kubernetes-based platforms integrated with networking and storage. Lead account technical strategy, develop standard methodologies and runbooks, and engage in POCs/POVs to support DevOps and platform architecture decisions.
Location: Tel Aviv
Workplace: Onsite
Employment Type: Full time
Job Function: Solutions Engineering & Sales Engineering
Seniority: Mid level

Key Responsibilities

  • •Advise on and help maintain large-scale AI/HPC infrastructure, including monitoring, logging, and workload orchestration with Kubernetes and Linux job schedulers.
  • •Provide consultative guidance and hands-on troubleshooting across the stack, from bare metal and OS through containers, networking, and storage.
  • •Assess customer environments and recommend optimized, production-ready Kubernetes-based container platforms integrated with enterprise networking and storage.
  • •Develop, refine, and document standard methodologies, operational guidelines, and high-quality customer/internal documentation (runbooks, onboarding materials, best practices).
  • •Act as a technical leader for assigned customer accounts and engage in POCs/POVs to validate new features, architectures, and upgrade approaches.

Key Requirements

  • •BS/MS/PhD in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or related fields, with 5+ years of experience managing scalable cloud environments and automation engineering roles.
  • •Proven expertise in networking fundamentals and data center architectures, including hands-on experience operating and troubleshooting HPC/AI clusters.
  • •Hands-on experience deploying, configuring, and optimizing NVIDIA GPU-accelerated infrastructure, including driver management, CUDA toolkit integration, and GPU workload profiling.
  • •Extensive Kubernetes experience for container orchestration, resource scheduling/scaling, and integration with GPU-accelerated and HPC environments.
  • •Deep knowledge of Linux (RedHat, Ubuntu), observability/monitoring, and automation using Python/Bash with Infrastructure-as-Code tools (e.g., Ansible, Terraform).
Experience:5+ yearsHPCAI/MLCloud infrastructureDevOpsKubernetes
Education:
Skills:ConsultingCommunicationTroubleshootingDocumentationTechnical leadership
Tech Stack:KubernetesLinuxPythonBashAnsibleTerraformInfrastructure-as-CodeGrafanaLokiPrometheusCUDARedHatUbuntuCI/CDSLURMMPIEnrootNVIDIA Base Command Manager (BCM)InfiniBandRoCE

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor