Senior HPC DevOps Engineer, NCS

NVIDIA
Tel Aviv
Full timeFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsEducation: bachelorsSkills: ["Problem-solving","Technical leadership","Communication","Innovation"]

Design, implement, and operate large-scale HPC/AI clusters, advancing AI and GPU computing platforms. Build infrastructure-as-code workflows, automate deployments, and streamline CI/CD pipelines. Develop automation for networking and operations, troubleshoot performance and reliability from bare metal through applications, and support R&D with POCs/POVs. Collaborate with HPC, OS, GPU compute, and systems specialists to architect and bring up high-performance platforms.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
5 days ago

Senior HPC DevOps Engineer, NCS

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 29 minutes agoStatus: Live

Job Summary

Design, implement, and operate large-scale HPC/AI clusters, advancing AI and GPU computing platforms. Build infrastructure-as-code workflows, automate deployments, and streamline CI/CD pipelines. Develop automation for networking and operations, troubleshoot performance and reliability from bare metal through applications, and support R&D with POCs/POVs. Collaborate with HPC, OS, GPU compute, and systems specialists to architect and bring up high-performance platforms.
Location: Tel Aviv
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Design, implement, and maintain large-scale HPC/AI clusters with monitoring, logging, and alerting.
  • •Use and develop tools for infrastructure as code to enable scalable, repeatable deployments.
  • •Develop and maintain CI/CD pipelines to automate and streamline deployments.
  • •Automate deployment, configuration management, and operational monitoring, including networking automations.
  • •Troubleshoot end-to-end issues from bare metal to application level to ensure reliability and efficiency.

Key Requirements

  • •5+ years of experience in related engineering roles.
  • •B.Sc. in Computer Science, Engineering, or a related field.
  • •Advanced proficiency in programming and scripting with solid object-oriented principles.
  • •Familiarity with Jenkins and configuration management tools such as Ansible and Puppet/Chef.
  • •Deep understanding of Kubernetes/container microservice technologies, plus hands-on experience with event streaming/message queues like Apache Kafka and storage solutions such as Lustre/GPFS/ZFS/XFS.
Experience:5+ yearsHPCAIGPU computing
Education:Bachelor's in Computer Science, Engineering
Skills:Problem-solvingTechnical leadershipCommunicationInnovation
Tech Stack:KubernetesJenkinsAnsiblePuppetChefApache KafkaLustreGPFSZFSXFSVMwareHyper-VKVMCitrixAWSAzureGoogle CloudSlurmCI/CDInfrastructure as Code (IaC)

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor