Senior Research-Ops & DevOps Engineer

NVIDIA
Israel
Workplace: HybridFull timeFunction: Research & Scientific (R&D)Experience: 5+ yearsEducation: bachelorsSkills: ["Problem-solving","Communication","Collaboration","Teamwork","Agile"]

Lead infrastructure and operations for NVIDIA’s Video/Multimedia Architecture & Algorithms team. Build and maintain on-prem GPU clusters and cloud bursts, develop distributed pipelines for large-scale regressions and ML workloads, manage CI/CD and development environments, and automate research workflows into reliable systems. This hybrid role requires collaboration with Architects and Algorithms Engineers and 4 days/week in the office.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
2 months ago

Senior Research-Ops & DevOps Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

Lead infrastructure and operations for NVIDIA’s Video/Multimedia Architecture & Algorithms team. Build and maintain on-prem GPU clusters and cloud bursts, develop distributed pipelines for large-scale regressions and ML workloads, manage CI/CD and development environments, and automate research workflows into reliable systems. This hybrid role requires collaboration with Architects and Algorithms Engineers and 4 days/week in the office.
Location: Israel
Workplace: Hybrid
Employment Type: Full time
Job Function: Research & Scientific (R&D)
Seniority: Sr. Manager level

Key Responsibilities

  • •Work closely with Architects and Algorithms Engineers to transform one-off research workflows into dependable, consistent, automated systems
  • •Stand up and operate the compute the group runs on — on-prem GPU clusters, cloud bursts, queues, schedulers (Slurm / Kubernetes), container images, environments
  • •Develop and build decentralized workflows for extensive regression testing and experiments across HW-simulations and ML workloads — and the dashboards that make sense of the results
  • •Lead the team’s CI/CD plus the dev environments, container images and tooling everyone in the group lives in every day
  • •Transform research workflows into reliable, repeatable, automated systems

Key Requirements

  • •B.Sc. in Computer Science or Electrical/Computer Engineering
  • •5+ years in a DevOps, SRE, MLOps, Research-Ops or platform-engineering role
  • •Strong Linux fundamentals — shell, processes, networking, filesystems, systemd, performance tools
  • •Strong hands-on experience in Python — production-quality code
  • •Hands-on experience with at least one major Cloud ecosystem (OCI, AWS, Azure, GCP) and Infrastructure as Code (Terraform, Pulumi or similar)
Experience:5+ yearsGPUCloudDevOpsMLOpsSRE
Education:Bachelor's
Skills:Problem-solvingCommunicationCollaborationTeamworkAgile
Tech Stack:LinuxPythonDockerKubernetesSlurmTerraformPulumiGitLab CIGitHub ActionsJenkinsPrometheusGrafanaOpenTelemetryCI/CDBatch pipelinesOCIAWSAzureGCP

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor