Senior Technical Support Engineer – Slurm

NVIDIA
Tel Aviv
Workplace: OnsiteFull timeFunction: Customer Support (Service Desk)Education: bachelorsSkills: ["Analytical skills","Research skills","Written communication","Verbal communication","Technical troubleshooting"]

Own customer support for Slurm, handling production issues from initial investigation through resolution for AI and HPC clusters. Diagnose scheduler and cluster problems across Slurm daemons and dependencies including Linux, networking, storage, authentication, databases, and GPUs. Tackle complex Slurm configuration and policy features, investigate performance/scalability using logs and debugging, and guide customers with upgrades and safe recovery practices.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
1 day ago

Senior Technical Support Engineer – Slurm

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 3 hours agoStatus: Live

Job Summary

Own customer support for Slurm, handling production issues from initial investigation through resolution for AI and HPC clusters. Diagnose scheduler and cluster problems across Slurm daemons and dependencies including Linux, networking, storage, authentication, databases, and GPUs. Tackle complex Slurm configuration and policy features, investigate performance/scalability using logs and debugging, and guide customers with upgrades and safe recovery practices.
Location: Tel Aviv
Workplace: Onsite
Employment Type: Full time
Job Function: Customer Support (Service Desk)
Seniority: Mid level

Key Responsibilities

  • •Own Slurm support cases end-to-end for customers running production AI and HPC clusters.
  • •Diagnose complex issues across Slurm daemons and areas like job scheduling, node management, resource allocation, accounting, authentication, and high availability.
  • •Solve Slurm configuration and policy problems involving partitions, reservations, priorities, fair-share, QoS, backfill, preemption, GRES/TRES, cgroups, and job constraints.
  • •Investigate performance, reliability, and scalability using logs, diagnostic data, reproductions, and source-level debugging when needed.
  • •Collaborate with engineering by producing clear technical descriptions, reproducible test cases, and well-supported defect reports.

Key Requirements

  • •BS degree in Computer Science, Engineering, or related field, or equivalent experience.
  • •5+ years of hands-on experience administering and supporting Slurm in production HPC or AI environments, including business-critical outages.
  • •Expert-level understanding of Slurm architecture, daemons, configuration, scheduling behavior, accounting, resource management, and failure modes.
  • •Ability to identify sophisticated Slurm incidents independently and guide them to sound resolution.
  • •In-depth Linux system administration and troubleshooting across systemd, cgroups, authentication, networking, and database-backed services.
Experience:HPCAIHigh-performance computingProduction supportSlurm
Education:Bachelor's in Computer Science, Engineering (or related field)
Skills:Analytical skillsResearch skillsWritten communicationVerbal communicationTechnical troubleshooting
Tech Stack:SlurmSlurmctldSlurmdSlurmdbdLinuxSystemdCgroupsMUNGELuaSPANKPyxisEnrootApptainerSingularityContainersGPUsNVIDIA Base Command ManagerBright Cluster ManagerParallel storageNetworking

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor