Senior HPC and LSF Operations Engineer

NVIDIA
Santa Clara
Workplace: HybridFull timeUSD 152,000 - 241,500 annuallyFunction: Solutions Engineering & Sales EngineeringExperience: 5+ yearsEducation: bachelorsSkills: ["Communication","Problem-solving","Team collaboration"]

Senior HPC and LSF Operations Engineer overseeing large-scale Linux-based compute infrastructure. You will manage and optimize job scheduling systems (LSF/Slurm) across multiple sites, drive efficiency through automation, improve observability and reliability, and partner with customer teams to translate requirements into measurable performance metrics.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
4 months ago

Senior HPC and LSF Operations Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 5 hours agoStatus: Live

Job Summary

Senior HPC and LSF Operations Engineer overseeing large-scale Linux-based compute infrastructure. You will manage and optimize job scheduling systems (LSF/Slurm) across multiple sites, drive efficiency through automation, improve observability and reliability, and partner with customer teams to translate requirements into measurable performance metrics.
Location: Santa Clara
Workplace: Hybrid
Employment Type: Full time
Job Function: Solutions Engineering & Sales Engineering
Seniority: Sr. Manager level

Key Responsibilities

  • •Manage, scale, and optimize job scheduling systems (LSF, Slurm, etc.) in a large-scale, multi-site environment supporting EDA and other compute-intensive workloads
  • •Analyze scheduler and infrastructure performance data to identify systemic bottlenecks and drive measurable improvements in utilization, throughput, and turnaround time
  • •Lead problem solving across scheduler, OS, and workload layers, ensuring timely resolution of service-impacting issues
  • •Identify recurring operational challenges and implement targeted automation or process improvements to reduce manual effort and prevent repeat incidents
  • •Help define and track reliable metrics and SLOs for service performance and reliability, partnering with customers to ensure expectations are realistic and measurable

Pay and Benefits

Salary: USD 152,000 - 241,500 annually
Equity and Bonus:Equity
Perks:Equity

Key Requirements

  • •Bachelor's degree in Computer Science or related field, or equivalent experience
  • •Minimum 5+ years of experience operating and supporting large-scale Linux-based compute infrastructure
  • •Strong hands-on experience supporting and tuning job scheduling systems (LSF, Slurm, etc.) in HPC or silicon design environments
  • •Proficiency in Linux systems administration (CentOS/RHEL)
  • •Strong problem solving skills and the ability to independently analyze complex system behavior under load
Experience:5+ yearsHPCLinuxLSFSlurmEDAComputing infrastructure
Education:Bachelor's in Computer Science
Skills:CommunicationProblem-solvingTeam collaboration
Languages:English
Tech Stack:LSFSlurmLinuxCentOSRHELDockerSingularityPodman

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor