Senior Compute Platform Engineer, LSF

NVIDIA
United States
Workplace: HybridFull timeUSD 184,000 - 356,500 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 8+ yearsSkills: ["Troubleshooting","Debugging","Performance analysis","Technical escalation","Systems thinking"]

Own scheduler behavior for NVIDIA’s federated IBM Spectrum LSF compute environment spanning 15–25 cells. Lead troubleshooting for MultiCluster forwarding and scheduling latency, design cell topology as the farm scales, and work with infrastructure-as-code to encode scheduler policy safely. Partner with CAD/methodology teams on demanding HPC workloads, including large memory and interactive-vs-batch contention, while escalating only after deep internals-level diagnostics.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
3 days ago

Senior Compute Platform Engineer, LSF

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 1 hour agoStatus: Live

Job Summary

Own scheduler behavior for NVIDIA’s federated IBM Spectrum LSF compute environment spanning 15–25 cells. Lead troubleshooting for MultiCluster forwarding and scheduling latency, design cell topology as the farm scales, and work with infrastructure-as-code to encode scheduler policy safely. Partner with CAD/methodology teams on demanding HPC workloads, including large memory and interactive-vs-batch contention, while escalating only after deep internals-level diagnostics.
Location: United States
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Own scheduler behavior across 15–25 federated LSF cells, including mbatchd and mbschd tuning and scheduling cycle analysis.
  • •Diagnose MultiCluster forwarding issues such as remote queue sizing, forwarding policy, and cross-cluster pend behavior.
  • •Set the technical design for cell topology and federation as the compute farm grows.
  • •Encode scheduler policy into a config schema in partnership with an IaC engineer for reliable behavior at scale.
  • •Partner with CAD and methodology teams to support demanding workloads (e.g., 500GB+ memory, interactive-vs-batch contention, tape-out crunch bursts).

Pay and Benefits

Salary: USD 184,000 - 356,500 annually
Equity and Bonus:Equity

Key Requirements

  • •BS or MS in Computer Science, Computer Engineering, or equivalent experience.
  • •8+ years in HPC or large-scale batch compute, including 5+ years on IBM Spectrum LSF.
  • •Demonstrated depth in LSF internals, with ability to debug scheduler behavior beyond documentation.
  • •Hands-on MultiCluster experience in a production, multi-site environment.
  • •Strong Linux systems fundamentals, plus scripting/programming in Python, Perl, and shell.
Experience:8+ yearsHPCLarge-scale batch computeIBM Spectrum LSFMultiClusterEDASemiconductor
Skills:TroubleshootingDebuggingPerformance analysisTechnical escalationSystems thinking
Tech Stack:LSFIBM Spectrum LSFMbatchdMbschdMultiClusterMultiCluster forwardingRTMEsubEexecElimSubmit wrappersLSF APIsLinuxPythonPerlShellIaCSlurmPBSGrid Engine

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor