Senior Compute Platform Engineer, LSF - EDA Infrastructure

NVIDIA
Austin
Workplace: HybridFull timeUSD 184,000 - 356,500 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 8+ yearsEducation: bachelorsSkills: ["Escalation handling","Deep debugging","Scheduling cycle analysis","Systems fundamentals"]

Own scheduler behavior for NVIDIA’s federated IBM Spectrum LSF compute environment, tuning mbatchd/mbschd and analyzing scheduling cycles as workloads and cells scale. Diagnose MultiCluster forwarding issues impacting user-perceived farm performance. Design cell topology/federation, collaborate on scheduler policy encoded as infrastructure-as-code, and partner with CAD/methodology teams on demanding EDA workloads in high-scale HPC environments.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
6 days ago

Senior Compute Platform Engineer, LSF - EDA Infrastructure

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 14 hours agoStatus: Live

Job Summary

Own scheduler behavior for NVIDIA’s federated IBM Spectrum LSF compute environment, tuning mbatchd/mbschd and analyzing scheduling cycles as workloads and cells scale. Diagnose MultiCluster forwarding issues impacting user-perceived farm performance. Design cell topology/federation, collaborate on scheduler policy encoded as infrastructure-as-code, and partner with CAD/methodology teams on demanding EDA workloads in high-scale HPC environments.
Location: Austin
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Own scheduler behavior across 15–25 federated LSF cells, including mbatchd/mbschd tuning and scheduling cycle analysis.
  • •Diagnose MultiCluster forwarding problems, including remote queue sizing, forwarding policy, and cross-cluster pending behavior.
  • •Set technical design for cell topology and federation as the farm grows, deciding what belongs in a cell vs. a new one.
  • •Work with the IaC engineer to encode scheduler policy into a config schema that remains functional under MultiCluster scale.
  • •Partner with CAD and methodology teams on high-demand EDA workloads, including very large memory jobs and interactive vs. batch contention.

Pay and Benefits

Salary: USD 184,000 - 356,500 annually
Equity and Bonus:Equity

Key Requirements

  • •BS or MS in Computer Science, Computer Engineering, or equivalent experience.
  • •8+ years in HPC or large-scale batch compute, including 5+ years with IBM Spectrum LSF.
  • •Demonstrated depth in LSF internals, including debugging beyond documentation and explaining a scheduling cycle end-to-end.
  • •Hands-on MultiCluster experience in a production, multi-site environment.
  • •Strong Linux systems fundamentals and scripting in Python, Perl, and shell.
Experience:8+ yearsHPCLarge-scale batch computeIBM Spectrum LSFMultiClusterSemiconductorEDA
Education:Bachelor's
Skills:Escalation handlingDeep debuggingScheduling cycle analysisSystems fundamentals
Tech Stack:IBM Spectrum LSFLSFMbatchdMbschdMultiClusterLinuxPythonPerlShellIaCSlurmPBSGrid EngineEsubEexecElimRTMLSF APIs

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor