Senior Infrastructure Automation Engineer, Compute Platform

NVIDIA
Santa Clara, Austin
Workplace: HybridFull timeUSD 184,000 - 356,500 annuallyFunction: QA, Test & Release EngineeringExperience: 6+ yearsEducation: mastersSkills: ["Ownership","Collaboration","Observability"]

Own NVIDIA’s config-as-code foundation for an EDA compute farm by designing the configuration schema for LSF cell deployment and building a pipeline that promotes scheduler configuration from merge to production with staged rollout and rollback. Eliminate configuration drift across a federated estate, implement regression testing to support scheduled LSF upgrades, and encode scheduler expertise into templates and policy using GitOps and automation tooling.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
3 days ago

Senior Infrastructure Automation Engineer, Compute Platform

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 1 hour agoStatus: Live

Job Summary

Own NVIDIA’s config-as-code foundation for an EDA compute farm by designing the configuration schema for LSF cell deployment and building a pipeline that promotes scheduler configuration from merge to production with staged rollout and rollback. Eliminate configuration drift across a federated estate, implement regression testing to support scheduled LSF upgrades, and encode scheduler expertise into templates and policy using GitOps and automation tooling.
Location: Santa Clara, Austin
Workplace: Hybrid
Employment Type: Full time
Job Function: QA, Test & Release Engineering
Seniority: Mid level

Key Responsibilities

  • •Design and own the configuration schema for LSF cell deployment so policy changes are written once, reviewed, tested, and applied identically.
  • •Build the deployment pipeline to move scheduler configuration from merge to production across a federated estate, including staged rollout and rollback.
  • •Eliminate configuration drift across scheduler cells and build tooling to keep drift eliminated.
  • •Stand up a regression suite that enables scheduled LSF upgrades.
  • •Collaborate with an LSF internals engineer to encode scheduler expertise into repository templates and policy.

Pay and Benefits

Salary: USD 184,000 - 356,500 annually
Equity and Bonus:Equity
Perks:Equity

Key Requirements

  • •BS or MS in Computer Science or equivalent experience.
  • •6+ years in infrastructure engineering with strong config management depth using Ansible, Salt, Puppet, Chef, or comparable tools.
  • •Real experience with GitOps at scale, including review workflow, environment promotion, drift detection, and safe rollback.
  • •Proficiency in Go, Python, and shell, and comfort building infrastructure tooling.
  • •Experience automating stateful, long-lived infrastructure that cannot be simply destroyed and recreated.
Experience:6+ yearsInfrastructure automationConfig managementGitOpsEDA computeScheduler automation
Education:Master's in Computer Science
Skills:OwnershipCollaborationObservability
Tech Stack:LSFSlurmAnsibleSaltPuppetChefGitOpsGitGoPythonShellCI/CDRollbackDrift detection

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor