NCX Senior Engineer

NVIDIA
Santa Clara, Seattle
Full timeUSD 184,000 - 356,500 annuallyFunction: Solutions Engineering & Sales EngineeringEducation: bachelorsSkills: ["Operational readiness","Automation","Troubleshooting","Collaboration","Hands-on execution"]

Join the DSX team to help NVIDIA Cloud Partners mature from initial cluster deployment into advanced Day 2 operations. Lead operational readiness, continuous infrastructure validation, and production observability across compute, GPU, networking (InfiniBand/RoCE), storage, Kubernetes, and AI workloads. Build automated detection and remediation workflows, refine GPU fleet lifecycle administration, and operationalize NVIDIA reference architectures into repeatable runbooks, SLOs, and measurable reliability standards.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
3 days ago

NCX Senior Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 12 hours agoStatus: Live

Job Summary

Join the DSX team to help NVIDIA Cloud Partners mature from initial cluster deployment into advanced Day 2 operations. Lead operational readiness, continuous infrastructure validation, and production observability across compute, GPU, networking (InfiniBand/RoCE), storage, Kubernetes, and AI workloads. Build automated detection and remediation workflows, refine GPU fleet lifecycle administration, and operationalize NVIDIA reference architectures into repeatable runbooks, SLOs, and measurable reliability standards.
Location: Santa Clara, Seattle
Employment Type: Full time
Job Function: Solutions Engineering & Sales Engineering
Seniority: Mid level

Key Responsibilities

  • •Lead NCP Day 2 operational readiness by setting up systems, procedures, automation, and operational methods for consistent management after deployment.
  • •Build continuous validation to detect GPU/CPU/storage/network degradation before it impacts training or inference workloads.
  • •Establish observability and telemetry including monitoring, alerting, dashboards, and operational signals across compute, GPU, networking, storage, Kubernetes, and AI workloads.
  • •Develop automated workflows to detect, isolate, drain, repair, validate, and return unhealthy infrastructure while minimizing disruption to customer workloads.
  • •Operationalize NVIDIA reference architectures into production operating practices, runbooks, automation, and measurable health signals/SLOs for infrastructure reliability readiness.

Pay and Benefits

Salary: USD 184,000 - 356,500 annually
Perks:Equity

Key Requirements

  • •8+ years in infrastructure engineering, Site Reliability Engineering, DevOps, or cloud/system engineering supporting large-scale production environments.
  • •Strong experience operating Linux-based distributed systems and cloud infrastructure in production.
  • •Deep understanding of Kubernetes, containers, cluster scheduling, and operational lifecycle of large multi-node environments.
  • •Strong production observability experience: metrics, logging, alerting, dashboards, and health checks guided by SLAs.
  • •Programming/automation experience with Python, Go, or shell scripting, plus automation for infrastructure lifecycle management and remediation.
Experience:Distributed systemsCloud infrastructureKubernetesDevOpsSite reliability engineeringProduction operationsAI trainingAI inferenceLarge-scale infrastructure
Education:Bachelor's
Skills:Operational readinessAutomationTroubleshootingCollaborationHands-on execution
Tech Stack:LinuxKubernetesContainersPythonGoShell scriptingPrometheusGrafanaOpenTelemetryAlertmanagerInfiniBandRoCECUDANVLinkNVSwitchDGXHGXGPU OperatorNetwork OperatorAI workloads

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor