NCX Senior Engineer

NVIDIA
Bengaluru, Hyderabad, Pune
Workplace: OnsiteFull timeFunction: Solutions Engineering & Sales EngineeringExperience: 8+ yearsSkills: ["Collaboration","Troubleshooting","Operational readiness"]

Drive Day 2 operational readiness for NVIDIA Cloud Partner (NCP) accelerated infrastructure within the DSX team. Partner with cloud and operations teams to build systems, automation, and runbooks that ensure reliable production operations. Continuously validate GPU/CPU/storage/network health, implement observability with SLO-driven metrics, and automate detection, remediation, and fleet lifecycle processes across large Kubernetes-based AI clusters.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
3 days ago

NCX Senior Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 2 hours agoStatus: Live

Job Summary

Drive Day 2 operational readiness for NVIDIA Cloud Partner (NCP) accelerated infrastructure within the DSX team. Partner with cloud and operations teams to build systems, automation, and runbooks that ensure reliable production operations. Continuously validate GPU/CPU/storage/network health, implement observability with SLO-driven metrics, and automate detection, remediation, and fleet lifecycle processes across large Kubernetes-based AI clusters.
Location: Bengaluru, Hyderabad, Pune
Workplace: Onsite
Employment Type: Full time
Job Function: Solutions Engineering & Sales Engineering
Seniority: Mid level

Key Responsibilities

  • •Lead NCP Day 2 operational readiness efforts by collaborating with cloud partners to establish systems, procedures, automation, and operational methods.
  • •Build continuous validation for GPU, CPU, storage, and network health across large-scale AI clusters to detect degradation early.
  • •Establish observability and operational telemetry, including telemetry, monitoring, alerting, dashboards, and operational signals for compute, GPU, networking, storage, Kubernetes, and AI workloads.
  • •Develop automated detection and remediation workflows to isolate, drain, repair, validate, and return unhealthy infrastructure while minimizing customer disruption.
  • •Operationalize NVIDIA reference architectures by translating NCP requirements into production operating practices, validation criteria, runbooks, and measurable operational standards.

Key Requirements

  • •BS, MS, or Ph.D. in Computer Science, Computer/Electrical Engineering, or a related technical field, or equivalent experience.
  • •8+ years in infrastructure engineering, Site Reliability Engineering, DevOps, cloud platform engineering, systems engineering, or similar roles supporting large-scale production environments.
  • •Strong experience operating Linux-based distributed systems and cloud infrastructure in production.
  • •Deep understanding of Kubernetes, containers, cluster scheduling, and multi-node environment operational lifecycle.
  • •Experience with observability and automation for infrastructure lifecycle management, failure detection/remediation, upgrades, and configuration management.
Experience:8+ yearsDistributed systemsCloud infrastructureDevOpsSite Reliability EngineeringAI training and inference
Education:
Skills:CollaborationTroubleshootingOperational readiness
Tech Stack:LinuxKubernetesContainersPythonGoShell scriptingPrometheusGrafanaOpenTelemetryAlertmanager

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor