Senior Engineer, NCX

NVIDIA
Germany, Czech Republic, Poland, Spain
Workplace: RemoteFull timePLN 292,500 - 650,000Function: Solutions Engineering & Sales EngineeringExperience: 8+ yearsEducation: bachelorsSkills: ["Collaboration","Operational readiness","Troubleshooting"]

Drive NVIDIA Cloud Partner (NCP) Day 2 operations for large-scale NVIDIA accelerated infrastructure. Lead operational readiness efforts, continuously validate GPU/CPU/storage/network health, and build observability with monitoring, alerting, dashboards, and SLO-driven health signals. Create automation for detection, remediation, lifecycle management, and fleet lifecycle administration across Kubernetes and multi-node environments. Translate reference architectures into production runbooks, playbooks, and operational standards.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
3 days ago

Senior Engineer, NCX

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 3 hours agoStatus: Live

Job Summary

Drive NVIDIA Cloud Partner (NCP) Day 2 operations for large-scale NVIDIA accelerated infrastructure. Lead operational readiness efforts, continuously validate GPU/CPU/storage/network health, and build observability with monitoring, alerting, dashboards, and SLO-driven health signals. Create automation for detection, remediation, lifecycle management, and fleet lifecycle administration across Kubernetes and multi-node environments. Translate reference architectures into production runbooks, playbooks, and operational standards.
Location: Germany, Czech Republic, Poland, Spain
Workplace: Remote
Employment Type: Full time
Job Function: Solutions Engineering & Sales Engineering
Seniority: Mid level

Key Responsibilities

  • •Lead NCP Day 2 operational readiness efforts by defining systems, procedures, automation, and operational methods for managing accelerated infrastructure after initial deployment.
  • •Continuously validate infrastructure health across GPU, CPU, storage, and network to identify degraded conditions before they impact critical training or inference workloads.
  • •Establish observability and operational telemetry including monitoring, alerting, dashboards, and operational signals across compute, GPUs, networking, storage, Kubernetes, and AI workloads.
  • •Build automated detection and remediation workflows to detect, isolate, drain, repair, validate, and return unhealthy infrastructure with minimal disruption to customer workloads.
  • •Operationalize NVIDIA reference architectures by translating NCP requirements into production runbooks, automation, validation criteria, health signals, and reusable operational frameworks.

Pay and Benefits

Salary: PLN 292,500 - 650,000

Key Requirements

  • •BS, MS, or Ph.D. in Computer Science, Computer/Electrical Engineering, or a related technical field, or equivalent experience.
  • •8+ years in infrastructure engineering, Site Reliability Engineering, DevOps, cloud platform engineering, systems engineering, or similar roles supporting large-scale production environments.
  • •Strong experience operating Linux-based distributed systems and cloud infrastructure in production.
  • •Deep understanding of Kubernetes, containers, cluster scheduling, and the operational lifecycle of multi-node environments.
  • •Strong observability and automation experience covering metrics/logging/alerting plus lifecycle management, remediation, upgrades, and configuration management.
Experience:8+ yearsInfrastructure engineeringSite reliability engineeringDevopsCloud platform engineeringSystems engineering
Education:Bachelor's in Computer Science, Computer/Electrical Engineering
Skills:CollaborationOperational readinessTroubleshooting
Tech Stack:LinuxKubernetesContainersPythonGoShell scriptingPrometheusGrafanaOpenTelemetryAlertmanagerInfiniBandRoCECUDANVLinkNVSwitchGPU OperatorNetwork OperatorKubernetes node maintenanceOS patching

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor