Principal Software Engineer, At-Scale Reliability and Fleet Intelligence — CSP Engagements

NVIDIA
United States
Workplace: OnsiteFull timeUSD 272,000 - 431,250 annuallyFunction: Software EngineeringExperience: 15+ yearsEducation: bachelorsSkills: []

Lead CSP-facing reliability initiatives for fleet-scale NVIDIA platforms, driving MTBI measurement, failure classification, and health monitoring architecture. Analyze CSP telemetry to identify cross-customer failure patterns, influence firmware/driver/hardware teams, and validate reliability improvements in real customer environments.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
2 months ago

Principal Software Engineer, At-Scale Reliability and Fleet Intelligence — CSP Engagements

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 12 hours agoStatus: Live

Job Summary

Lead CSP-facing reliability initiatives for fleet-scale NVIDIA platforms, driving MTBI measurement, failure classification, and health monitoring architecture. Analyze CSP telemetry to identify cross-customer failure patterns, influence firmware/driver/hardware teams, and validate reliability improvements in real customer environments.
Location: United States
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Director level

Key Responsibilities

  • •Drive reliability work streams with CSP engineering teams — ensuring shared understanding of MTBI measurement methodology, failure classification, and health monitoring architecture.
  • •Gather and synthesize CSP fleet reliability data — identify failure patterns that appear across multiple customers and champion improvements back into NVIDIA's firmware, driver, and hardware teams.
  • •Define consistent MTBI measurement methodology that works across different CSP monitoring environments and operational practices.
  • •Conduct fleet-scale failure pattern analysis using statistical methods (Pareto, survival analysis, Weibull) to classify failures as systemic, environmental, or configuration-specific.
  • •Drive fleet health monitoring integration architecture — ensure NVIDIA's health agents, telemetry, and reporting align with CSP operational workflows and automation.

Pay and Benefits

Salary: USD 272,000 - 431,250 annually
Equity and Bonus:Equity
Perks:Equity

Key Requirements

  • •15+ years of experience in systems software at datacenter scale, or reliability engineering with focus on at-scale challenges.
  • •BS or MS in Computer Science, Electrical Engineering, Statistics, or related field (or equivalent experience).
  • •Deep expertise in multi-NUMA, rack-scale system software and firmware; MTBF/MTBI calculation, Pareto analysis, root cause classification.
  • •Experience with fleet-level telemetry and observability systems: time-series databases, anomaly detection, health scoring, event correlation.
  • •Understanding of hardware failure modes in large-scale GPU/accelerator deployments and prioritization across compute, interconnect, memory, power, and thermal domains.
Experience:15+ yearsDatacenterHyperscaleGPUAccelerator
Education:Bachelor's

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor