Senior Server RAS Engineer

NVIDIA
Bengaluru
Workplace: OnsiteFull timeFunction: Hospitality & Food ServiceExperience: 10+ yearsSkills: ["Problem-solving","Communication","Collaboration","Detail-oriented","Self-starter"]

Senior engineer focusing on reliability, availability, and serviceability (RAS) for NVIDIA’s data center GPUs and Grace systems. You’ll design, architect, and deliver server-level RAS features, define requirements for scale-out environments, develop fault detection/recovery mechanisms, and collaborate with hardware, software, and customer teams to ensure high reliability and minimal downtime across enterprise AI infrastructure.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
6 months ago

Senior Server RAS Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 2 hours agoStatus: Live

Job Summary

Senior engineer focusing on reliability, availability, and serviceability (RAS) for NVIDIA’s data center GPUs and Grace systems. You’ll design, architect, and deliver server-level RAS features, define requirements for scale-out environments, develop fault detection/recovery mechanisms, and collaborate with hardware, software, and customer teams to ensure high reliability and minimal downtime across enterprise AI infrastructure.
Location: Bengaluru
Workplace: Onsite
Employment Type: Full time
Job Function: Hospitality & Food Service

Key Responsibilities

  • •Design, architect, and deliver server-level RAS for NVIDIA’s data center products.
  • •Define RAS requirements to ensure industry standards and customer expectations for scale-out environments.
  • •Develop fault detection, isolation, and recovery mechanisms to ensure system resilience and minimize downtime.
  • •Evaluate and select technologies to optimize reliability, availability, and serviceability (MTBF, MTTR, TCO).
  • •Collaborate with customers, vendors and suppliers to integrate their RAS-related solutions into the overall system architecture.

Key Requirements

  • •BS, MS, or PhD or equivalent experience in EE/CS or related field with demonstrated experience of 10+ years
  • •Strong Python programming in Linux environment; strong understanding of Linux kernel internals; strong code review skills
  • •Extensive knowledge in system-level architecture, reliability engineering, and fault tolerance mechanisms for complex computing systems
  • •Proficient in scale-out architectures; hands-on experience is a plus
  • •Proficiency in system-level simulation tools and methodologies (e.g., fault injection, reliability block diagrams, failure rate analysis)
Experience:10+ yearsAI computingData centerGPU
Skills:Problem-solvingCommunicationCollaborationDetail-orientedSelf-starter
Tech Stack:PythonLinuxKernelSystem-level simulationFault injection

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor