Hardware Failure Analysis Engineer

X AI
Memphis
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningExperience: 2+ yearsSkills: ["Communication","Problem-solving","Prioritization","Collaboration","Data-driven reliability thinking"]

Serve as a hardware failure analysis expert for a data center environment, combining firmware analysis, hardware specification review, and rigorous diagnostics to improve reliability. Proactively scan firmware for security vulnerabilities, investigate and prove complex (including grey/intermittent) failures as true hardware defects, manage RMA processes with vendors, and collaborate with data center operations technicians. Build monitoring automation and document RCAs and reliability models while participating in on-call incident response.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
X AI
X AI
2 days ago

Hardware Failure Analysis Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 6 hours agoStatus: Live

Job Summary

Serve as a hardware failure analysis expert for a data center environment, combining firmware analysis, hardware specification review, and rigorous diagnostics to improve reliability. Proactively scan firmware for security vulnerabilities, investigate and prove complex (including grey/intermittent) failures as true hardware defects, manage RMA processes with vendors, and collaborate with data center operations technicians. Build monitoring automation and document RCAs and reliability models while participating in on-call incident response.
Location: Memphis
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Mid level

Key Responsibilities

  • •Analyze firmware packages and hardware specifications for compatibility, performance, and reliability for data center releases.
  • •Run security scanning and CVE/vulnerability analysis on firmware and related components, and flag safety issues before deployment.
  • •Investigate and diagnose hardware failures, proving ambiguous/intermittent issues as true hardware defects through rigorous testing and data analysis.
  • •Manage vendor relationships and RMA claims, including negotiations and driving resolutions when needed.
  • •Develop monitoring tools/scripts to detect hardware anomalies early, document RCAs and reliability models, and support on-call incident response.

Key Requirements

  • •Bachelor’s degree in Systems Engineering, Electrical Engineering, Computer Science, or related field (or equivalent experience).
  • •2+ years of experience in hardware reliability engineering, preferably in high-performance computing or data center environments.
  • •Proven expertise in firmware analysis, hardware specifications review, and release validation.
  • •Strong RMA experience, including filing claims, vendor negotiations, and pushing for resolutions beyond standard protocols.
  • •Ability to diagnose and prove complex hardware failures, including grey/intermittent issues, using tools/diagnostic software.
Experience:2+ yearsHigh-performance computingData centerAI/ML infrastructureSupercomputingStartupsTechnology
Education:
Skills:CommunicationProblem-solvingPrioritizationCollaborationData-driven reliability thinking
Certifications:CRECompTIA Server+
Languages:English
Tech Stack:FirmwareSecurity scanningCVEPythonBashCC++JavaRustLogic analyzersDiagnostic softwareRMA

Company Brief

X AI
Develops advanced artificial intelligence models and research aimed at building safe, general AI and understanding the fundamental nature of the universe. Focuses on large-scale AI systems, research publications, and building foundational AI capabilities.
Industry: AI & Machine Learning
Company Size: Medium (51 to 250 employees)
Growth: Early Stage Startup
Headquarters: San Francisco, United States
Founded: 2023
Website