Site Reliability Engineer - Data Center

X AI
Memphis
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 2+ yearsSkills: ["Problem-solving","Prioritization","Communication"]

Serve as a hardware-focused Site Reliability Engineer for data center operations, acting as an expert on firmware and hardware specifications. You’ll analyze firmware packages for compatibility, performance, and reliability, run security/Vulnerability (CVE) scanning, and proactively flag safety issues. Diagnose and prove hardware failures (including intermittent “grey” failures), manage RMA/vendor resolution, collaborate with operations technicians, build monitoring tooling, and document RCAs and reliability models during on-call incidents.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
X AI
X AI
1 day ago

Site Reliability Engineer - Data Center

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 17 hours agoStatus: Live

Job Summary

Serve as a hardware-focused Site Reliability Engineer for data center operations, acting as an expert on firmware and hardware specifications. You’ll analyze firmware packages for compatibility, performance, and reliability, run security/Vulnerability (CVE) scanning, and proactively flag safety issues. Diagnose and prove hardware failures (including intermittent “grey” failures), manage RMA/vendor resolution, collaborate with operations technicians, build monitoring tooling, and document RCAs and reliability models during on-call incidents.
Location: Memphis
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Analyze firmware packages and hardware specifications for upcoming releases to ensure compatibility, performance, and reliability in the data center environment.
  • •Run security scanning and CVE/vulnerability analysis on firmware and related components, and flag electrical/thermal/power-safety and fail-safe issues before deployment.
  • •Investigate and diagnose hardware failures, including ambiguous or intermittent “grey failures,” using rigorous testing and data analysis to prove hardware defects.
  • •Manage vendor relationships and RMA processes, initiating claims, negotiating beyond standard processes when needed, and driving vendor accountability for resolutions.
  • •Collaborate with data center operations technicians to troubleshoot, repair, and optimize hardware systems; develop monitoring tools/scripts to detect anomalies early; document RCAs and participate in on-call incident response.

Key Requirements

  • •Bachelor’s degree in Systems Engineering, Electrical Engineering, Computer Science, or related field (or equivalent experience).
  • •2+ years of hardware reliability engineering experience, preferably in high-performance computing or data center environments.
  • •Proven expertise in firmware analysis, hardware specification review, and release validation.
  • •Strong RMA process experience, including filing claims and negotiating for resolutions.
  • •Ability to diagnose and prove complex hardware failures, including grey/intermittent issues, using diagnostic tools and data analysis.
Experience:2+ yearsHigh-performance computingData centerSupercomputingAI/ML infrastructure
Education:
Skills:Problem-solvingPrioritizationCommunication
Certifications:CRECompTIA Server+
Languages:English
Tech Stack:FirmwareCC++JavaRustPythonBashCVESecurity scanning

Company Brief

X AI
Develops advanced artificial intelligence models and research aimed at building safe, general AI and understanding the fundamental nature of the universe. Focuses on large-scale AI systems, research publications, and building foundational AI capabilities.
Industry: AI & Machine Learning
Company Size: Medium (51 to 250 employees)
Growth: Early Stage Startup
Headquarters: San Francisco, United States
Founded: 2023
Website