Reliability Engineer

Intel
Santa Clara
Workplace: OnsiteFull timeUSD 122,440 - 232,190 annuallyFunction: Data Analytics & Business IntelligenceExperience: 4-6 yearsEducation: mastersSkills: ["Root-cause analysis","Data-driven problem-solving","Cross-functional collaboration"]

Define and own pod-level reliability specifications for AI hardware data centers, translating system/SLA needs into compute, memory, storage, network, power, and cooling targets. Lead FMEA, root-cause analysis, and fleet failure-data analytics to update specs and corrective actions. Architect RAS features and graceful degradation, partner on redundancy and disaster-recovery readiness, and establish HALT/HASS, burn-in, and qualification processes while tracking field returns and KPIs.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Intel
Intel
3 days ago

Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 14 hours agoStatus: Live

Job Summary

Define and own pod-level reliability specifications for AI hardware data centers, translating system/SLA needs into compute, memory, storage, network, power, and cooling targets. Lead FMEA, root-cause analysis, and fleet failure-data analytics to update specs and corrective actions. Architect RAS features and graceful degradation, partner on redundancy and disaster-recovery readiness, and establish HALT/HASS, burn-in, and qualification processes while tracking field returns and KPIs.
Location: Santa Clara
Workplace: Onsite
Employment Type: Full time
Job Function: Data Analytics & Business Intelligence
Seniority: Mid level

Key Responsibilities

  • •Define and maintain pod-level reliability/availability specs and targets (MTBF, AFR, RAS) across compute, memory, storage, network, power, and cooling.
  • •Translate system/SLA requirements into pod/subsystem reliability specs and flow requirements down to silicon, platform, and facilities teams.
  • •Lead FMEA, root-cause analysis, and pod fleet failure-data analytics to drive corrective actions and spec updates.
  • •Architect RAS features (ECC, memory mirroring, predictive failure, telemetry) and graceful degradation/redundancy against pod-level specs.
  • •Partner with facilities on pod power/cooling redundancy, thermal margins, and disaster-recovery readiness.

Pay and Benefits

Salary: USD 122,440 - 232,190 annually
Equity and Bonus:Equity
Perks:Health InsuranceRetirementPaid Leave

Key Requirements

  • •BS/MS/PhD in EE/ME Reliability or related, and/or 4–6 years of experience in reliability.
  • •Experience authoring and owning reliability specifications and requirement flow-down.
  • •Strong RAS, FMEA, and statistical reliability skills (Weibull, FIT).
  • •Experience with large-scale fleet telemetry and thermal/power redundancy.
  • •Experience architecting reliability features and reliability targets for complex hardware systems.
Experience:4-6 yearsAI hardwareData centerAI cluster operationsFleet telemetryReliability engineering
Education:Master's in EE/ME Reliability
Skills:Root-cause analysisData-driven problem-solvingCross-functional collaboration
Tech Stack:PythonSQLWeibullFIT

Company Brief

Intel
Designs and manufactures semiconductor chips, processors, and related hardware for PCs, data centers, networking, and embedded applications, while providing software and services to accelerate computing across industries globally.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1968
Glassdoor
Glassdoor: 3.8
WebsiteLinkedInGlassdoor