Head of Platform Product Reliability

Etched
San Jose
Workplace: OnsiteFull timeFunction: Administration & Executive AssistanceExperience: 10+ yearsSkills: ["Technical judgment","Cross-functional communication","Stakeholder influence","Root-cause analysis","Program leadership"]

Lead end-to-end reliability engineering for AI server, accelerator platform, rack, and datacenter infrastructure. Own reliability strategy from design requirements through fleet deployment, including EVT/DVT/PVT gates, ALT/AST, environmental and HALT/HASS programs, root-cause investigations, and system reliability modeling (MTBF, FIT, Weibull). Partner cross-functionally and with contract manufacturers/ODM/JDMs to enforce reliability commitments, build fleet telemetry visibility, and drive product release readiness.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Etched
Etched
4 months ago

Head of Platform Product Reliability

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 18 hours agoStatus: Live

Job Summary

Lead end-to-end reliability engineering for AI server, accelerator platform, rack, and datacenter infrastructure. Own reliability strategy from design requirements through fleet deployment, including EVT/DVT/PVT gates, ALT/AST, environmental and HALT/HASS programs, root-cause investigations, and system reliability modeling (MTBF, FIT, Weibull). Partner cross-functionally and with contract manufacturers/ODM/JDMs to enforce reliability commitments, build fleet telemetry visibility, and drive product release readiness.
Location: San Jose
Workplace: Onsite
Employment Type: Full time
Job Function: Administration & Executive Assistance
Seniority: Sr. Director level

Key Responsibilities

  • •Define and own end-to-end reliability strategy for AI servers, accelerator platforms, rack systems, and datacenter infrastructure from requirements through field deployment.
  • •Establish reliability requirements, qualification standards, and validation methodologies across product generations.
  • •Institutionalize full lifecycle reliability engineering processes including qualification gates, accelerated life/stress testing, environmental testing, HALT/HASS, and stress testing (vibration/shock/transport/power/thermal/soak).
  • •Lead root-cause investigations and drive corrective actions across hardware, firmware, thermal, and mechanical domains.
  • •Build fleet reliability infrastructure (telemetry analysis, field feedback loops, monitoring frameworks) and lead reliability signoff and release readiness reviews.

Pay and Benefits

Perks:Health InsuranceDentalVisionHousing SubsidyRelocationMeal AllowanceWellness Stipend

Key Requirements

  • •BS, MS, or PhD in Electrical Engineering, Mechanical Engineering, Reliability Engineering, or a related technical field.
  • •10+ years of reliability engineering experience in hardware-centric organizations, focused on complex systems rather than component-level work.
  • •Experience leading reliability programs for AI accelerator/GPU-class compute systems, hyperscale/cloud server infrastructure, or networking/storage/rack-scale infrastructure.
  • •Deep system-level failure mechanism knowledge (thermal, power delivery, mechanical, connector/interconnect failures) and how design affects long-term field reliability.
  • •Hands-on experience with FMEA, Weibull analysis, HALT/HASS, qualification planning, failure analysis, and reliability statistics/modeling.
Experience:10+ yearsHardwareAI infrastructureDatacenterHyperscaleGPU compute
Skills:Technical judgmentCross-functional communicationStakeholder influenceRoot-cause analysisProgram leadership
Tech Stack:EVTDVTPVTALTASTHALTHASSFMEAWeibull modelingMTBFFIT rate analysisComponent deratingReliability growth trackingTelemetry analysis pipelinesFleet reliability monitoringWeibull lifetime modelingMTBF projections

Company Brief

Etched
Designs transformer‑specialized AI inference ASICs (product: Sohu) to accelerate large‑language‑model workloads, working with TSMC for fabrication and targeting energy‑efficient inference performance.
Industry: Hardware Devices
Company Size: Medium (51 to 250 employees)
Growth: Early Stage Startup
Valuation: Unicorn (USD 1B+)
Funding: Series A
Headquarters: San Jose, United States
Founded: 2022
WebsiteLinkedInGlassdoor