Principal Engineer, AI Inference Reliability

Cerebras
United States, Canada
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningExperience: 7+ yearsSkills: ["Communication","Cross-functional leadership","Mentoring","Incident response","Reliability focus"]

Own the reliability mission for AI inference services, spanning client SDKs, public-cloud multi-region deployments, and wafer-scale systems in specialized data centers. Define SLOs and incident-response frameworks, design reliability mechanisms for fault detection, graceful degradation, failover, throttling, and recovery, and lead postmortems and prevention loops. Build reliability dashboards and tooling for chaos testing and distributed fault injection, and mentor engineers to embed reliability across every layer of the inference stack.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cerebras
Cerebras
10 months ago

Principal Engineer, AI Inference Reliability

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 16 hours agoStatus: Live

Job Summary

Own the reliability mission for AI inference services, spanning client SDKs, public-cloud multi-region deployments, and wafer-scale systems in specialized data centers. Define SLOs and incident-response frameworks, design reliability mechanisms for fault detection, graceful degradation, failover, throttling, and recovery, and lead postmortems and prevention loops. Build reliability dashboards and tooling for chaos testing and distributed fault injection, and mentor engineers to embed reliability across every layer of the inference stack.
Location: United States, Canada
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Mid level

Key Responsibilities

  • •Define and drive reliability strategy by establishing SLOs and aligning across engineering.
  • •Design and implement reliability mechanisms for fault detection, graceful degradation, failover, throttling, and recovery across regions and data centers.
  • •Lead large-scale incident management, including postmortems, root-cause analysis, and prevention loops.
  • •Architect for reliability and observability, influencing system design for redundancy, durability, and debuggability.
  • •Develop reliability tooling for chaos testing, load simulation, and distributed fault injection; monitor and communicate reliability metrics through dashboards and alerts.

Key Requirements

  • •Bachelor's or master's degree in computer science or related field.
  • •7+ years of experience in backend, infrastructure, or reliability engineering for large-scale distributed systems.
  • •Strong programming skills in at least one backend language such as Python, C++, Go, or Rust.
  • •Deep experience with reliability principles including SLO/SLI/SLA design, incident response, and postmortem culture.
  • •Excellent communication and cross-functional leadership skills.
Experience:7+ years
Education:
Skills:CommunicationCross-functional leadershipMentoringIncident responseReliability focus
Tech Stack:PythonC++GoRust

Company Brief

Cerebras
Designs and builds wafer-scale AI accelerators and systems for large-scale deep learning workloads, delivering specialized hardware and software to accelerate model training and inference for enterprises and research institutions.
Industry: Hardware Devices
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Headquarters: Sunnyvale, United States
Founded: 2016
WebsiteLinkedIn