Principal SRE - AI Inference

Cerebras
Sunnyvale
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 15+ yearsSkills: ["Technical judgment","Cross-functional influence","Mentorship","Program leadership","Clear communication"]

Define and drive the technical architecture for scaling an AI inference fleet through self-service delivery, shared observability, capacity orchestration, rollout safety, and operational automation. Lead the transition to a unified capacity management and production control plane covering reliable capacity planning, workload placement, validation, and operational decision-making across large-scale inference infrastructure. Mentor senior SREs, prioritize high-leverage automation, and improve reliability using SLO/SLI practices and incident-response learnings.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cerebras
Cerebras
2 months ago

Principal SRE - AI Inference

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 16 hours agoStatus: Live

Job Summary

Define and drive the technical architecture for scaling an AI inference fleet through self-service delivery, shared observability, capacity orchestration, rollout safety, and operational automation. Lead the transition to a unified capacity management and production control plane covering reliable capacity planning, workload placement, validation, and operational decision-making across large-scale inference infrastructure. Mentor senior SREs, prioritize high-leverage automation, and improve reliability using SLO/SLI practices and incident-response learnings.
Location: Sunnyvale
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Define and implement a strategy for delivering and running software reliably and at scale across multiple datacenters and cloud-based solutions.
  • •Architect self-service platforms and internal tooling for safe triggering and observation of critical workflows with minimal handoffs.
  • •Define and evolve reliability practices for inference workloads, including SLOs/SLIs, error budgets, blameless postmortems, chaos testing, and capacity forecasting across multi-datacenter and on-prem environments.
  • •Mentor senior SREs, support critical incident escalations, and prioritize high-leverage automation based on production pain points.
  • •Measure and drive impact using metrics such as toil reduction, deployment velocity, SLO compliance, MTTR, and adoption of self-service workflows.

Key Requirements

  • •15+ years in SRE, infrastructure engineering, or platform engineering with a record of setting technical direction and delivering reliability improvements at large scale in demanding production environments.
  • •Deep experience with large-scale compute fleets, internal control planes, schedulers, orchestration systems, capacity management, and reliability automation.
  • •Experience defining cross-team architecture for production control planes, capacity orchestration, fleet management, or self-service infrastructure with clear operational ownership.
  • •Strong ability to converge fragmented workflows, tools, and teams into coherent architectures that improve reliability, efficiency, and operational leverage.
  • •Hands-on experience with production observability, incident response, and SLO-based reliability management across metrics, logs, traces, alerting, dashboards, and review loops.
Experience:15+ yearsAI inferenceFAANGHyperscalerFrontier AICompute fleetsPlatform engineering
Skills:Technical judgmentCross-functional influenceMentorshipProgram leadershipClear communication
Tech Stack:BazelSLOSLIObservabilityMetricsLogsTraces

Company Brief

Cerebras
Designs and builds wafer-scale AI accelerators and systems for large-scale deep learning workloads, delivering specialized hardware and software to accelerate model training and inference for enterprises and research institutions.
Industry: Hardware Devices
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Headquarters: Sunnyvale, United States
Founded: 2016
WebsiteLinkedIn