Site Reliability Engineer - Ops & Automation

Cerebras
United States, Canada
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 2+ yearsSkills: ["Hands-on operational execution","Automation mindset","Collaboration","Communication","Impact measurement"]

Own real production systems for a rapidly growing AI inference service, helping operate the current environment and bring up new capacity in high-stakes setups. Build and extend continuous delivery and self-service pipelines, improve telemetry/observability/alerting for reliability at scale, and reduce operational toil with reusable automation and internal developer tools. Collaborate with Cluster Ops and development teams on SLOs, post-mortems, and capacity planning (no 24/7 on-call).

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cerebras
Cerebras
11 months ago

Site Reliability Engineer - Ops & Automation

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 16 hours agoStatus: Live

Job Summary

Own real production systems for a rapidly growing AI inference service, helping operate the current environment and bring up new capacity in high-stakes setups. Build and extend continuous delivery and self-service pipelines, improve telemetry/observability/alerting for reliability at scale, and reduce operational toil with reusable automation and internal developer tools. Collaborate with Cluster Ops and development teams on SLOs, post-mortems, and capacity planning (no 24/7 on-call).
Location: United States, Canada
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Remain hands-on with operational execution, including releases, capacity changes, and cluster upgrades, while building robust continuous delivery and self-service capabilities.
  • •Develop self-service continuous delivery pipelines for key workflows using Kubernetes, Bazel, Prometheus/Grafana/InfluxDB, Python, and Go.
  • •Build reusable automation and internal developer tools to minimize operational toil and cross-team friction.
  • •Develop and extend telemetry, observability, and alerting solutions to ensure operational reliability at scale.
  • •Collaborate with Cluster Ops and development teams to identify high-impact automation opportunities and improve reliability practices like SLOs, post-mortems, and capacity planning.

Key Requirements

  • •2-4+ years in SRE with a strong operations or automation focus.
  • •Production Kubernetes experience.
  • •Build tools and automation with Python or Go.
  • •Use Prometheus and Grafana with observability-driven workflows.
  • •Measure and communicate impact using reliability metrics, operational toil, and velocity gains.
Experience:2+ yearsSREAI inferenceKubernetesContinuous deliveryObservability
Skills:Hands-on operational executionAutomation mindsetCollaborationCommunicationImpact measurement
Tech Stack:KubernetesBazelPrometheusGrafanaInfluxDBPythonGoGitOpsArgo CDFluxContinuous delivery pipelinesSelf-service

Company Brief

Cerebras
Designs and builds wafer-scale AI accelerators and systems for large-scale deep learning workloads, delivering specialized hardware and software to accelerate model training and inference for enterprises and research institutions.
Industry: Hardware Devices
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Headquarters: Sunnyvale, United States
Founded: 2016
WebsiteLinkedIn