IT SRE Team Lead

Cerebras
Sunnyvale
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 8+ yearsSkills: ["Organizational skills","Multitasking","Prioritization","Detail-oriented","Communication"]

Lead the reliability function for internal IT systems by defining SLOs, error budgets, and operational health reporting. Build and run a small IT SRE team focused on automation, observability, and incident response across identity, endpoints, collaboration, SaaS, and internal networking. Drive infrastructure-as-code and GitOps practices, implement automation to eliminate toil, instrument services with monitoring and on-call workflows, and partner with security and networking teams to keep the environment fast, stable, and secure as the company scales.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cerebras
Cerebras
5 months ago

IT SRE Team Lead

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 16 hours agoStatus: Live

Job Summary

Lead the reliability function for internal IT systems by defining SLOs, error budgets, and operational health reporting. Build and run a small IT SRE team focused on automation, observability, and incident response across identity, endpoints, collaboration, SaaS, and internal networking. Drive infrastructure-as-code and GitOps practices, implement automation to eliminate toil, instrument services with monitoring and on-call workflows, and partner with security and networking teams to keep the environment fast, stable, and secure as the company scales.
Location: Sunnyvale
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Manager level

Key Responsibilities

  • •Define and own the reliability strategy for internal IT systems, including SLOs, error budgets, and operational health reporting.
  • •Build and lead an IT SRE team focused on automation, observability, and incident response for corporate systems.
  • •Design and implement automation to eliminate manual IT work across provisioning, access management, patching, and lifecycle operations.
  • •Instrument internal services and SaaS integrations with monitoring, alerting, and on-call workflows, then run incident response with root cause analysis and durable remediation.
  • •Drive infrastructure-as-code and GitOps practices and partner with security and networking teams on identity, access, and network reliability.

Key Requirements

  • •Minimum 8 years of experience in SRE, DevOps, or IT engineering roles, with at least 2 years in a leadership capacity.
  • •Direct hands-on experience building and deploying AI agents for triage and bug fixes.
  • •Strong software engineering background with hands-on experience in Python, Go, or similar, and comfort writing production-grade automation.
  • •Deep experience with identity platforms (Okta, Entra), endpoint management (Jamf, Intune), and SaaS integration patterns.
  • •Hands-on experience with infrastructure-as-code (Terraform) and CI/CD pipelines for IT systems.
Experience:8+ years
Skills:Organizational skillsMultitaskingPrioritizationDetail-orientedCommunication
Tech Stack:PythonGoOktaEntraJamfIntuneTerraformCI/CDGitOpsAI agents

Company Brief

Cerebras
Designs and builds wafer-scale AI accelerators and systems for large-scale deep learning workloads, delivering specialized hardware and software to accelerate model training and inference for enterprises and research institutions.
Industry: Hardware Devices
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Headquarters: Sunnyvale, United States
Founded: 2016
WebsiteLinkedIn