Site Reliability Engineer

Cognition
San Francisco
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Problem-solving","Ownership","Communication","Collaboration"]

Join a compact, high-skill team building AI-driven software and assume ownership of production reliability and platform engineering. You’ll define and own SLOs/SLIs, lead incident response and on-call improvements, manage CI/CD pipelines and release infrastructure, and implement infrastructure as code to scale across cloud environments. Expect a culture that treats reliability as a craft and partners closely with product and engineering to prevent outages.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cognition
Cognition
11 months ago

Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 6 hours agoStatus: Live

Job Summary

Join a compact, high-skill team building AI-driven software and assume ownership of production reliability and platform engineering. You’ll define and own SLOs/SLIs, lead incident response and on-call improvements, manage CI/CD pipelines and release infrastructure, and implement infrastructure as code to scale across cloud environments. Expect a culture that treats reliability as a craft and partners closely with product and engineering to prevent outages.
Location: San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Own production reliability concepts for Devin and Windsurf, including defining and monitoring SLOs, SLIs, and error budgets.
  • •Lead incident response and on-call rotations; conduct blameless postmortems and implement durable improvements.
  • •Own platform engineering, CI/CD pipelines, release infrastructure, and internal tooling to enable rapid, safe shipping.
  • •Manage Infrastructure as Code and cloud environments to ensure reproducible, auditable deployments and prevent drift.
  • •Collaborate with product and engineering teams to embed reliability from the start and reduce toil through automation.

Key Requirements

  • •Deep experience running production systems at scale: SLOs, error budgets, on-call rotations, and incident command
  • •Strong software engineering fundamentals; SRE at Cognition means writing real code, not just configuring tools
  • •Proficiency with cloud infrastructure (AWS, GCP, or Azure), container orchestration (Kubernetes), and infrastructure as code (Terraform or equivalent)
  • •Experience building and owning CI/CD pipelines and deployment infrastructure for fast-moving product teams
  • •Strong observability instincts: knows how to instrument systems, build useful dashboards, and design alerts that surface signal without generating noise
Experience:Artificial intelligenceSoftware engineeringDeveloper tools
Skills:Problem-solvingOwnershipCommunicationCollaboration
Tech Stack:AWSGCPAzureKubernetesTerraformCI/CDObservabilityMonitoring

Company Brief

Cognition
Builds Devin, an autonomous AI software engineer and AI-native developer tools (Windsurf, DeepWiki) to automate software engineering workflows for enterprise customers.
Industry: Developer Tools
Company Size: Medium (51 to 250 employees)
Revenue: USD 50M to 100M
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series C
Headquarters: San Francisco, United States
Founded: 2023
Glassdoor
Glassdoor: 4.7
WebsiteLinkedInGlassdoor