Site Reliability Engineer

Gamma
San Francisco
Workplace: HybridFull timeUSD 230,000 - 310,000 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsSkills: ["Incident management","Communication","Leadership","Collaboration","Problem-solving"]

This role owns the reliability and performance of Gamma’s production backend on AWS, building observability, automation, and tooling to reduce toil. You’ll lead incident response, drive systemic reliability improvements, and collaborate with engineering on architecture reviews and SLO/SLI design to scale a SaaS platform serving millions of users.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Gamma
Gamma
10 months ago

Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

This role owns the reliability and performance of Gamma’s production backend on AWS, building observability, automation, and tooling to reduce toil. You’ll lead incident response, drive systemic reliability improvements, and collaborate with engineering on architecture reviews and SLO/SLI design to scale a SaaS platform serving millions of users.
Location: San Francisco
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Sr. Manager level

Key Responsibilities

  • •Own the reliability, availability, and performance of Gamma's production systems across our AWS infrastructure
  • •Build observability infrastructure from the ground up: metrics, logging, tracing, and alerting that give the team genuine visibility into system health before users feel the impact
  • •Design and ship automation that reduces toil, makes deployments safer, and gets us back on our feet faster when things go wrong
  • •Lead incident response and blameless post-mortems, then follow through on the systemic fixes that keep the same issues from coming back
  • •Partner with engineering teams on architecture reviews, SLO and SLI design, and reliability best practices that scale with the product

Pay and Benefits

Salary: USD 230,000 - 310,000 annually

Key Requirements

  • •5+ years in site reliability engineering, DevOps, or systems engineering with deep, hands-on AWS expertise
  • •Strong programming skills in Python, Go, or TypeScript/Node.js, applied to building real tools and automation
  • •Solid experience with infrastructure-as-code (Terraform, CloudFormation) and end-to-end observability solutions
  • •Track record of making systems meaningfully more reliable through automation, smarter monitoring, and architectural improvements
  • •Deep understanding of networking, distributed systems, containerization (Docker, Kubernetes), and database performance at scale
Experience:5+ yearsSaaSCloud
Skills:Incident managementCommunicationLeadershipCollaborationProblem-solving
Certifications:AWS
Languages:English
Tech Stack:AWSPythonGoTypeScriptNode.jsTerraformCloudFormationDockerKubernetesKafka

Company Brief

Gamma
Gamma (gamma.app) is an AI-powered visual storytelling platform that generates presentations, websites, documents, and social content from text or briefs — combining AI design agents, smart layouts, and collaboration tools for fast, polished outputs.
Industry: SaaS
Company Size: Medium (51 to 250 employees)
Revenue: USD 100M to 250M
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series B
Headquarters: San Francisco, United States
Founded: 2020
WebsiteLinkedIn