Site Reliability Engineer (SRE)

Retool
San Francisco
Workplace: HybridFull timeUSD 163,710 - 306,000 annuallyFunction: DevOps, Cloud & InfrastructureSkills: ["Clear written communication","Bias toward automation","Operational judgment","Leadership through ambiguity","Customer-focused collaboration"]

Own reliability across Retool Cloud and customer-managed deployment paths, including provisioning, upgrades, migrations, configuration changes, and production escalations. Build automation that replaces manual infrastructure work (Terraform runs, upgrade workflows, secret rotations) and improve observability by turning health signals into clear status and actions. Design safer deployment/upgrade/rollback paths and partner with product engineers while writing runbooks and migration guides.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Retool
Retool
2 months ago

Site Reliability Engineer (SRE)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 10 minutes agoStatus: Live

Job Summary

Own reliability across Retool Cloud and customer-managed deployment paths, including provisioning, upgrades, migrations, configuration changes, and production escalations. Build automation that replaces manual infrastructure work (Terraform runs, upgrade workflows, secret rotations) and improve observability by turning health signals into clear status and actions. Design safer deployment/upgrade/rollback paths and partner with product engineers while writing runbooks and migration guides.
Location: San Francisco
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Own reliability across Retool Cloud, managed single tenant, BYOC, and self-hosted deployment paths, including provisioning, upgrades, migrations, configuration changes, and production escalations.
  • •Build automation to make manual infrastructure work repeatable, including Terraform runs, customer environment updates, upgrade workflows, secret rotations, and migration steps.
  • •Improve observability for Retool Cloud, self-hosted customers, and internal operators by translating health signals into clear status, likely causes, and recommended actions.
  • •Design safer deployment, upgrade, and rollback paths so cloud and managed customers can stay current.
  • •Write docs, runbooks, design notes, and migration guides to help engineers and customers understand complex systems; partner with product engineers on infrastructure requirements for new products.

Pay and Benefits

Salary: USD 163,710 - 306,000 annually
Perks:Health InsuranceDentalVision401k

Key Requirements

  • •Deep experience operating production infrastructure in AWS.
  • •Strong Kubernetes fundamentals and ability to debug Kubernetes issues.
  • •Real Terraform or infrastructure-as-code experience to support automated infrastructure changes.
  • •Operational judgment around databases, especially Postgres.
  • •Experience building or operating observability systems and improving reliability for customer-facing SaaS.
Experience:SaaSCustomer-facing systems
Skills:Clear written communicationBias toward automationOperational judgmentLeadership through ambiguityCustomer-focused collaboration
Tech Stack:AWSKubernetesHelmDocker ComposeTerraformPostgresNetworkingGoPythonTypeScriptJavaRuby

Company Brief

Retool
Provides a low-code platform for building internal tools and business applications, enabling developers and teams to assemble UIs on top of databases and APIs to accelerate operational workflows.
Industry: Developer Tools
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series D
Headquarters: San Francisco, United States
Founded: 2017
WebsiteLinkedIn