Senior Site Reliability Engineer

Sanity
New York, Canada, United States, America/Montreal
Workplace: RemoteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsSkills: ["Analytical thinking","Communication","Incident response","Mentoring","Collaboration"]

Build and operate the shared infrastructure foundations that power Sanity’s AI Content Operating System at real-world scale. Partner with development teams to diagnose distributed systems, ensure observability, and improve reliability through better alerting, dashboards, paging standards, incident response, and safe deployment practices. Help modernize edge, caching, and gateway layers on Fastly while mentoring engineers through code and design reviews.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Sanity
Sanity
2 months ago

Senior Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 5 hours agoStatus: Live

Job Summary

Build and operate the shared infrastructure foundations that power Sanity’s AI Content Operating System at real-world scale. Partner with development teams to diagnose distributed systems, ensure observability, and improve reliability through better alerting, dashboards, paging standards, incident response, and safe deployment practices. Help modernize edge, caching, and gateway layers on Fastly while mentoring engineers through code and design reviews.
Location: New York, Canada, United States, America/Montreal
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Sr. Manager level

Key Responsibilities

  • •Design, build, and operate shared platform foundations including GCP infrastructure, Kubernetes, networking/routing, CI/CD, and observability.
  • •Diagnose and troubleshoot complex distributed systems running at high request volume.
  • •Ensure observability and analyze behavior of the stack.
  • •Modernize edge, caching, and gateway layers onto Fastly and tighten observability across the platform.
  • •Raise the reliability bar via improved dashboards, alert severity, paging standards, on-call readiness, incident response, and production readiness automation.

Pay and Benefits

Perks:Health InsuranceEquity

Key Requirements

  • •5+ years of experience as part of an SRE on-call rotation.
  • •Experience with SRE/DevOps tools, processes, and culture.
  • •Strong experience managing scalable, highly available, cloud-based applications with customer-facing uptime expectations.
  • •Hands-on experience with Kubernetes for orchestrating and managing containerized applications in cloud environments.
  • •Experience building CI/CD pipelines and using an observability stack (e.g., Prometheus).
Experience:5+ yearsSREDevOpsCloudDistributed systemsHigh availability
Skills:Analytical thinkingCommunicationIncident responseMentoringCollaboration
Tech Stack:KubernetesPrometheusElasticSearchPostgreSQLNATSKongFastlyGoogle Cloud PlatformGCPCI/CDNetworkingRoutingObservability

Company Brief

Sanity
Provides a headless CMS and structured content platform that enables developers and teams to build, manage, and deliver content via real-time APIs, customizable studio, and rich tooling for digital experiences across web and mobile.
Industry: Developer Tools
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Headquarters: Oslo, Norway
Founded: 2015
WebsiteLinkedIn