Incident Operations Specialist

Zapier
San Francisco
Workplace: RemoteFull timeUSD 119,000 - 178,500 annuallyFunction: DevOps, Cloud & InfrastructureSkills: ["Attention to detail","Written communication","Prioritization","Problem-solving","Cross-functional collaboration"]

Own Zapier’s incident program operations—maintaining incident.io, PagerDuty, Slack-based workflows, and on-call escalation paths. Build and improve repeatable AI-powered workflows for summarization, postmortem drafts, triage, and data hygiene. Operate reporting systems using Databricks, Grafana, and Looker, ensure data quality, and keep playbooks up to date. Collaborate cross-functionally to keep incident tooling reliable and operations visible at scale.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Zapier
Zapier
1 month ago

Incident Operations Specialist

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 12 hours agoStatus: Live

Job Summary

Own Zapier’s incident program operations—maintaining incident.io, PagerDuty, Slack-based workflows, and on-call escalation paths. Build and improve repeatable AI-powered workflows for summarization, postmortem drafts, triage, and data hygiene. Operate reporting systems using Databricks, Grafana, and Looker, ensure data quality, and keep playbooks up to date. Collaborate cross-functionally to keep incident tooling reliable and operations visible at scale.
Location: San Francisco
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Maintain the reliability and configuration of incident tooling (incident.io, PagerDuty, Slack-based workflows), integrations, on-call rotations, and escalation paths.
  • •Design and ship repeatable AI-powered workflows for incident operations such as thread summarization, postmortem draft generation, follow-up triage, severity classification, and data hygiene.
  • •Build and sustain a community of practice for Incident Commanders and Support Leads, including touchpoints and coaching.
  • •Operate and troubleshoot dashboards and reports using Databricks, Grafana, and Looker, ensuring data quality and metric accuracy.
  • •Maintain incident playbooks, templates, and documentation; participate in incidents and postmortem reviews to drive continuous improvement.

Pay and Benefits

Salary: USD 119,000 - 178,500 annually

Key Requirements

  • •Experience in incident response, technical operations, or a reliability-adjacent role.
  • •Hands-on familiarity with incident tooling (incident.io, PagerDuty or equivalent) and Slack-based workflow automation.
  • •Comfortable building with SQL and reporting tools (Databricks, Looker, Grafana).
  • •Able to diagnose operational issues using logs, system context, and integration debugging.
  • •Demonstrably uses AI in daily work, including building repeatable AI-powered workflows with verification and judgment in incident contexts.
Skills:Attention to detailWritten communicationPrioritizationProblem-solvingCross-functional collaboration
Tech Stack:Incident.ioPagerDutySlackDatabricksGrafanaLookerSQLDatadogPrometheusOpensearchGraylogGitLabCodaGoogle WorkspaceJiraZendeskCursorZapier AIClaudeCopilot

Company Brief

Zapier
Provides a no-code automation platform that connects apps and automates workflows for businesses and individuals, enabling integrations between thousands of web applications to streamline repetitive tasks and improve productivity.
Industry: SaaS
Company Size: Enterprise (1,001+ employees)
Growth: Established Company
Founded: 2011
WebsiteLinkedIn