Staff Software Engineer I - SRE

Confluent
Bengaluru
Workplace: RemoteFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Communication","Leadership","Problem-solving","Collaboration","Writing"]

Proactive reliability-focused engineer responsible for reducing incident recurrence and improving the observability and incident-response of a large, multi-cloud streaming platform. You’ll split time between engineering work to build automation and tooling, and leadership activities such as training incident commanders, coordinating post-mortems, and driving enterprise-wide reliability standards across a follow-the-sun team of 800-1000 engineers.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Confluent
Confluent
7 months ago

Staff Software Engineer I - SRE

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 12 hours agoStatus: Live

Job Summary

Proactive reliability-focused engineer responsible for reducing incident recurrence and improving the observability and incident-response of a large, multi-cloud streaming platform. You’ll split time between engineering work to build automation and tooling, and leadership activities such as training incident commanders, coordinating post-mortems, and driving enterprise-wide reliability standards across a follow-the-sun team of 800-1000 engineers.
Location: Bengaluru
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Sr. Manager level

Key Responsibilities

  • •Proactive reliability engineering: analyze systemic failure patterns and design improvements that prevent incident recurrence; define and maintain SLO/SLA frameworks; use error budgets to guide reliability investments; build tooling and automation to reduce incident response toil; own Rootly configuration and integrations with PagerDuty, Jira, Confluence, and Slack; analyze reliability data to identify improvements and build dashboards.
  • •Incident management program: own standards and continuous improvement of incident response; define incident commander criteria and manage rotation; serve as escalation IC when needed; develop and deliver training for engineering teams; coach teams through post-mortems and corrective actions.
  • •Customer root cause analysis (CRCA): edit/review customer-facing incident documents; drive turnaround SLAs while maintaining accuracy; ensure clear explanation of what happened, why, and how to prevent recurrence.
  • •Cross-team leadership: partner with engineering leaders to elevate reliability practices; be the go-to expert teams proactively engage for guidance.
  • •Design scalable reliability standards to reduce reactive workload over time and drive proactive improvements across teams.

Key Requirements

  • •10+ years in SRE, incident management, or reliability engineering.
  • •Cloud experience with at least one of AWS, GCP, or Azure.
  • •Deep expertise with incident management tooling (Rootly, PagerDuty, or similar platforms).
  • •Strong understanding of distributed systems and failure modes at scale; Kafka/event streaming expertise preferred, or demonstrated rapid mastery of complex systems.
  • •Deep experience with observability: metrics, logging, tracing; Kubernetes and container orchestration experience; CI/CD pipelines and release processes; systems thinking; familiarity with SLO/SLA frameworks.
Skills:CommunicationLeadershipProblem-solvingCollaborationWriting
Languages:English
Tech Stack:RootlyPagerDutyKafkaEvent streamingKubernetesCI/CDGitHubConfluenceJiraSlackObservabilityMetricsLoggingTracing

Company Brief

Confluent
Provides a cloud-native data streaming platform built on Apache Kafka that enables organizations to harness real-time data for event streaming, analytics, and AI applications across hybrid and multi-cloud environments.
Industry: Data Infrastructure
Company Size: Enterprise (1,001+ employees)
Revenue: USD 500M to 1B
Growth: Public Company
Valuation: Unicorn (USD 1B+)
Funding: IPO / Publicly Listed
Headquarters: Mountain View, United States
Founded: 2014
Glassdoor
Glassdoor: 3.6
WebsiteLinkedInGlassdoor