Site Reliability Engineer

Razorpay
Bengaluru
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 10+ yearsSkills: ["Communication","Judgment","Incident leadership","Blameless postmortems","On-call ownership","Pragmatism"]

Be one of Razorpay’s founding SREs embedded with payment platform teams to raise payment availability from three nines to four/five nines. Define SLIs, SLOs, and error budgets for key payment flows, own release lifecycles with canary and automated rollback, and lead incident response with blameless postmortems. Reduce toil with automation, harden distributed-system failure modes, and drive production readiness using reliability data and observability.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Razorpay
Razorpay
1 day ago

Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 17 hours agoStatus: Live

Job Summary

Be one of Razorpay’s founding SREs embedded with payment platform teams to raise payment availability from three nines to four/five nines. Define SLIs, SLOs, and error budgets for key payment flows, own release lifecycles with canary and automated rollback, and lead incident response with blameless postmortems. Reduce toil with automation, harden distributed-system failure modes, and drive production readiness using reliability data and observability.
Location: Bengaluru
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Define SLIs, SLOs, and error budgets for critical payment flows (authorization, capture, refunds, settlements, webhooks).
  • •Own the release lifecycle for payment services with progressive rollouts (canary, staged, feature-flagged) and automated rollback triggers.
  • •Carry the pager for payment-critical services, lead incident command during outages, and drive blameless postmortems with action items that ship.
  • •Eliminate toil by building automation for failover, capacity management, load shedding, and degradation to prevent recurring known failures.
  • •Harden payment flows against distributed-systems failure modes and run production readiness reviews using error budget data.

Key Requirements

  • •10+ years of engineering experience, including 5+ years operating large-scale distributed systems in production with business-critical reliability requirements.
  • •Strong software engineering skills in Go, Java, or Python, with experience building tools and services.
  • •In-depth understanding of distributed systems failure modes and mitigation patterns (e.g., circuit breakers, backpressure, bulkheading, graceful degradation, idempotency).
  • •Solid fundamentals in Linux internals, networking, and databases under load (e.g., replication, failover, connection pool exhaustion, lock contention).
  • •Hands-on experience designing deployment pipelines (e.g., canary analysis, automated rollback, feature flags) and demonstrating genuine on-call ownership.
Experience:10+ years
Skills:CommunicationJudgmentIncident leadershipBlameless postmortemsOn-call ownershipPragmatism
Languages:English
Tech Stack:GoJavaPythonLinuxNetworkingDatabasesPrometheusGrafanaOpenTelemetryDatadogCoralogixClickhouseKubernetesService meshFeature flagsCanaryProgressive rolloutAutomated rollbackStructured loggingMetrics

Company Brief

Razorpay
Provides payment processing, banking, lending, and financial infrastructure products for Indian businesses. Razorpay offers payment gateway services, payouts, banking-as-a-service, credit, and developer-focused APIs to help companies accept and manage online payments.
Industry: Fintech Infrastructure
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series E+
Headquarters: Bengaluru, India
Founded: 2014
WebsiteLinkedIn