SRE Lead

Fiserv
Dublin
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Reliability","Incident response","Resilience engineering","Operability","Collaboration"]

Own production readiness and operational health for a high-volume payment platform modernization, migrating from legacy PL/SQL to cloud-native, event-driven microservices on AWS. Define observability and SLIs/SLOs, build Dynatrace/APM, Splunk logging, and Prometheus metrics, and lead incident response with blameless post-mortems. Drive resilience engineering (chaos, failover, load tests), support operability in service designs, and automate operational toil during legacy-to-new transitions.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Fiserv
Fiserv
1 month ago

SRE Lead

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 22 hours agoStatus: Live

Job Summary

Own production readiness and operational health for a high-volume payment platform modernization, migrating from legacy PL/SQL to cloud-native, event-driven microservices on AWS. Define observability and SLIs/SLOs, build Dynatrace/APM, Splunk logging, and Prometheus metrics, and lead incident response with blameless post-mortems. Drive resilience engineering (chaos, failover, load tests), support operability in service designs, and automate operational toil during legacy-to-new transitions.
Location: Dublin
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Sr. Manager level

Key Responsibilities

  • •Define and own production readiness standards (observability, SLOs, runbooks, resource limits, graceful degradation, policy compliance).
  • •Build and maintain observability across the platform, including Dynatrace APM, Splunk log pipelines via Fluent Bit, and Prometheus metrics with tracing and dashboards.
  • •Define and implement SLIs/SLOs and build the measurement infrastructure with engineering teams.
  • •Own incident response practices, including detection/response and blameless post-mortems for reliability improvements.
  • •Drive resilience and operational confidence through chaos testing, failover validation, load testing, and guidance on retry/timeouts/circuit breakers/back-pressure.

Key Requirements

  • •Experience in SRE, production engineering, or operations for high-volume distributed systems with improved reliability.
  • •Hands-on observability skills including distributed tracing (OpenTelemetry), metrics (Prometheus/Grafana), structured logging, and APM tools (Dynatrace/Datadog).
  • •Experience defining and measuring SLIs/SLOs to drive engineering decisions.
  • •Knowledge of resilience patterns for distributed systems (circuit breakers, retries, timeouts, bulkheads, back-pressure) and validating them.
  • •Comfort with cloud infrastructure and event-driven architectures on AWS and Kubernetes, including Kafka/operational challenges.
Experience:FintechPaymentsMicroservicesEvent-driven systemsCloudDistributed systems
Skills:ReliabilityIncident responseResilience engineeringOperabilityCollaboration
Tech Stack:PL/SQLAWSConfluent CloudAurora PostgreSQLEKSKubernetesJavaQuarkusDynatraceSplunkFluent BitPrometheusGrafanaOpenTelemetryArgo RolloutsKafkaInfrastructure-as-code

Company Brief

Fiserv
Provides payments, processing services, risk management, and core banking technology to financial institutions, merchants, and businesses worldwide, enabling digital payments, account processing, and financial services integration.
Industry: Fintech Infrastructure
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Brookfield, United States
Founded: 1984
WebsiteLinkedIn