Lead Site Reliability Engineer

Discover Financial
Mexico City
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 6+ yearsEducation: bachelorsSkills: ["Troubleshooting","Debugging","Incident response","Root cause analysis","Documentation"]

Build and run reliability for batch settlement systems that process every credit and debit transaction, ensuring settlement cycles complete accurately, on time, and with SOX/PCI-DSS readiness. Improve observability (dashboards, alerts, anomaly detection) and automate operational toil such as certificate rotation and compliance artifacts. Partner with UK-based settlement engineers, lead incident response with durable root-cause fixes, and help run audit-ready operations with minimal manual effort.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Discover Financial
Discover Financial
2 months ago

Lead Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 15 minutes agoStatus: Live

Job Summary

Build and run reliability for batch settlement systems that process every credit and debit transaction, ensuring settlement cycles complete accurately, on time, and with SOX/PCI-DSS readiness. Improve observability (dashboards, alerts, anomaly detection) and automate operational toil such as certificate rotation and compliance artifacts. Partner with UK-based settlement engineers, lead incident response with durable root-cause fixes, and help run audit-ready operations with minimal manual effort.
Location: Mexico City
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Sr. Manager level

Key Responsibilities

  • •Own reliability for batch settlement systems, ensuring cycle completion windows, data integrity, and failure detection before downstream impact.
  • •Build and improve observability for settlement pipelines using dashboards, alerts, and anomaly detection to reduce reliance on tribal knowledge.
  • •Automate operational toil including certificate rotation, environment provisioning, compliance artifact generation, and manual validation steps.
  • •Partner with UK-based settlement engineers to build domain expertise and support compliance-related windows and SLA adherence.
  • •Participate in incident management and implement durable fixes via root cause analysis to prevent recurrence.

Key Requirements

  • •6+ years of experience in SRE, production operations, or reliability engineering.
  • •5+ years of experience in at least one: Java, Python, or Go.
  • •4+ years of experience with cloud-native technologies (AWS, Microsoft Azure, or Google Cloud Platform).
  • •3+ years with container orchestration services including Docker or Kubernetes.
  • •Experience with Shell/Bash scripting and 3+ years of Unix/Linux system administration.
Experience:6+ yearsFinancial services
Education:Bachelor's
Skills:TroubleshootingDebuggingIncident responseRoot cause analysisDocumentation
Languages:English
Tech Stack:JavaPythonGoShell scriptingBashSQLAWSMicrosoft AzureGoogle Cloud PlatformDockerKubernetesOpenShiftDatadogObserveHashiCorp VaultCI/CDAPI automationClaude CodeCopilot CLITCP

Company Brief

Discover Financial
Provides consumer banking products, credit cards, personal loans, and payment services through the Discover brand. It operates a major U.S. financial network and serves individuals and merchants with lending and digital payment solutions.
Industry: Retail Banking
Company Size: Enterprise (1,001+ employees)
Revenue: USD 10M to 25M
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Riverwoods, United States
Founded: 1985
WebsiteLinkedIn