Principal Associate SRE

Discover Financial
Mexico City
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureExperience: 4+ yearsEducation: bachelorsSkills: ["Troubleshooting","Debugging","Independent problem-solving","Incident response","Blameless postmortem contributions"]

Build and maintain Site Reliability Engineering capabilities for payment-critical systems, focusing on settlement reliability, alert signal/observability, and reliability automation. You’ll develop automation tooling using Python, Java, and shell scripts, troubleshoot production issues across hybrid on-prem and AWS, and tune monitoring in Datadog and Observe. Participate in on-call rotations, manage secrets/certificates, and deliver improvements through CI/CD and API automation.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Discover Financial
Discover Financial
2 months ago

Principal Associate SRE

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 17 minutes agoStatus: Live

Job Summary

Build and maintain Site Reliability Engineering capabilities for payment-critical systems, focusing on settlement reliability, alert signal/observability, and reliability automation. You’ll develop automation tooling using Python, Java, and shell scripts, troubleshoot production issues across hybrid on-prem and AWS, and tune monitoring in Datadog and Observe. Participate in on-call rotations, manage secrets/certificates, and deliver improvements through CI/CD and API automation.
Location: Mexico City
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Sr. Manager level

Key Responsibilities

  • •Build and maintain reliability tooling such as observability dashboards, automated alerts, runbooks, and remediation scripts.
  • •Develop automation solutions to eliminate manual operational processes, including certificate rotation and compliance artifact generation.
  • •Troubleshoot and debug complex production issues across distributed systems spanning on-prem data centers and AWS, and implement durable fixes.
  • •Tune observability by configuring monitoring in Datadog and Observe and reducing unactionable alert volume.
  • •Participate in on-call incident response and contribute to blameless postmortems, using CI/CD and automation agents to accelerate engineering output.

Key Requirements

  • •Professional English fluency.
  • •Bachelor’s degree.
  • •Background in SRE, production operations, or reliability engineering.
  • •At least 4 years of DevOps Engineering experience (internship experience does not apply).
  • •4+ years experience in Java, Python, or Go, plus 2+ years cloud-native experience and 2+ years container orchestration (Docker or Kubernetes).
Experience:4+ yearsSREDevOpsProduction operationsReliability engineeringCloud-nativeContainersDistributed systemsPayments
Education:Bachelor's
Skills:TroubleshootingDebuggingIndependent problem-solvingIncident responseBlameless postmortem contributions
Languages:English
Tech Stack:PythonJavaGoShell scriptingBashAWSMicrosoft AzureGoogle Cloud PlatformDockerKubernetesOpenShiftCI/CD pipelinesAPI automation frameworksDatadogObserveHashiCorp VaultClaude CodeCopilot CLIUnixLinux

Company Brief

Discover Financial
Provides consumer banking products, credit cards, personal loans, and payment services through the Discover brand. It operates a major U.S. financial network and serves individuals and merchants with lending and digital payment solutions.
Industry: Retail Banking
Company Size: Enterprise (1,001+ employees)
Revenue: USD 10M to 25M
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Riverwoods, United States
Founded: 1985
WebsiteLinkedIn