Site Reliability Engineer, Observability

Ripple
Chicago, New York
Workplace: HybridFull timeUSD 160,000 - 200,000 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 7+ yearsSkills: ["Coaching","Mentoring","Troubleshooting","Stakeholder communication","Cross-functional collaboration"]

Build observability and reliability for enterprise treasury infrastructure at scale across Azure and AWS. Design and implement monitoring, alerting, dashboards, and NRQL queries in New Relic, including SLOs/SLIs and error budgets. Lead alert signal quality and observability cost optimization while developing Terraform-based IaC for monitoring resources. Establish and run incident management foundations using Incident.IO, partnering with product and platform teams to improve operational maturity.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Ripple
Ripple
3 months ago

Site Reliability Engineer, Observability

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 14 hours agoStatus: Live

Job Summary

Build observability and reliability for enterprise treasury infrastructure at scale across Azure and AWS. Design and implement monitoring, alerting, dashboards, and NRQL queries in New Relic, including SLOs/SLIs and error budgets. Lead alert signal quality and observability cost optimization while developing Terraform-based IaC for monitoring resources. Establish and run incident management foundations using Incident.IO, partnering with product and platform teams to improve operational maturity.
Location: Chicago, New York
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Design and implement observability in New Relic (dashboards, alerting, dashboards) across Azure and AWS, including NRQL query authoring and SLO/SLI management.
  • •Develop and maintain Terraform IaC for provisioning and managing monitoring and observability infrastructure, and provide IaC governance across teams.
  • •Administer incident management using Incident.IO (alert routing, notification workflows, integrations) and build incident management foundations (on-call, escalation, playbooks).
  • •Reduce alert noise and improve signal quality by tuning thresholds and eliminating false positives to ensure actionable alerts.
  • •Respond to and debrief production incidents and enable stream-aligned teams through workshops, documentation, and hands-on guidance.

Pay and Benefits

Salary: USD 160,000 - 200,000 annually
Equity and Bonus:Equity
Perks:Health InsuranceRetirementLearning BudgetParental LeaveWellness StipendMobile Stipend

Key Requirements

  • •7+ years in Site Reliability Engineering, DevOps, or Platform Engineering with strong observability and production operations focus.
  • •Expert hands-on experience with New Relic (APM, Infrastructure, Logs, Synthetics, Alerts) and strong NRQL proficiency.
  • •Expertise defining and implementing SLOs/SLIs and error budgets for reliability management.
  • •Hands-on experience with incident management platforms including Incident.IO, PagerDuty, OpsGenie, or similar.
  • •Strong Terraform and infrastructure tooling experience, including PowerShell for Windows environments and Azure (App Services, Virtual Machines, Azure SQL, networking, monitoring).
Experience:7+ yearsSREDevOpsPlatform EngineeringObservabilityIncident ManagementFinTech
Skills:CoachingMentoringTroubleshootingStakeholder communicationCross-functional collaboration
Languages:English
Tech Stack:New RelicAPMInfrastructureLogsSyntheticsAlertsNRQLSLOs/SLIsTerraformAzure DevOpsAzureAWSWindowsPowerShellIncident.IOSlackOpsGenieDistributed tracingStructured loggingRED/USE

Company Brief

Ripple
Builds blockchain-based payment and liquidity infrastructure for financial institutions and enterprises. Its products support cross-border payments, crypto liquidity, and digital asset settlement.
Industry: Fintech Infrastructure
Company Size: Enterprise (1,001+ employees)
Growth: Established Company
Headquarters: San Francisco, United States
Founded: 2012
WebsiteLinkedIn