Site Reliability Engineer

Moniepoint Group
South Africa
Workplace: RemoteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsSkills: ["Incident response","Troubleshooting","Cross-functional coordination","Stakeholder communication","Automation mindset"]

Own end-to-end service reliability by participating in on-call rotations, triaging incidents, and acting as Incident Commander during major outages. Build visibility with dashboards, alerts, and application instrumentation, then improve resilience by defining SLIs/SLOs and implementing automation to reduce operational toil. Investigate escalated customer issues, especially complex performance and reliability problems, while partnering with development teams across cloud and distributed systems environments.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Moniepoint Group
Moniepoint Group
2 months ago

Site Reliability Engineer

âś“ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 3 hours agoStatus: Live

Job Summary

Own end-to-end service reliability by participating in on-call rotations, triaging incidents, and acting as Incident Commander during major outages. Build visibility with dashboards, alerts, and application instrumentation, then improve resilience by defining SLIs/SLOs and implementing automation to reduce operational toil. Investigate escalated customer issues, especially complex performance and reliability problems, while partnering with development teams across cloud and distributed systems environments.
Location: South Africa
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Participate in on-call rotations to detect, triage, and drive resolution of service and reliability issues across all environments.
  • •Act as Incident Commander during major incidents, coordinating cross-functional teams and providing timely stakeholder updates.
  • •Create and maintain dashboards and alerts, and partner with development teams to instrument code for end-to-end visibility.
  • •Develop automation to eliminate manual operational toil across both applications and infrastructure.
  • •Implement and track SLIs and SLOs, and investigate escalated customer complaints related to performance, reliability, or complex system behavior.

Pay and Benefits

Perks:PensionHealth InsuranceAnnual Bonus

Key Requirements

  • •Minimum 5 years of experience as an SRE (or similar) supporting enterprise applications, with coding proficiency in Java, Go, or Python.
  • •Strong understanding of distributed systems, microservices architecture, and software design patterns.
  • •Hands-on experience with Kubernetes; you’ve managed applications on GCP, AWS, or Azure and can troubleshoot container issues.
  • •Experience building dashboards/alerts in Grafana and using APM tools such as Datadog, New Relic, or SigNoz; solid understanding of metrics, logs, and traces.
  • •Proficiency in SQL (e.g., PostgreSQL or MySQL), including writing complex queries for debugging and basic database performance knowledge.
Experience:5+ years
Skills:Incident responseTroubleshootingCross-functional coordinationStakeholder communicationAutomation mindset
Tech Stack:JavaGoPythonKubernetesGCPAWSAzureGrafanaDatadogNew RelicSigNozSQLPostgreSQLMySQL

Company Brief

Moniepoint Group
Provides digital payments, merchant services, agent banking, card issuance, and business banking solutions across Nigeria and other African markets, serving SMEs, merchants, and financial institutions with technology-first transaction and banking infrastructure.
Industry: Fintech Infrastructure
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Headquarters: Lagos, Nigeria
Founded: 2015
WebsiteLinkedIn