Site Reliability Engineer

Moniepoint Group
India
Workplace: RemoteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 3-6 yearsSkills: ["Incident response","Cross-functional coordination","Analytical thinking","Process improvement","Writing production code"]

Engineer reliability for a highly distributed, hyper-growth platform by defining SLOs and error budgets, leading incident response, and building automation and self-healing mechanisms. Own on-call rotations and act as Incident Commander during major events. Improve observability with high-cardinality metrics and distributed traces, partner with product engineering to bake in reliability patterns from day one, and validate resilience through load testing and chaos engineering.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Moniepoint Group
Moniepoint Group
2 months ago

Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 12 hours agoStatus: Live
Reposted: similar role first listed 3 weeks ago

Job Summary

Engineer reliability for a highly distributed, hyper-growth platform by defining SLOs and error budgets, leading incident response, and building automation and self-healing mechanisms. Own on-call rotations and act as Incident Commander during major events. Improve observability with high-cardinality metrics and distributed traces, partner with product engineering to bake in reliability patterns from day one, and validate resilience through load testing and chaos engineering.
Location: India
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Participate in on-call rotations as the primary technical lead and act as Incident Commander during major severity incidents, coordinating cross-functional teams and providing clear status updates.
  • •Instrument systems with high-cardinality metrics and distributed traces, and define, measure, and defend Service Level Objectives (SLOs) and Error Budgets with product owners.
  • •Write production-ready code in Java, Go, or Python to build internal tooling, automation platforms, and self-healing mechanisms to reduce manual intervention.
  • •Partner with Product Engineering during design to ensure new services include reliability, scalability, and observability patterns such as circuit breakers, rate limiting, backpressure, and fallback strategies.
  • •Analyze performance and traffic patterns to model capacity needs, and run load testing and chaos engineering to verify resilience under failure conditions.

Pay and Benefits

Perks:PensionHealth InsuranceAnnual Bonus

Key Requirements

  • •3–6 years of experience in SRE or backend engineering with strong ability to write clean, performant, tested code in Java, Go, Rust, or Python.
  • •Deep understanding of distributed systems architecture and design patterns, including microservices fundamentals and event-driven architectures.
  • •Extensive experience with Google Cloud Platform (GCP) or similar cloud providers (AWS/Azure), including running production workloads on Kubernetes (GKE/EKS).
  • •Experience designing observability strategies using OpenTelemetry, Prometheus, New Relic, Datadog, or SigNoz.
  • •Familiarity with operating/tuning production data stores (PostgreSQL, MySQL) and streaming platforms (Kafka, RabbitMQ) in high-throughput environments.
Experience:3-6 yearsSREBackend engineeringDistributed systemsMicroservicesEvent-driven architecturesCloudKubernetesObservabilityHigh-throughput systems
Skills:Incident responseCross-functional coordinationAnalytical thinkingProcess improvementWriting production code
Tech Stack:JavaGoRustPythonGoogle Cloud Platform (GCP)AWSAzureKubernetesGKEEKSOpenTelemetryPrometheusNew RelicDatadogSigNozPostgreSQLMySQLKafkaRabbitMQ

Company Brief

Moniepoint Group
Provides digital payments, merchant services, agent banking, card issuance, and business banking solutions across Nigeria and other African markets, serving SMEs, merchants, and financial institutions with technology-first transaction and banking infrastructure.
Industry: Fintech Infrastructure
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Headquarters: Lagos, Nigeria
Founded: 2015
WebsiteLinkedIn