Site Reliability Engineer

Lucidya
Riyadh
Workplace: RemoteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 3+ yearsSkills: ["Ownership","Incident management","Methodical troubleshooting","Clear communication","Automation mindset"]

Own the reliability of cloud infrastructure and ensure the platform stays stable and scalable as real-time customer data volumes grow. Design fault-tolerant systems, eliminate single points of failure, and manage workloads across AWS, GCP, or Azure using Terraform and Kubernetes. Build observability with tools like Prometheus, Grafana, Datadog, and ELK, drive incident response and root-cause analysis, and automate repetitive operational work to improve deployment reliability.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Lucidya
Lucidya
4 days ago

Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 11 hours agoStatus: Live

Job Summary

Own the reliability of cloud infrastructure and ensure the platform stays stable and scalable as real-time customer data volumes grow. Design fault-tolerant systems, eliminate single points of failure, and manage workloads across AWS, GCP, or Azure using Terraform and Kubernetes. Build observability with tools like Prometheus, Grafana, Datadog, and ELK, drive incident response and root-cause analysis, and automate repetitive operational work to improve deployment reliability.
Location: Riyadh
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Design, maintain, and continuously improve highly available, fault-tolerant, scalable cloud infrastructure.
  • •Identify and eliminate single points of failure, and keep production systems stable under increasing scale and load.
  • •Own and optimize cloud environments and manage workloads across AWS, GCP, or Azure.
  • •Operate Kubernetes in production, troubleshoot issues quickly, and ensure smooth deployments and upgrades.
  • •Build observability, define meaningful alerting, lead root cause analysis for incidents, and automate repetitive operations.

Key Requirements

  • •3 years of experience in SRE, DevOps, or infrastructure engineering at scale.
  • •Hands-on experience with AWS, GCP, or Azure and understanding distributed systems behavior.
  • •Production experience with Kubernetes and the ability to troubleshoot it effectively.
  • •Comfort using Terraform (or similar IaC), Docker/Kubernetes, and CI/CD pipeline concepts.
  • •Implemented monitoring/observability tools such as Prometheus, Grafana, Datadog, or ELK and know how to reduce alert noise.
Experience:3+ years
Skills:OwnershipIncident managementMethodical troubleshootingClear communicationAutomation mindset
Tech Stack:AWSGCPAzureTerraformDockerKubernetesEKSGKEPrometheusGrafanaDatadogELKJenkinsGitHub ActionsBitbucketPythonBash

Company Brief

Lucidya
Provides an Arabic-first customer experience and social listening platform that helps brands monitor conversations, analyze sentiment, and manage customer interactions across digital channels using AI-driven insights.
Industry: SaaS
Company Size: Medium (51 to 250 employees)
Growth: Growth Stage Startup
Headquarters: Riyadh, Saudi Arabia
Founded: 2016
WebsiteLinkedIn