Senior Site Reliability Engineer

Mozn
Cairo
Workplace: RemoteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 3+ yearsSkills: ["Incident response","Root cause analysis","Debugging","Building trust","Stakeholder communication"]

Own hands-on SRE operations while building LLM-based agents that automate reliability toil over time. Carry an on-call rotation, investigate and fix application-level incidents, and ship targeted fixes into application repositories. Design and evaluate safe agent workflows integrated with Kubernetes, cloud APIs, and observability systems (Prometheus/Grafana/ELK/Datadog) using tools like Claude Code and OpenAI Codex, with strong guardrails and auditability.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Mozn
Mozn
2 days ago

Senior Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 1 minute agoStatus: Live

Job Summary

Own hands-on SRE operations while building LLM-based agents that automate reliability toil over time. Carry an on-call rotation, investigate and fix application-level incidents, and ship targeted fixes into application repositories. Design and evaluate safe agent workflows integrated with Kubernetes, cloud APIs, and observability systems (Prometheus/Grafana/ELK/Datadog) using tools like Claude Code and OpenAI Codex, with strong guardrails and auditability.
Location: Cairo
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Carry a normal on-call rotation and serve as a hands-on incident responder for workflows without an agent.
  • •Debug application-level reliability by reading service code, identifying real root causes, and shipping fixes/PRs in application repos.
  • •Design, build, and ship LLM-based agents integrated with Kubernetes, cloud APIs, and observability/alerting systems.
  • •Define agent tool interfaces and build safe wrappers around APIs/scripts with guardrails and human-in-the-loop approvals.
  • •Own agent evaluation using test/backtest suites against historical incidents and report reliability/impact metrics (MTTD/MTTR/MTTX, false positives/negatives).

Pay and Benefits

Perks:Health Insurance

Key Requirements

  • •3+ years building production software with LLMs for agentic workflows (tool/function calling, multi-step planning, RAG).
  • •Hands-on experience shipping work with an agentic coding tool such as Claude Code, OpenAI Codex, or Kimi K2/K3.
  • •Strong Python for building agent tooling, API wrappers, and orchestration.
  • •Real SRE experience as a primary on-call responder: incident response and root cause analysis under pressure.
  • •Solid hands-on Kubernetes and cloud provider experience (AWS/GCP/OCI/Azure) plus observability tools (Prometheus, Grafana, Datadog, ELK).
Experience:3+ yearsLLMsAgentic workflowsProduction softwareKubernetesCloudObservability
Skills:Incident responseRoot cause analysisDebuggingBuilding trustStakeholder communication
Tech Stack:PythonLLMRAGClaude CodeOpenAI CodexKimi K2Kimi K3KubernetesAWSGCPOCIAzurePrometheusGrafanaELKDatadogPagerDutySlackTerraformAnsible

Company Brief

Mozn
Builds enterprise AI products and data platforms that help organizations automate decision-making, improve risk management, and extract insights from large datasets. The company focuses on applied machine learning for regulated and data-intensive industries.
Industry: AI & Machine Learning
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Headquarters: Riyadh, Saudi Arabia
Founded: 2017
WebsiteLinkedIn