Senior SRE Engineer (MLOps) - AI

Salla
Makkah
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Reliability mindset","Debugging","Cross-team collaboration","Trade-off communication","Operational judgment"]

Own production reliability for SRE/agentic AI and ML systems, building SLOs, alerts, dashboards, runbooks, and running incident follow-ups. Create end-to-end observability across latency, errors, traces, tool calls, cost, and user impact. Design safe release strategies for models, prompts, agents, tools, and configuration, and provide operational support for inference APIs, queues, and retrieval layers on Kubernetes/EKS. Establish guardrails and cost governance for agent tool-calling.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Salla
Salla
2 months ago

Senior SRE Engineer (MLOps) - AI

âś“ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

Own production reliability for SRE/agentic AI and ML systems, building SLOs, alerts, dashboards, runbooks, and running incident follow-ups. Create end-to-end observability across latency, errors, traces, tool calls, cost, and user impact. Design safe release strategies for models, prompts, agents, tools, and configuration, and provide operational support for inference APIs, queues, and retrieval layers on Kubernetes/EKS. Establish guardrails and cost governance for agent tool-calling.
Location: Makkah
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Own reliability for ML and agentic AI services in production, including SLOs, dashboards, alerts, runbooks, and incident follow-ups.
  • •Build observability across the AI stack covering latency, errors, traces, tool calls, cost, and user impact.
  • •Design safe-release patterns for models, prompts, agents, tools, and configuration (canary, rollback, feature flags, evaluation-gates).
  • •Provide operational support for inference APIs, queues, retrieval layers, and AI workflows running on Kubernetes/EKS.
  • •Establish ownership, traceability, guardrails, and security controls for agent tool-calling, including defenses against prompt injection and untrusted-data risks.

Key Requirements

  • •4+ years in SRE, platform engineering, DevOps, or production infrastructure, operating distributed systems in production—not only in demos.
  • •Hands-on production experience with Kubernetes and cloud-native systems.
  • •Familiarity with deploying ML projects.
  • •Strong CI/CD, GitOps, observability, and incident response experience.
  • •Infrastructure-as-code, secrets management, and networking experience; ability to build automation/tooling in Python or a similar language.
Experience:MLOpsML platformsDistributed systemsKubernetesCloud-native
Skills:Reliability mindsetDebuggingCross-team collaborationTrade-off communicationOperational judgment
Tech Stack:KubernetesEKSCI/CDGitOpsObservabilityRunbooksIncident responseInfrastructure-as-codeSecrets managementNetworkingPythonOpenTelemetryPrometheusGrafanaMLflowKServeRayLiteLLMVLLMLangGraph

Company Brief

Salla
Provides an e-commerce platform that helps merchants create, manage, and grow online stores. The product includes storefront setup, payments, order management, shipping integrations, and tools for sales and customer engagement.
Industry: SaaS
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Headquarters: Riyadh, Saudi Arabia
Founded: 2016
WebsiteLinkedIn