Senior AI Backend Engineer - Agent Evaluation & Quality

Salla
Makkah
Workplace: RemoteFull timeFunction: Software EngineeringSkills: ["System design","Testing mindset","Measurement mindset","Calibration","Collaboration"]

Own the evaluation stack for production multi-agent systems, building LLM-as-judge systems, simulators, and per-PR regression harnesses that enforce quality through CI. Calibrate evaluations against human labels, turn production failures into improving test sets, and partner with product to define measurable quality criteria. Grow into agent development by hardening the underlying agents alongside the systems that evaluate them.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Salla
Salla
4 weeks ago

Senior AI Backend Engineer - Agent Evaluation & Quality

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

Own the evaluation stack for production multi-agent systems, building LLM-as-judge systems, simulators, and per-PR regression harnesses that enforce quality through CI. Calibrate evaluations against human labels, turn production failures into improving test sets, and partner with product to define measurable quality criteria. Grow into agent development by hardening the underlying agents alongside the systems that evaluate them.
Location: Makkah
Workplace: Remote
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Own and build the evaluation stack, including LLM-as-judge systems and per-agent/per-failure-mode quality measurement.
  • •Create release gates with per-PR evaluation harnesses and regression detection wired into CI.
  • •Develop user simulators to generate test coverage and adversarial cases before production exposure.
  • •Continuously improve evaluations by piping production failures back into evaluation sets.
  • •Collaborate with product to translate “what good looks like” into measurable evaluation criteria and contribute to agent development components as needed.

Key Requirements

  • •5+ years software engineering with recent, hands-on LLM/agent work.
  • •Strong production engineering fundamentals, including Python or TypeScript, clean API/system design, testing, and CI/CD.
  • •Hands-on LLM/agent experience building with agents, RAG, tool/function calling, and orchestration frameworks such as LangGraph or LangChain.
  • •A measurement mindset: metrics, calibration, and experiments to quantify agent performance and failures.
  • •Production experience running LLM systems with reliability, latency, cost, and observability considerations.
Experience:LLM/agentsRAGOrchestration frameworksProduction systems
Skills:System designTesting mindsetMeasurement mindsetCalibrationCollaboration
Languages:Arabic
Tech Stack:PythonTypeScriptCI/CDLLM-as-judgeRAGTool/function callingLangGraphLangChainCI

Company Brief

Salla
Provides an e-commerce platform that helps merchants create, manage, and grow online stores. The product includes storefront setup, payments, order management, shipping integrations, and tools for sales and customer engagement.
Industry: SaaS
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Headquarters: Riyadh, Saudi Arabia
Founded: 2016
WebsiteLinkedIn