Senior AI Backend Engineer - Agent Evaluation & Quality
Makkah
Workplace: RemoteFull timeFunction: Software EngineeringSkills: ["System design","Testing mindset","Measurement mindset","Calibration","Collaboration"]Own the evaluation stack for production multi-agent systems, building LLM-as-judge systems, simulators, and per-PR regression harnesses that enforce quality through CI. Calibrate evaluations against human labels, turn production failures into improving test sets, and partner with product to define measurable quality criteria. Grow into agent development by hardening the underlying agents alongside the systems that evaluate them.
Loading
Loading job details...
Preparing the role view and application actions.

