AI Evaluation Engineer

SiteGround
Sofia
Workplace: HybridFull timeFunction: Data Science & Machine LearningSkills: ["Data-driven mindset","Experiment design","Root-cause analysis","Safety mindset","Healthy scepticism"]

Own how SiteGround’s LLM-powered AI products perform in production. You’ll build and maintain evaluation datasets with ground truth, define task-specific metrics, run prompt/agent experiments, and ensure changes are measurable—not guessed. Partner with backend engineers and product teams to root-cause quality and cost issues via trace analysis and observability, iterating on prompts, agent behavior, and safety guardrails across OpenAI, Gemini, and Anthropic ecosystems.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
SiteGround
SiteGround
2 days ago

AI Evaluation Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 3 minutes agoStatus: Live

Job Summary

Own how SiteGround’s LLM-powered AI products perform in production. You’ll build and maintain evaluation datasets with ground truth, define task-specific metrics, run prompt/agent experiments, and ensure changes are measurable—not guessed. Partner with backend engineers and product teams to root-cause quality and cost issues via trace analysis and observability, iterating on prompts, agent behavior, and safety guardrails across OpenAI, Gemini, and Anthropic ecosystems.
Location: Sofia
Workplace: Hybrid
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Mid level

Key Responsibilities

  • •Build high-quality evaluation datasets with ground truth and define task-specific metrics (correctness, task completion, efficiency, cost, token usage, safety).
  • •Run model bake-offs and before/after prompt experiments so prompt changes are measured and repeatable.
  • •Design evaluations for hard cases without clean ground truth, including agentic workflows, tool-calling correctness, and structured-output validity.
  • •Own production observability using trace analysis, dashboards, error classification, and session-level review; root-cause failures and translate findings into prompt improvements and engineering tickets.
  • •Design and continuously improve system prompts and agent behavior, including tool orchestration, multi-turn flows, and safety safeguards; select models balancing quality, latency, and cost.

Pay and Benefits

Perks:Annual BonusHealth InsuranceMedical Check-upsFree ParkingMetro ShuttleGym MembershipFree BreakfastMeal AllowancePaid Volunteering

Key Requirements

  • •Shipped LLM-powered features or agents to production (not just prototypes), with strong prompt engineering experience.
  • •Built evaluation datasets and metrics with a data-driven mindset, including measuring results before concluding something works.
  • •Comfort with Python and integrating/consuming REST APIs.
  • •Familiarity with LLM tooling and patterns such as function/tool calling, RAG, agentic workflows, and observability platforms.
  • •Experience with evaluation design/statistics (e.g., A/B testing) and red-teaming/adversarial testing (e.g., prompt injection, jailbreaks, structured-output abuse).
Experience:LLMAI productsObservabilityAgentic workflows
Skills:Data-driven mindsetExperiment designRoot-cause analysisSafety mindsetHealthy scepticism
Tech Stack:PythonREST APIsOpenAIGoogle GeminiAnthropicFunction/tool callingRAGAgentic workflowsObservability platformsLangfuseLangchainN8nGoogle ADKSQLPandasDuckDBNotebooksA/B testing

Company Brief

SiteGround
Provides web hosting, managed WordPress hosting, cloud hosting, and related domain and site management services for businesses and individuals, emphasizing speed, security, and customer support.
Industry: Cloud Computing
Company Size: Large (251 to 1,000 employees)
Growth: Established Company
Headquarters: Sofia, Bulgaria
Founded: 2004
WebsiteLinkedIn