AI QA & Evaluation Engineer

Elastic
Bengaluru
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningSkills: ["Attention to detail","Written communication","Collaboration","Technical recommendations","Documentation"]

Design and execute test strategies and rubric-based evaluations for Elastic’s GenAI solutions, ensuring accuracy, reliability, bias, robustness, and regression quality. Build automated validation suites for agentic and multi-agent workflows across CI/CD pipelines, validate data quality and output grounding, and document hallucinations and model performance. Partner with IT, Engineering, Data, Risk, and Compliance to implement governance-ready testing frameworks.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Elastic
Elastic
2 days ago

AI QA & Evaluation Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 5 hours agoStatus: Live

Job Summary

Design and execute test strategies and rubric-based evaluations for Elastic’s GenAI solutions, ensuring accuracy, reliability, bias, robustness, and regression quality. Build automated validation suites for agentic and multi-agent workflows across CI/CD pipelines, validate data quality and output grounding, and document hallucinations and model performance. Partner with IT, Engineering, Data, Risk, and Compliance to implement governance-ready testing frameworks.
Location: Bengaluru
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Mid level

Key Responsibilities

  • •Design and implement comprehensive test strategies for AI/ML systems, including accuracy, bias, robustness, and regression testing.
  • •Create rubric-based evaluation tasks with prompts, supporting files, and grading rubrics to assess AI performance on functional workflows.
  • •Automate validation suites for agentic/multi-agent systems, integration testing, and ML model CI/CD pipelines.
  • •Validate that AI/ML models consume accurate, authorized, and properly structured data sources, ensuring data quality across training and inference.
  • •Observe and report on AI agent behaviors, including documenting performance and hallucinations, and refine evaluation tasks and rubrics based on feedback.

Pay and Benefits

Perks:Health InsurancePaid Leave

Key Requirements

  • •Proficiency in Python, TypeScript, or other programming languages used for AI and test automation.
  • •Experience designing rubric-based evaluation, grading against criteria, or building structured scoring frameworks.
  • •Direct experience with LLM evaluation frameworks and benchmarking tools such as LangSmith and Confident AI.
  • •Knowledge of the GenAI stack, including Retrieval Augmented Generation (RAG).
  • •Experience with cloud platforms (Azure, GCP, AWS) and DevOps/automation/CI/CD tools like GitHub and Terraform.
Experience:GenAIAI/MLCloudDevOpsLLM evaluation
Skills:Attention to detailWritten communicationCollaborationTechnical recommendationsDocumentation
Languages:English
Tech Stack:PythonTypeScriptLangSmithConfident AIRetrieval Augmented Generation (RAG)Azure OpenAIVertex AIChatGPT EnterpriseAzureGCPAWSGitHubTerraformCI/CDDevOpsLLM evaluation frameworks

Company Brief

Elastic
Builds the Elastic Stack (Elasticsearch, Kibana, Beats, Logstash) and provides search, observability, and security solutions that enable organizations to search, analyze, and protect data in real time across applications, infrastructure, and enterprises.
Industry: Data Infrastructure
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Amsterdam, Netherlands
Founded: 2012
WebsiteLinkedIn