AI Evaluation Engineer (QA)

Appnovation
New York, Austin, Miami, Dallas
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningExperience: 4+ yearsEducation: bachelorsSkills: ["Detail-oriented","Analytical","Communication","Collaboration","Problem-solving"]

Build and run large-scale AI/LLM evaluation to prove answer quality improvements over time. You’ll design evaluation and load/quality tests, define test plans and quality gates, and create an automated regression/eval harness integrated into CI/CD. Using statistical methods, you’ll measure factual grounding and accuracy lift (before/after) and report quality metrics clearly to technical and non-technical stakeholders.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Appnovation
Appnovation
1 day ago

AI Evaluation Engineer (QA)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: Just nowStatus: Live

Job Summary

Build and run large-scale AI/LLM evaluation to prove answer quality improvements over time. You’ll design evaluation and load/quality tests, define test plans and quality gates, and create an automated regression/eval harness integrated into CI/CD. Using statistical methods, you’ll measure factual grounding and accuracy lift (before/after) and report quality metrics clearly to technical and non-technical stakeholders.
Location: New York, Austin, Miami, Dallas
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Mid level

Key Responsibilities

  • •Run evaluations at scale across large question sets, from human UAT batches to millions of automated evals.
  • •Measure factual grounding and accuracy lift (before/after) using statistical methods.
  • •Build a metrics framework to show how much better answers get over time.
  • •Design load and quality tests, and define/maintain test plans, test cases, and quality gates.
  • •Automate regression and evaluation suites, integrate them into CI/CD, and report results to stakeholders while collaborating with engineering to verify fixes.

Key Requirements

  • •Bachelor’s degree in a technical field or equivalent experience.
  • •4+ years in QA/test engineering with exposure to data/ML systems.
  • •Strong Python and data-science techniques for measuring factual grounding and answer quality.
  • •Experience with LLM evaluation frameworks and statistical analysis.
  • •Ability to build automated, large-scale evaluation harnesses and run lighter human-in-the-loop A/B tests.
Experience:4+ years
Education:Bachelor's
Skills:Detail-orientedAnalyticalCommunicationCollaborationProblem-solving
Tech Stack:PythonCI/CD

Company Brief

Appnovation
Provides digital transformation, design, and engineering services, helping enterprises build customer experiences, digital platforms, and integrations using open-source and cloud technologies across industries.
Industry: Consulting
Company Size: Enterprise (1,001+ employees)
Growth: Established Company
Funding: Bootstrapped
Headquarters: Vancouver, Canada
Founded: 2007
WebsiteLinkedIn