AI Evaluation Engineer (QA)

Appnovation
Toronto, São Paulo
Full timeFunction: Data Science & Machine LearningExperience: 4+ yearsEducation: bachelorsSkills: ["Analytical skills","Communication skills","Problem-solving","Decision-making","Detail-oriented","Collaboration"]

Own large-scale AI evaluation for answer quality: design and run evaluation suites across small human UAT batches and up to millions of automated tests. Use statistical methods to measure factual grounding and accuracy lift, build a metrics framework for quality improvement, and implement load/quality tests and automated regression integrated into CI/CD. Partner with engineering to reproduce, triage, verify fixes, and continuously improve QA coverage.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Appnovation
Appnovation
1 day ago

AI Evaluation Engineer (QA)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 12 hours agoStatus: Live

Job Summary

Own large-scale AI evaluation for answer quality: design and run evaluation suites across small human UAT batches and up to millions of automated tests. Use statistical methods to measure factual grounding and accuracy lift, build a metrics framework for quality improvement, and implement load/quality tests and automated regression integrated into CI/CD. Partner with engineering to reproduce, triage, verify fixes, and continuously improve QA coverage.
Location: Toronto, São Paulo
Employment Type: Full time
Job Function: Data Science & Machine Learning

Key Responsibilities

  • •Run evaluations at scale across large question sets, from human UAT batches to automated evaluations.
  • •Statistically measure factual grounding and accuracy lift (before/after).
  • •Build and maintain a metrics framework to show quality improvement over time.
  • •Design load and quality tests as the corpus scales, including defining test plans, cases, and quality gates.
  • •Automate regression/evaluation suites and integrate them into CI/CD while reporting metrics to stakeholders and collaborating with engineering to verify fixes.

Key Requirements

  • •Bachelor’s degree in a technical field or equivalent experience.
  • •4+ years in QA/test engineering with exposure to data/ML systems.
  • •Strong Python and data-science techniques for measuring factual grounding and answer quality.
  • •Experience with LLM evaluation frameworks and statistical analysis.
  • •Ability to build automated, large-scale evaluation harnesses and lighter human-in-the-loop A/B tests.
Experience:4+ yearsData/ML systemsLLM evaluationLife Sciences
Education:Bachelor's
Skills:Analytical skillsCommunication skillsProblem-solvingDecision-makingDetail-orientedCollaboration
Languages:English
Tech Stack:PythonCI/CD

Company Brief

Appnovation
Provides digital transformation, design, and engineering services, helping enterprises build customer experiences, digital platforms, and integrations using open-source and cloud technologies across industries.
Industry: Consulting
Company Size: Enterprise (1,001+ employees)
Growth: Established Company
Funding: Bootstrapped
Headquarters: Vancouver, Canada
Founded: 2007
WebsiteLinkedIn