AI Evaluation Engineer (QA)
New York, Austin, Miami, Dallas
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningExperience: 4+ yearsEducation: bachelorsSkills: ["Detail-oriented","Analytical","Communication","Collaboration","Problem-solving"]Build and run large-scale AI/LLM evaluation to prove answer quality improvements over time. You’ll design evaluation and load/quality tests, define test plans and quality gates, and create an automated regression/eval harness integrated into CI/CD. Using statistical methods, you’ll measure factual grounding and accuracy lift (before/after) and report quality metrics clearly to technical and non-technical stakeholders.
Loading
Loading job details...
Preparing the role view and application actions.

