AI Eval / Testing (Eval Engineer)
NTT
Dallas
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningExperience: 10+ yearsSkills: ["Analytical","Debugging","Attention to detail","Interpersonal skills","Collaboration"]Build and maintain automated evaluation and testing pipelines for generative AI models and applications. Define AI quality metrics and KPIs (e.g., factuality, safety, bias, latency, cost), implement CI/CD release gates, and create adversarial tests and golden datasets to measure drift. Develop LLM-as-a-judge grading frameworks and tracing for observability, then troubleshoot issues with product and engineering partners to deliver safe, accurate, and trusted AI experiences.

