Senior Data Scientist, AI Scoring & Evaluation

Workera
Anywhere
Workplace: RemoteFull timeFunction: Data Science & Machine LearningExperience: 4+ yearsSkills: ["Independent execution","Analytical thinking","Statistical analysis","Communication","Autonomy"]

Own the end-to-end scoring trust layer that turns signals into defensible, fair skill data for Workera’s assessment-driven AI. You’ll design and continuously evaluate scoring quality (calibration, bias diagnostics, drift detection, and pre-release gates), publish KPIs, and lead investigations when scores look wrong. Work cross-functionally with assessment science and the scoring pipeline teams, and translate AI governance and InfoSec frameworks (e.g., GDPR, the EU AI Act) into audit-ready evidence.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Workera
Workera
2 days ago

Senior Data Scientist, AI Scoring & Evaluation

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 11 hours agoStatus: Live

Job Summary

Own the end-to-end scoring trust layer that turns signals into defensible, fair skill data for Workera’s assessment-driven AI. You’ll design and continuously evaluate scoring quality (calibration, bias diagnostics, drift detection, and pre-release gates), publish KPIs, and lead investigations when scores look wrong. Work cross-functionally with assessment science and the scoring pipeline teams, and translate AI governance and InfoSec frameworks (e.g., GDPR, the EU AI Act) into audit-ready evidence.
Location: Anywhere
Workplace: Remote
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Mid level

Key Responsibilities

  • •Own scoring quality end to end, including evaluator design, rubric anchoring, calibration, and accuracy benchmarking.
  • •Build and run a continuous evaluation harness with gold sets, bias diagnostics, drift detection, and a pre-release gate for scoring changes.
  • •Define and publish assessment quality KPIs on dashboards (e.g., reliability, classification accuracy, bias indicators, latency, cost).
  • •Lead analytics investigations, including performance studies, impact simulations, and root-cause analysis when scores are wrong.
  • •Coordinate measurement improvements end to end with engineering (spec, validation, and quality bar), and serve as a liaison between AI governance and InfoSec for audit evidence.

Key Requirements

  • •Drive production data science work independently.
  • •4+ years (or 3+ with demonstrated end-to-end ownership) on production data projects, ideally in an AI startup or fast-moving product environment.
  • •Proficiency with Python and SQL for running analyses.
  • •Strong statistical/analytical skills to design defensible studies (e.g., scoring reliability, accuracy).
  • •Experience working with LLM-based systems in production, including prompt design and evaluating outputs against human judgment.
Experience:4+ yearsAIData scienceLLMAssessmentEducational/learning data
Skills:Independent executionAnalytical thinkingStatistical analysisCommunicationAutonomy
Languages:English
Tech Stack:PythonSQLLLMsPrompt designDashboardsDrift detectionCalibrationBias diagnosticsGDPREU AI Act

Company Brief

Workera
Provides AI-driven skills assessment and personalized learning pathways for data science, machine learning, and AI professionals and teams, helping organizations identify skill gaps and recommend tailored training to accelerate capability development.
Industry: Corporate Training
Website