Senior Machine Learning Engineer - Model Evaluations, Public Sector

Scale AI
San Francisco, St. Louis, New York, Washington
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningSkills: ["Collaboration","Communication","Problem-solving","Attention to detail","Teamwork"]

Lead the design, implementation, and scaling of automated evaluation pipelines for ML models in public-sector deployments, focusing on functional, performance, robustness, and safety metrics. Collaborate with cross-functional teams to create evaluation datasets, benchmarks, and stress tests for LLMs, agents, and multimodal systems, ensuring reliable, safe operation in defense, intelligence, and federal missions.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Scale AI
Scale AI
9 months ago

Senior Machine Learning Engineer - Model Evaluations, Public Sector

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 18 hours agoStatus: Live

Job Summary

Lead the design, implementation, and scaling of automated evaluation pipelines for ML models in public-sector deployments, focusing on functional, performance, robustness, and safety metrics. Collaborate with cross-functional teams to create evaluation datasets, benchmarks, and stress tests for LLMs, agents, and multimodal systems, ensuring reliable, safe operation in defense, intelligence, and federal missions.
Location: San Francisco, St. Louis, New York, Washington
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Sr. Manager level

Key Responsibilities

  • •Develop and maintain automated evaluation pipelines for ML models across functional, performance, robustness, and safety metrics, including LLM-judge–based evaluations.
  • •Design test datasets and benchmarks to measure generalization, bias, explainability, and failure modes.
  • •Build evaluation frameworks for LLM agents, including infrastructure for scenario-based and environment-based testing.
  • •Conduct comparative analyses of model architectures, training procedures, and evaluation outcomes.
  • •Implement tools for continuous monitoring, regression testing, and quality assurance for ML systems.

Pay and Benefits

Perks:Health InsuranceDentalVisionRetirement BenefitsLearning StipendCommuter BenefitsPaid Leave

Key Requirements

  • •Experience in computer vision, deep learning, reinforcement learning, or NLP in production settings.
  • •Strong programming skills in Python; experience with TensorFlow or PyTorch.
  • •Background in algorithms, data structures, and object-oriented programming.
  • •Experience with LLM pipelines, simulation environments, or automated evaluation systems.
  • •Ability to convert research insights into measurable evaluation criteria.
Skills:CollaborationCommunicationProblem-solvingAttention to detailTeamwork
Languages:English
Tech Stack:PythonTensorFlowPyTorchLLMNLPReinforcement learningComputer visionAWSGCP

Eligibility

Security Clearance:Or ability to obtain a security clearance

Company Brief

Scale AI
Provides data labeling, annotation, and infrastructure services to accelerate machine learning and AI development. Supplies high-quality training data, tooling, and APIs for customers in autonomous vehicles, mapping, robotics, and enterprise AI applications.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series E+
Headquarters: San Francisco, United States
Founded: 2016
WebsiteLinkedIn