Senior Machine Learning Research Engineer

Scale AI
San Francisco, New York
Workplace: OnsiteFull timeUSD 216,000 - 270,000 annuallyFunction: Data Science & Machine LearningExperience: 5+ yearsSkills: ["Experimentation","Collaboration","Mentoring","Statistical thinking","Cross-functional communication"]

Own the end-to-end lifecycle of production agent reliability and continuous improvement. Build observability to understand agent behavior in production, design evaluation methodologies and scalable metrics, and develop ML systems that detect drift, anomalies, and misalignment. Run rigorous experiments to validate improvements before deployment, working with software engineers and product, customers, and other teams to translate enterprise and government requirements into dependable platform capabilities.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Scale AI
Scale AI
2 months ago

Senior Machine Learning Research Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 minutes agoStatus: Live

Job Summary

Own the end-to-end lifecycle of production agent reliability and continuous improvement. Build observability to understand agent behavior in production, design evaluation methodologies and scalable metrics, and develop ML systems that detect drift, anomalies, and misalignment. Run rigorous experiments to validate improvements before deployment, working with software engineers and product, customers, and other teams to translate enterprise and government requirements into dependable platform capabilities.
Location: San Francisco, New York
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Mid level

Key Responsibilities

  • •Build and contribute to observability that instruments agent behavior in production to see what agents do.
  • •Design scalable evaluation methodologies and metrics for agentic applications and enable them to run automatically.
  • •Develop and own ML systems to detect drift, anomalies, and misalignment in production agent behavior from prototype to scale.
  • •Design and run rigorous experiments to validate model and agent performance improvements before shipping.
  • •Collaborate with product managers, customers, data annotators, Forward Deployed Engineers, and other teams to translate requirements into platform capabilities.

Pay and Benefits

Salary: USD 216,000 - 270,000 annually
Equity and Bonus:Equity
Perks:Health InsuranceDentalVisionRetirement BenefitsLearning BudgetPaid LeaveCommuter Benefits

Key Requirements

  • •5+ years as an ML engineer or applied scientist on production ML or LLM-powered systems.
  • •Strong grounding in at least two areas: evaluation/monitoring/continuous learning infrastructure, agent system design, or new methods/reward models/model training and fine-tuning.
  • •Hands-on experience with LLMs and agent architectures, including tool use, planning, and multi-agent orchestration.
  • •Ability to partner with software engineers to productionize research and experimental work.
  • •A rigorous experimentation mindset with clear hypotheses and statistically grounded results.
Experience:5+ yearsLLM
Skills:ExperimentationCollaborationMentoringStatistical thinkingCross-functional communication
Languages:English
Tech Stack:LLMsAgent architecturesMulti-agent orchestrationTool usePlanningObservabilityEvaluation frameworksDrift detection

Company Brief

Scale AI
Provides data labeling, annotation, and infrastructure services to accelerate machine learning and AI development. Supplies high-quality training data, tooling, and APIs for customers in autonomous vehicles, mapping, robotics, and enterprise AI applications.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series E+
Headquarters: San Francisco, United States
Founded: 2016
WebsiteLinkedIn