Forward Deployed Machine Learning Engineer

Protege
United States
Workplace: RemoteFull timeFunction: Data Science & Machine LearningSkills: ["Bias to action","High ambiguity tolerance","Strong written communication","Urgency"]

Build and scale Protege’s Benchmarks and Evaluations vertical by defining strong evaluation criteria, designing and implementing benchmarks with researchers, and creating the infrastructure to run and measure model performance. Own backend components including data pipelines, execution environments, storage, and orchestration, plus sandboxed environments for agentic evals. Turn live engagements into repeatable eval patterns and customer-ready engineering work end to end.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Protege
Protege
21 hours ago

Forward Deployed Machine Learning Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

Build and scale Protege’s Benchmarks and Evaluations vertical by defining strong evaluation criteria, designing and implementing benchmarks with researchers, and creating the infrastructure to run and measure model performance. Own backend components including data pipelines, execution environments, storage, and orchestration, plus sandboxed environments for agentic evals. Turn live engagements into repeatable eval patterns and customer-ready engineering work end to end.
Location: United States
Workplace: Remote
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Mid level

Key Responsibilities

  • •Partner with the GM and early customers to define what strong evals look like across domains.
  • •Design and build benchmarks with Protege researchers.
  • •Build the backend infrastructure for the vertical (data pipelines, execution environments, storage, orchestration).
  • •Stand up sandboxed environments for agentic evals requiring tools, code execution, or multi-step tasks.
  • •Identify repeatable eval patterns and product opportunities from live engagements, then ship iterations of the eval infrastructure.

Key Requirements

  • •4+ years of engineering experience.
  • •Hands-on ML work evaluating models.
  • •Have previously owned backend and infrastructure.
  • •Comfort working with urgency to meet market pace and volume.
  • •Strong written communication.
Experience:AIMachine learningBenchmarksEvaluationsLLMsResearch orgsFrontier labsAgentic systems
Skills:Bias to actionHigh ambiguity toleranceStrong written communicationUrgency
Tech Stack:Machine learningLLMsBenchmarksEvalsData pipelinesExecution environmentsStorageOrchestrationAgentic systemsRL environmentsCode execution sandboxes

Company Brief

Protege
Provides a privacy‑first platform that connects data holders with vetted AI developers to license and deliver high‑quality, multimodal training and evaluation datasets across healthcare, media, audio, and motion capture.
Industry: Data Infrastructure
Company Size: Small (11 to 50 employees)
Growth: Growth Stage Startup
Funding: Series A
Headquarters: New York City, United States
Founded: 2024
WebsiteLinkedIn