Evaluations Engineering - Member of Technical Staff

Simile
San Francisco
Workplace: OnsiteFull timeUSD 200,000 - 400,000 annuallyFunction: Software EngineeringSkills: ["Ownership","Communication","Collaboration","System design"]

Build evaluation infrastructure and tooling for a foundation model that simulates human behavior. You’ll design reliable, reproducible systems to run evaluations across datasets and model versions, strengthen evaluation data schemas with provenance and access controls, and automate validation, survey operations, and human data workflows. Partner across Evals, Modeling, Product Engineering, and Data Operations to deliver execution pipelines, orchestration, and interfaces that help teams compare models and detect regressions.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Simile
Simile
1 month ago

Evaluations Engineering - Member of Technical Staff

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 14 hours agoStatus: Live

Job Summary

Build evaluation infrastructure and tooling for a foundation model that simulates human behavior. You’ll design reliable, reproducible systems to run evaluations across datasets and model versions, strengthen evaluation data schemas with provenance and access controls, and automate validation, survey operations, and human data workflows. Partner across Evals, Modeling, Product Engineering, and Data Operations to deliver execution pipelines, orchestration, and interfaces that help teams compare models and detect regressions.
Location: San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Build evaluation execution infrastructure with services, pipelines, and orchestration to run evaluations across datasets, model versions, populations, and use cases.
  • •Strengthen evaluation data systems by designing relational schemas, versioning, provenance, permissions, and quality controls for reproducible results.
  • •Automate validation and data collection by streamlining customer validations, survey deployment, response ingestion, and ground-truth integration.
  • •Build human data workflows with labeling and review tools for external experts and operators to contribute judgments.
  • •Develop evaluation tooling by building interfaces to manage evals, compare models, investigate results, and identify regressions.

Pay and Benefits

Salary: USD 200,000 - 400,000 annually
Perks:Health InsuranceDentalVisionEquity

Key Requirements

  • •Several years of experience building and maintaining production-quality software with strong system design, testing, debugging, and maintainability.
  • •Experience building backend services, data pipelines, automation workflows, and relational data models.
  • •Ability to take ambiguous projects end-to-end from technical design through deployment and adoption across data, backend, and interface layers.
  • •Strong intuition for evaluation infrastructure reliability, including versioning, provenance, reproducibility, holdout integrity, noisy ground truth, and model comparisons.
  • •Familiarity with modern ML/LLM model-development and evaluation workflows to partner effectively with researchers.
Experience:AI simulationLLM evalsHuman dataBehavioral science
Skills:OwnershipCommunicationCollaborationSystem design

Company Brief

Simile
Develops embedding and similarity search tools to power semantic search, retrieval-augmented generation, and similarity matching for applications. Provides APIs and developer tools to integrate vector search and multimodal retrieval capabilities into products.
Industry: Developer Tools
Website