Evaluations - Member of Technical Staff

Simile
San Francisco
Workplace: OnsiteFull timeUSD 200,000 - 400,000 annuallyFunction: Data Science & Machine LearningSkills: ["Analytical reasoning","Hands-on ownership","Rapid prototyping","Data-driven decision-making","Rigor"]

Build the measurement layer for Simile’s AI simulations of human behavior, designing evals, metrics, rubrics, datasets, dashboards, and workflows that make simulation quality accurate, trustworthy, and decision-relevant. Partner with modeling teams to diagnose regressions and maintain eval suites, create product/app applied evals, make uncertainty and ground truth legible, and automate end-to-end evaluation systems using agentic coding tools.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Simile
Simile
2 months ago

Evaluations - Member of Technical Staff

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 16 hours agoStatus: Live

Job Summary

Build the measurement layer for Simile’s AI simulations of human behavior, designing evals, metrics, rubrics, datasets, dashboards, and workflows that make simulation quality accurate, trustworthy, and decision-relevant. Partner with modeling teams to diagnose regressions and maintain eval suites, create product/app applied evals, make uncertainty and ground truth legible, and automate end-to-end evaluation systems using agentic coding tools.
Location: San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning

Key Responsibilities

  • •Design the measurement layer for behavioral simulation using evals, metrics, rubrics, datasets, dashboards, and workflows tied to customer decision contexts.
  • •Evaluate model updates and diagnose regressions while maintaining stable eval suites that reflect customer-relevant capabilities.
  • •Build applied/product evals for qualitative outputs, retrieval, survey generation, AI-generated research reports, and other customer-facing surfaces.
  • •Develop rigorous approaches to compare simulated outputs to human data, customer studies, and behavioral datasets, including uncertainty and calibration reasoning.
  • •Automate evaluation workflows by building internal tools, labeling systems, and grader/release-quality pipelines using agentic coding tools.

Pay and Benefits

Salary: USD 200,000 - 400,000 annually
Equity and Bonus:Equity
Perks:Health InsuranceDentalVision

Key Requirements

  • •Strong evaluation taste: can explain what an eval measures, its limits, how it can be gamed, and how it should influence decisions.
  • •Fluency with modern LLM training, post-training, evaluation, and iterative improvement; can interpret model outputs and judge whether changes helped.
  • •Statistical judgment for noisy data, uncertainty, sampling/distributions, calibration, confidence intervals, validity, bias/variance, and population-level interpretation.
  • •Ability to build eval tooling quickly using Python, SQL, R, notebooks, LLM APIs, and agentic coding tools (e.g., Codex, Claude Code, Cursor).
  • •Hands-on ownership: independently drive ambiguous evaluation work end-to-end until the system is reliable and useful.
Experience:LLM evalsApplied MLResearch engineeringBehavioral scienceHuman data
Skills:Analytical reasoningHands-on ownershipRapid prototypingData-driven decision-makingRigor
Tech Stack:PythonSQLRNotebooksLLM APIsCodexClaude CodeCursor

Company Brief

Simile
Develops embedding and similarity search tools to power semantic search, retrieval-augmented generation, and similarity matching for applications. Provides APIs and developer tools to integrate vector search and multimodal retrieval capabilities into products.
Industry: Developer Tools
Website