Product Designer, Evals & Prompts

Anthropic
San Francisco
Workplace: OnsiteFull timeUSD 305,000 - 385,000 annuallyFunction: Design (Product/UX/UI/Visual)Education: bachelorsSkills: ["Communication","Collaboration","Attention to detail"]

Own the evaluation layer behind prompts and tools for Claude, designing prompts that drive correct behaviors and building automated graders that measure them across model launches. You’ll create low-code eval tooling for designers, translate rubrics into regression suites, and analyze transcripts to fix prompt issues. Partner with surface owners and engineers to support releases with trustworthy harnesses and prompt migrations grounded in reliable eval results.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Anthropic
Anthropic
2 days ago

Product Designer, Evals & Prompts

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

Own the evaluation layer behind prompts and tools for Claude, designing prompts that drive correct behaviors and building automated graders that measure them across model launches. You’ll create low-code eval tooling for designers, translate rubrics into regression suites, and analyze transcripts to fix prompt issues. Partner with surface owners and engineers to support releases with trustworthy harnesses and prompt migrations grounded in reliable eval results.
Location: San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Design (Product/UX/UI/Visual)

Key Responsibilities

  • •Write and revise prompts behind Claude’s tools, features, and behaviors on a product surface, testing and shipping prompt fixes
  • •Build graders and automated evals from designers’ rubrics, then analyze transcripts and rerun on the next model
  • •Create visual, low-code eval tools so designers can assemble comparison sets, create graders from plain-English rubrics, and review results
  • •Support model releases by testing each surface, writing prompt fixes/migrations, and writing prompts for features launching with the new model
  • •Stand up and scale the eval harness to keep evals green across models and package eval-driven training signals where prompting can’t fix behavior

Pay and Benefits

Salary: USD 305,000 - 385,000 annually
Perks:Paid LeaveParental Leave

Key Requirements

  • •Production-quality Python
  • •Experience building and maintaining evaluation pipelines for LLM products, including graders, rubrics, comparison sets, and regression suites across models
  • •Experience building internal tools with a real interface for people who do not write code
  • •Experience standing up test harnesses, sandboxing tool calls, and pinning settings for comparable runs
  • •Experience shipping prompts (or closely working with people who do) and understanding failures across model versions
Experience:LLMAI safetyToolingModel launches
Education:Bachelor's
Skills:CommunicationCollaborationAttention to detail
Languages:English
Tech Stack:PythonLLMPromptsEvalsGradersRubricsComparison setsRegression suitesTest harnessA/B testing

Company Brief

Anthropic
Develops large-scale AI systems and safety research to create reliable, steerable, and interpretable AI assistants and models for commercial and research applications.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series C
Headquarters: San Francisco, United States
Founded: 2021
WebsiteLinkedIn