Director, Research - AI Evals

Figma
San Francisco, New York
Workplace: HybridFull timeUSD 258,000 - 348,000 annuallyFunction: Research & Scientific (R&D)Experience: 10+ yearsSkills: ["Communication","Stakeholder management","Strategic thinking","Leadership"]

Own how Figma measures the quality of its AI-powered experiences. Define what “good” means, build evaluation frameworks, rubrics, golden datasets, and quality bars that combine human judgment with automated/model-based methods (including LLM-as-judge). Partner with engineering to create repeatable evaluation pipelines and regression testing, and deliver dashboards/readouts that support go/no-go decisions while leading a small team.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Figma
Figma
1 month ago

Director, Research - AI Evals

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 15 hours agoStatus: Live

Job Summary

Own how Figma measures the quality of its AI-powered experiences. Define what “good” means, build evaluation frameworks, rubrics, golden datasets, and quality bars that combine human judgment with automated/model-based methods (including LLM-as-judge). Partner with engineering to create repeatable evaluation pipelines and regression testing, and deliver dashboards/readouts that support go/no-go decisions while leading a small team.
Location: San Francisco, New York
Workplace: Hybrid
Employment Type: Full time
Job Function: Research & Scientific (R&D)
Seniority: Director level

Key Responsibilities

  • •Own AI evaluation methods and operations for Figma’s AI-powered experiences, defining quality dimensions and turning results into decision-ready signal.
  • •Build and maintain evaluation frameworks, rubrics, golden datasets, and quality bars combining human evaluation with automated/model-based approaches.
  • •Partner with engineering to set up repeatable evaluation pipelines and regression testing so evaluation is built into the AI shipping process.
  • •Create readouts and dashboards that enable go/no-go and prioritization decisions.
  • •Manage a small team and socialize a shared definition of quality so evaluation standards are adopted across teams.

Pay and Benefits

Salary: USD 258,000 - 348,000 annually
Equity and Bonus:Equity
Perks:Health InsuranceDentalVisionRetirementParental LeaveLearning BudgetRemote WorkCell Phone

Key Requirements

  • •10+ years of experience in product, research, applied research, or a closely related field, including 2+ years of management experience.
  • •Direct, hands-on experience owning the evaluation of AI/LLM-powered products.
  • •Expertise designing and running AI evaluations, including human evaluation programs, rubrics, golden datasets, and inter-rater reliability, with judgment on automated/model-based approaches (e.g., LLM-as-judge).
  • •Strength in qualitative and quantitative methods with comfort working with data, metrics, and reasoning about model behavior.
  • •Proven ability to identify the riskiest assumptions behind ambiguous quality questions and design right-sized evaluations with executive buy-in.
Experience:10+ yearsAI/LLMApplied research
Skills:CommunicationStakeholder managementStrategic thinkingLeadership
Tech Stack:AILLM-as-judgeBraintrustLangSmithDeepEvalRegression testing

Company Brief

Figma
Design collaboration platform for teams that enables interface design, prototyping, and real-time collaboration in the browser. Figma streamlines product design workflows, handoff to developers, and design system management for organizations of all sizes.
Industry: Enterprise Software
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Decacorn (USD 10B+)
Funding: Series E+
Headquarters: San Francisco, United States
Founded: 2012
WebsiteLinkedIn