Applied AI Researcher, Benchmarking

Distyl
San Francisco, New York
Workplace: HybridFull timeFunction: Research & Scientific (R&D)Skills: ["Problem-solving","Critical thinking","Collaboration","Analytical thinking","Communication"]

Design and run evaluation frameworks for benchmarking intelligent systems, focusing on metrics like reasoning depth, interaction quality, reliability, and operational impact. Develop rigorous methodologies for adversarial testing, longitudinal tracking, and human-in-the-loop assessment to quantify emergent capabilities and set industry standards.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Distyl
Distyl
10 months ago

Applied AI Researcher, Benchmarking

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 5 hours agoStatus: Live

Job Summary

Design and run evaluation frameworks for benchmarking intelligent systems, focusing on metrics like reasoning depth, interaction quality, reliability, and operational impact. Develop rigorous methodologies for adversarial testing, longitudinal tracking, and human-in-the-loop assessment to quantify emergent capabilities and set industry standards.
Location: San Francisco, New York
Workplace: Hybrid
Employment Type: Full time
Job Function: Research & Scientific (R&D)

Key Responsibilities

  • •Define and implement evaluation frameworks that capture benchmarking metrics and reflect real-world complexity
  • •Investigate new paradigms for evaluating intelligent systems, including adversarial robustness testing and human-in-the-loop assessment
  • •Quantify emergent capabilities and establish methodologies that influence internal research priorities and industry standards
  • •Prototype and prototype-driven experimentation to demonstrate effectiveness to senior stakeholders at Fortune 500 level
  • •Collaborate with cross-functional teams to ensure benchmarks align with client problems and business impact

Pay and Benefits

Equity and Bonus:Equity
Perks:MedicalDentalVision401kCommuter BenefitsMeal AllowanceEquity

Key Requirements

  • •Experience designing and running evaluations: built or maintained benchmarks, test suites, or experimental frameworks to measure model or system performance
  • •Statistical and analytical rigor: design fair, reproducible experiments and extract signal from noisy results
  • •Experience building with models, not just building models: expertise in compound AI systems, agentic collaboration, ensemble methods, ReAct, graph-of-thoughts
  • •Proven track record of research results: publications or demonstrable work
  • •Uses AI daily: familiarity with tools like ChatGPT, Cursor, and Perplexity to accelerate workflows
Experience:AIEnterprise AIResearch
Skills:Problem-solvingCritical thinkingCollaborationAnalytical thinkingCommunication
Languages:English
Tech Stack:ChatGPTCursorPerplexityEnsemblingReActGraph-of-thoughtsPrototypingModels

Company Brief

Distyl
Builds enterprise-grade AI systems and integration services (Distillery) to help Fortune 500 companies become AI-native, delivering measurable operational impact across healthcare, telecom, manufacturing, and finance.
Industry: AI & Machine Learning
Company Size: Medium (51 to 250 employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series B
Headquarters: San Francisco, United States
Founded: 2022
WebsiteLinkedInGlassdoor