Member of Technical Staff, North Modelling (Evals)

Cohere
London, Toronto, EU
Workplace: HybridFull timeFunction: Data Science & Machine LearningSkills: ["Self-directed","Practical execution","Data-driven judgment","Failure analysis","Collaboration"]

Build evaluation systems and feedback loops for North, Cohere’s enterprise AI workspace. Own eval strategy across agent workflows, tool use, and knowledge work, translating real usage, production failures, and privacy-preserving logs into high-quality evals. Use eval insights to guide model selection, patches, and regular updates, partnering with product, customer-facing, and modeling teams to define what “good” means for North users.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cohere
Cohere
2 days ago

Member of Technical Staff, North Modelling (Evals)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 1 hour agoStatus: Live

Job Summary

Build evaluation systems and feedback loops for North, Cohere’s enterprise AI workspace. Own eval strategy across agent workflows, tool use, and knowledge work, translating real usage, production failures, and privacy-preserving logs into high-quality evals. Use eval insights to guide model selection, patches, and regular updates, partnering with product, customer-facing, and modeling teams to define what “good” means for North users.
Location: London, Toronto, EU
Workplace: Hybrid
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Mid level

Key Responsibilities

  • •Own North eval strategy for agent workflows, tool use, enterprise knowledge work, and human-AI interactions.
  • •Build high-quality evals from user feedback, production failures, privacy-preserving usage logs, internal dogfooding, and customer needs.
  • •Create systems that continuously convert what the platform learns from users and feature teams into up-to-date evals.
  • •Extract actionable insight from eval results and product context, then turn model failures and capability gaps into recommendations for central modeling teams.
  • •Define what “good” means across the product surface and guide model selection, patches, and regular model updates.

Pay and Benefits

Perks:Health InsuranceDental401kPensionParental LeaveLearning BudgetHome OfficePaid Leave

Key Requirements

  • •Improve LLM-powered, agent-powered, or AI-product systems through evals, feedback loops, data curation, prompting, model adaptation, or model selection.
  • •Apply evaluation craft: representative tasks, precise rubrics, clean data, failure analysis, and regression tracking while avoiding false-confidence metrics.
  • •Use applied MLE judgment to reason about model behavior, eval validity, product outcomes, and production tradeoffs.
  • •Translate messy qualitative signals from users and product teams into measurement other modeling teams can act on.
  • •Work effectively with users and product teams while maintaining self-directed, practical execution on open-ended problems.
Experience:Enterprise AILLMAgentic workflowsMachine learningModel evaluation
Skills:Self-directedPractical executionData-driven judgmentFailure analysisCollaboration

Company Brief

Cohere
Builds security-first foundation models and enterprise AI products (LLMs, retrieval, agent platforms) for regulated industries, enabling customizable, private deployments across cloud and on-premises for real-world business applications.
Industry: AI & Machine Learning
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series D
Headquarters: Toronto, Canada
Founded: 2019
Glassdoor
Glassdoor: 2.9
WebsiteLinkedInGlassdoor