Member of Technical Staff - RL Environments

Cohere
London
Workplace: HybridFull timeFunction: Data Science & Machine LearningSkills: []

Build reinforcement learning (RL) environments that replicate real enterprise work, then run and improve AI agents inside them. You’ll design tasks, plausible inputs, and verifiers/reward mechanisms; train and evaluate agents; and work with modeling and product teams to close capability gaps. You’ll also partner with external vendors and automate eval workflows to measure agent performance and iterate on both agents and environments.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cohere
Cohere
1 day ago

Member of Technical Staff - RL Environments

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 11 hours agoStatus: Live

Job Summary

Build reinforcement learning (RL) environments that replicate real enterprise work, then run and improve AI agents inside them. You’ll design tasks, plausible inputs, and verifiers/reward mechanisms; train and evaluate agents; and work with modeling and product teams to close capability gaps. You’ll also partner with external vendors and automate eval workflows to measure agent performance and iterate on both agents and environments.
Location: London
Workplace: Hybrid
Employment Type: Full time
Job Function: Data Science & Machine Learning

Key Responsibilities

  • •Build new RL environments targeting different agentic capabilities and industry areas.
  • •Train and evaluate agents in those environments.
  • •Integrate the full system pieces (tasks, data, tool implementations, verifiers) so they work together.
  • •Collaborate with modeling and product to identify agent performance gaps and improve both agents and environments.
  • •Work with external vendors and create tools to ensure high quality task, data, and verifier quality while automating capability-gap discovery and performance measurement.

Pay and Benefits

Perks:Health InsuranceDental401kPensionParental LeaveLearning BudgetPaid LeaveHome OfficeMeal AllowanceTravel Allowance

Key Requirements

  • •You have engineered and optimized agents for specific industry use cases.
  • •You can review agent trajectories to identify failure points and address them via model training or harness engineering.
  • •You obsess over measuring agentic capabilities and turning it into a repeatable evaluation process.
  • •You translate desired agent outcomes into verifier implementations and tune reward designs.
  • •You can design and run annotation workflows and build synthetic data pipelines to scale eval and training efforts.
Tech Stack:RL environmentsReinforcement learningAI agentsSynthetic dataAgent trajectoriesVerifiersReward designAnnotation workflowsModel trainingAgent evaluation

Company Brief

Cohere
Builds security-first foundation models and enterprise AI products (LLMs, retrieval, agent platforms) for regulated industries, enabling customizable, private deployments across cloud and on-premises for real-world business applications.
Industry: AI & Machine Learning
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series D
Headquarters: Toronto, Canada
Founded: 2019
Glassdoor
Glassdoor: 2.9
WebsiteLinkedInGlassdoor