Research Engineer (Reinforcement Learning)

LiveKit
North America, EMEA
Workplace: RemoteFull timeUSD 135,000 - 300,000 annuallyFunction: Education & TrainingSkills: ["Collaboration","Experimenting","Data quality focus","Analytical thinking","Risk-aware design"]

Build post-training capabilities for voice and text agents, including training environments, synthetic data pipelines, verifiers, and release evaluations. Own end-to-end training experiments and clearly communicate what improved model behavior. Select and adapt open-weight base models, ensure trained behavior holds up across multi-session voice/text use, and ship models into production while monitoring and iterating from real usage.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
LiveKit
LiveKit
6 days ago

Research Engineer (Reinforcement Learning)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 14 hours agoStatus: Live

Job Summary

Build post-training capabilities for voice and text agents, including training environments, synthetic data pipelines, verifiers, and release evaluations. Own end-to-end training experiments and clearly communicate what improved model behavior. Select and adapt open-weight base models, ensure trained behavior holds up across multi-session voice/text use, and ship models into production while monitoring and iterating from real usage.
Location: North America, EMEA
Workplace: Remote
Employment Type: Full time
Job Function: Education & Training

Key Responsibilities

  • •Build training environments and verifiers used to train post-training models.
  • •Own the synthetic data pipeline from generation through quality gates.
  • •Run training experiments end to end and explain what changed model behavior.
  • •Build evaluations and the criteria a release must pass.
  • •Ship trained models into production and continuously improve them based on real usage.

Pay and Benefits

Salary: USD 135,000 - 300,000 annually
Equity and Bonus:Equity
Perks:Health InsuranceDentalVision

Key Requirements

  • •Strong Python engineering experience.
  • •Experience carrying a model from raw data through to production.
  • •Ability to treat data as the product (coverage, diversity, leakage).
  • •Experience designing against reward hacking/exploitation of weak rewards.
  • •Comfort with GPUs and understanding training limitations.
Experience:Reinforcement learningVoice AILLM post-trainingAgent training
Skills:CollaborationExperimentingData quality focusAnalytical thinkingRisk-aware design
Tech Stack:PythonGPUsVLLMSGLangFSDPGRPOTRLVerlOpenRLHFOpen-weight modelsQwenLlamaLoRAFine-tuningReward designReinforcement learning

Company Brief

LiveKit
LiveKit provides an open-source realtime cloud and developer platform for building, deploying, and scaling voice, video, and physical AI agents with low-latency WebRTC infrastructure and edge services for production voice/video applications.
Industry: Cloud Computing
Company Size: Medium (51 to 250 employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series C
Headquarters: San Jose, United States
Founded: 2021
Glassdoor
Glassdoor: 5.0
WebsiteLinkedInGlassdoor