Reinforcement Learning Engineer - Ingénieur(e) en apprentissage par renforcement

NBC Universal
Montreal
Workplace: HybridFull timeFunction: Education & TrainingEducation: phdSkills: ["Cross-functional coordination","Attention to detail","Debugging","Mathematical reasoning","Optimization-focused thinking"]

Design and build high-fidelity 2D/3D simulation environments to train reinforcement learning agents. Engineer reward functions and policy architectures, implement and optimize RL algorithms such as PPO, SAC, and Offline RL, and handle high-dimensional 3D observation spaces. Collaborate with ML, annotation, and TPM teams to align data, simulation, and training requirements, and close the sim-to-real “reality gap” using domain randomization and adaptation for reliable real-world performance.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NBC Universal
NBC Universal
2 months ago

Reinforcement Learning Engineer - Ingénieur(e) en apprentissage par renforcement

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live
Reposted: similar role first listed 2 months ago

Job Summary

Design and build high-fidelity 2D/3D simulation environments to train reinforcement learning agents. Engineer reward functions and policy architectures, implement and optimize RL algorithms such as PPO, SAC, and Offline RL, and handle high-dimensional 3D observation spaces. Collaborate with ML, annotation, and TPM teams to align data, simulation, and training requirements, and close the sim-to-real “reality gap” using domain randomization and adaptation for reliable real-world performance.
Location: Montreal
Workplace: Hybrid
Employment Type: Full time
Job Function: Education & Training
Seniority: Mid level

Key Responsibilities

  • •Coordinate with ML, annotation engineers, and TPMs to define data, simulation, and training requirements.
  • •Build and maintain high-fidelity 2D/3D simulation environments (Unity, Unreal, Isaac Sim) for RL agent training.
  • •Design and tune complex reward functions aligned to product goals and safety constraints.
  • •Develop and optimize RL algorithms (PPO, SAC, Offline RL) for high-dimensional 3D observation spaces.
  • •Analyze the sim-to-real reality gap and apply domain randomization or adaptation to improve real-world reliability.

Key Requirements

  • •Master’s or PhD in Robotics, Computer Science, AI, or a related field with a focus on Reinforcement Learning, Imitation Learning, or Online Machine Learning.
  • •Proven experience as an RL Engineer or Research Engineer in a fast-paced environment.
  • •Strong mathematical background for understanding Markov Decision Processes (MDPs) and gradient-based optimization.
  • •Fluency with Python, Git, and Unix shell environments.
  • •Deep familiarity with RL frameworks such as Ray RLlib, Stable Baselines3, or CleanRL, plus experience with physics engines (MuJoCo, Bullet) or 3D game engines.
Experience:RoboticsGame developmentAerospaceReinforcement learningSimulationOffline reinforcement learning
Education:PhD / Doctorate in Robotics, Computer Science, AI, or related field (Reinforcement Learning/Imitation Learning/Online ML)
Skills:Cross-functional coordinationAttention to detailDebuggingMathematical reasoningOptimization-focused thinking
Languages:EnglishFrench
Tech Stack:PythonGitUnix shellRay RLlibStable Baselines3CleanRLUnityUnrealIsaac SimMuJoCoBulletJiraConfluenceSlackExperiment trackingDomain randomizationSim-to-realMDPPPOSAC

Company Brief

NBC Universal
NBCUniversal is a global media and entertainment company producing and distributing film, television, news, sports and streaming content, and operating theme parks and consumer experiences across a portfolio of well-known brands including NBC, Universal Pictures and Peacock.
Industry: Film & Television
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: New York City, United States
Founded: 2004
Glassdoor
Glassdoor: 4.0
WebsiteLinkedInGlassdoor