Research Engineer, Infrastructure, RL Systems

Thinking Machines Lab
San Francisco
Workplace: HybridFull timeUSD 350,000 - 475,000 annuallyFunction: Research & Scientific (R&D)Education: bachelorsSkills: ["Collaboration","Debugging","Initiative"]

Design and build core infrastructure for scalable reinforcement-learning training and post-training workloads. You’ll improve reliability and throughput for distributed RL pipelines, develop observability and monitoring tools, and collaborate with researchers to translate RL algorithms into production-grade training systems. The role also includes building evaluation/benchmarking infrastructure for helpfulness, safety, and factuality, and sharing learnings via documentation, open-source, or technical reports.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Thinking Machines Lab
Thinking Machines Lab
1 day ago

Research Engineer, Infrastructure, RL Systems

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 9 hours agoStatus: Live

Job Summary

Design and build core infrastructure for scalable reinforcement-learning training and post-training workloads. You’ll improve reliability and throughput for distributed RL pipelines, develop observability and monitoring tools, and collaborate with researchers to translate RL algorithms into production-grade training systems. The role also includes building evaluation/benchmarking infrastructure for helpfulness, safety, and factuality, and sharing learnings via documentation, open-source, or technical reports.
Location: San Francisco
Workplace: Hybrid
Employment Type: Full time
Job Function: Research & Scientific (R&D)

Key Responsibilities

  • •Design, build, and optimize infrastructure for large-scale reinforcement learning and post-training workloads.
  • •Improve reliability and scalability of RL training pipelines, distributed RL workloads, and training throughput.
  • •Develop shared monitoring and observability tools to support high uptime, debuggability, and reproducibility for RL systems.
  • •Collaborate with researchers to translate algorithmic ideas into production-grade training pipelines.
  • •Build evaluation and benchmarking infrastructure to measure progress on helpfulness, safety, and factuality.

Pay and Benefits

Salary: USD 350,000 - 475,000 annually
Perks:Health InsuranceDentalVisionPaid LeaveParental LeaveRelocation

Key Requirements

  • •Bachelor’s degree (or equivalent experience) in computer science, electrical engineering, statistics, machine learning, physics, robotics, or similar fields.
  • •Strong engineering skills to write performant, maintainable code and debug complex codebases.
  • •Understanding of deep learning frameworks and system architectures (e.g., PyTorch, JAX).
  • •Ability to thrive in a highly collaborative, cross-functional environment with many partners and subject-matter experts.
  • •A bias for action with initiative to work across stacks and teams to ensure shipping.
Experience:Reinforcement learningLarge language modelsDistributed training
Education:Bachelor's
Skills:CollaborationDebuggingInitiative
Tech Stack:PyTorchJAXKubernetesSlurmPrometheusGrafanaOpenTelemetry

Eligibility

Work Authorization:Sponsorship available.

Company Brief

Thinking Machines Lab
Develops enterprise AI solutions, custom large language models, and ML platforms to help organizations deploy intelligent applications. Services include data engineering, model development, and AI consulting for scale and production readiness.
Industry: AI & Machine Learning
Company Size: Medium (51 to 250 employees)
Growth: Growth Stage Startup
Headquarters: Mumbai, India
Website