Member of Technical Staff - RL Training Framework

X AI
Palo Alto
Workplace: OnsiteFull timeUSD 180,000 - 440,000Function: Education & TrainingSkills: ["Communication","Problem-solving","Prioritization","Engineering excellence"]

Help build and improve the RL training framework by designing and implementing systems that support RL workloads, from small experiments to production training. Profile, debug, and optimize end-to-end training performance, while enhancing scalability and observability across the RL stack. Work in the RL infrastructure team on large-scale distributed systems, diving into unfamiliar areas and improving efficiency using Python and one or more of JAX, Rust, or C++.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
X AI
X AI
2 months ago

Member of Technical Staff - RL Training Framework

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 6 hours agoStatus: Live
Reposted: similar role first listed 1 year ago

Job Summary

Help build and improve the RL training framework by designing and implementing systems that support RL workloads, from small experiments to production training. Profile, debug, and optimize end-to-end training performance, while enhancing scalability and observability across the RL stack. Work in the RL infrastructure team on large-scale distributed systems, diving into unfamiliar areas and improving efficiency using Python and one or more of JAX, Rust, or C++.
Location: Palo Alto
Workplace: Onsite
Employment Type: Full time
Job Function: Education & Training

Key Responsibilities

  • •Design and implement systems backing all RL workloads, from small-scale ablations to production training runs.
  • •Profile, debug, and optimize end-to-end training performance.
  • •Improve scalability and observability of the RL stack.
  • •Support RL infrastructure development within a small, hands-on team environment.

Pay and Benefits

Salary: USD 180,000 - 440,000
Perks:EquityHealth InsuranceVisionDental401kLife Insurance

Key Requirements

  • •Experience building, debugging, and optimizing efficiency of large-scale distributed systems.
  • •Comfortable diving into unfamiliar areas and solving problems across the full stack.
  • •Proficiency in Python, Jax, Rust, and/or C++.
  • •Ability to build RL training and training infrastructure at scale.
  • •Strong knowledge of reinforcement learning techniques and RL numerics.
Skills:CommunicationProblem-solvingPrioritizationEngineering excellence
Languages:English
Tech Stack:PythonJaxRustC++LLM training infrastructureReinforcement learningRL numericsDistributed systems

Company Brief

X AI
Develops advanced artificial intelligence models and research aimed at building safe, general AI and understanding the fundamental nature of the universe. Focuses on large-scale AI systems, research publications, and building foundational AI capabilities.
Industry: AI & Machine Learning
Company Size: Medium (51 to 250 employees)
Growth: Early Stage Startup
Headquarters: San Francisco, United States
Founded: 2023
Website