Machine Learning Systems Research Engineer, Agent Post-training - Enterprise GenAI

Scale AI
San Francisco, New York
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningExperience: 1-3 yearsEducation: mastersSkills: ["Communication","Collaboration","Problem-solving","Teamwork","Writing"]

Join Scale AI’s Enterprise ML Research Lab as a Machine Learning Systems Research Engineer focused on building and optimizing the Agent post-training training framework. You’ll post-train state-of-the-art models, collaborate with ML teams on multi-agent RL research, and contribute to scalable GPU cluster architectures for enterprise AI deployments, delivering robust post-training recipes and next-gen agent training algorithms.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Scale AI
Scale AI
10 months ago

Machine Learning Systems Research Engineer, Agent Post-training - Enterprise GenAI

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 19 hours agoStatus: Live

Job Summary

Join Scale AI’s Enterprise ML Research Lab as a Machine Learning Systems Research Engineer focused on building and optimizing the Agent post-training training framework. You’ll post-train state-of-the-art models, collaborate with ML teams on multi-agent RL research, and contribute to scalable GPU cluster architectures for enterprise AI deployments, delivering robust post-training recipes and next-gen agent training algorithms.
Location: San Francisco, New York
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning

Key Responsibilities

  • •Build, profile and optimize the training and inference framework.
  • •Post-train state-of-the-art models, developed both internally and from the community, to define stable post-training recipes for our enterprise engagements.
  • •Collaborate with ML teams to accelerate research and development, and enable them to develop the next generation of models and data curation.
  • •Create a next-gen agent training algorithm for multi-agent/multi-tool rollouts.
  • •Support large-scale training and integrate state-of-the-art technologies to optimize our ML system.

Pay and Benefits

Perks:Health InsuranceDentalVisionRetirement BenefitsLearning BudgetPaid LeaveCommuter Benefits

Key Requirements

  • •1-3 years of LLM training in a production environment
  • •Strong software engineering skills with CUDA, PyTorch, transformers, flash attention
  • •Experience with post-training methods like RLHF/RLVR and related algorithms such as PPO/GRPO
  • •Ability to operate the architecture of a modern GPU cluster and multi-node LLM training/inference
  • •PhD or Masters in Computer Science or a related field
Experience:1-3 yearsGenerative AIEnterprise AIAI research
Education:Master's
Skills:CommunicationCollaborationProblem-solvingTeamworkWriting
Languages:English
Tech Stack:CUDAPyTorchTransformersFlash attention

Company Brief

Scale AI
Provides data labeling, annotation, and infrastructure services to accelerate machine learning and AI development. Supplies high-quality training data, tooling, and APIs for customers in autonomous vehicles, mapping, robotics, and enterprise AI applications.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series E+
Headquarters: San Francisco, United States
Founded: 2016
WebsiteLinkedIn