Tech Lead Manager- MLRE, ML Systems

Scale AI
San Francisco, New York
Workplace: OnsiteFull timeUSD 252,000 - 315,000 annuallyFunction: Administration & Executive AssistanceSkills: ["Communication","Problem-solving","Cross-functional collaboration","Software engineering","CUDA","PyTorch","Transformers","Flash attention","Distributed systems","RLHF","PPO","GRPO"]

Lead and optimize Scale AI’s ML research platform, overseeing multi-node LLM training/inference and distributed ML system development. Collaborate with ML researchers and engineers to accelerate model research, integration of state-of-the-art techniques, and data curation, while guiding a cross-functional team in building a scalable training framework.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Scale AI
Scale AI
10 months ago

Tech Lead Manager- MLRE, ML Systems

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

Lead and optimize Scale AI’s ML research platform, overseeing multi-node LLM training/inference and distributed ML system development. Collaborate with ML researchers and engineers to accelerate model research, integration of state-of-the-art techniques, and data curation, while guiding a cross-functional team in building a scalable training framework.
Location: San Francisco, New York
Workplace: Onsite
Employment Type: Full time
Job Function: Administration & Executive Assistance

Key Responsibilities

  • •Build, profile and optimize our training and inference framework.
  • •Collaborate with ML and research teams to accelerate their research and development, and enable them to develop the next generation of models and data curation.
  • •Research and integrate state-of-the-art technologies to optimize our ML system.
  • •Lead and mentor a cross-functional team to scale the platform for broader ML workloads and post-training workflows.
  • •Drive system-level improvements and performance tuning for multi-node LLM training/inference.

Pay and Benefits

Salary: USD 252,000 - 315,000 annually
Perks:Health InsuranceDentalVisionRetirement BenefitsLearning BudgetCommuter Benefits

Key Requirements

  • •Strong software engineering skills, proficient in frameworks and tools such as CUDA, PyTorch, transformers, flash attention.
  • •Experience with multi-node LLM training and inference.
  • •Experience with developing large-scale distributed ML systems.
  • •Experience with post-training methods like RLHF/RLVR and related algorithms like PPO/GRPO, etc.
  • •Passionate about system optimization and have strong written and verbal communication skills for cross-functional collaboration.
Experience:AI/MLLLMDistributed systems
Skills:CommunicationProblem-solvingCross-functional collaborationSoftware engineeringCUDAPyTorchTransformersFlash attentionDistributed systemsRLHFPPOGRPO
Languages:English
Tech Stack:CUDAPyTorchTransformersFlash attentionDistributed systemsMulti-node trainingRLHFRLVRPPOGRPO

Company Brief

Scale AI
Provides data labeling, annotation, and infrastructure services to accelerate machine learning and AI development. Supplies high-quality training data, tooling, and APIs for customers in autonomous vehicles, mapping, robotics, and enterprise AI applications.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series E+
Headquarters: San Francisco, United States
Founded: 2016
WebsiteLinkedIn