ML Research Engineer, ML Systems

Scale AI
San Francisco, Seattle, New York
Workplace: OnsiteFull timeUSD 218,400 - 273,000 annuallyFunction: Research & Scientific (R&D)Skills: ["Communication","Team collaboration","Problem-solving","Software engineering","Distributed systems"]

Join Scale AI’s ML Platform team to build, profile, and optimize a distributed training and inference framework for large language models. You’ll collaborate with ML researchers to accelerate model development, integrate state-of-the-art technologies, and help shape next-generation LLM training, inference, and data curation across a cross-functional group.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Scale AI
Scale AI
1 year ago

ML Research Engineer, ML Systems

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 5 hours agoStatus: Live

Job Summary

Join Scale AI’s ML Platform team to build, profile, and optimize a distributed training and inference framework for large language models. You’ll collaborate with ML researchers to accelerate model development, integrate state-of-the-art technologies, and help shape next-generation LLM training, inference, and data curation across a cross-functional group.
Location: San Francisco, Seattle, New York
Workplace: Onsite
Employment Type: Full time
Job Function: Research & Scientific (R&D)

Key Responsibilities

  • •Build, profile and optimize training and inference framework for LLMs.
  • •Collaborate with ML teams to accelerate research and development and enable next generation models and data curation.
  • •Research and integrate state-of-the-art technologies to optimize the ML system.
  • •Communicate findings and contribute to cross-functional team discussions to align on platform improvements.
  • •Assist in designing scalable, robust ML infrastructure that supports scalable experiments and evaluation.

Pay and Benefits

Salary: USD 218,400 - 273,000 annually
Perks:Health InsuranceDentalVisionLearning BudgetPaid LeaveTravel Allowance

Key Requirements

  • •Experience with multi-node LLM training and inference
  • •Experience with developing large-scale distributed ML systems
  • •Strong software engineering skills, proficient in CUDA, PyTorch, transformers, and related tools
  • •Ability to communicate effectively and operate in a cross-functional team
  • •Strong excitement about system optimization
Experience:Artificial intelligenceMachine learning
Skills:CommunicationTeam collaborationProblem-solvingSoftware engineeringDistributed systems
Languages:English
Tech Stack:CUDAPyTorchTransformersFlashAttentionLLMDistributed ML systems

Company Brief

Scale AI
Provides data labeling, annotation, and infrastructure services to accelerate machine learning and AI development. Supplies high-quality training data, tooling, and APIs for customers in autonomous vehicles, mapping, robotics, and enterprise AI applications.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series E+
Headquarters: San Francisco, United States
Founded: 2016
WebsiteLinkedIn