Machine Learning Scientist - Open Source Lead

Arena
San Francisco
Workplace: RemoteFull timeFunction: Data Science & Machine LearningEducation: phdSkills: ["Collaboration","Communication","Problem-solving","Experimentation","Presentation"]

Lead open-source ML research efforts, designing and running experiments to evaluate AI models, curating datasets and reproducible benchmarks, and releasing code to strengthen transparency and community engagement, while collaborating with engineers, product teams, and external researchers to push open research into production-ready workflows.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Arena
Arena
8 months ago

Machine Learning Scientist - Open Source Lead

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 5 hours agoStatus: Live

Job Summary

Lead open-source ML research efforts, designing and running experiments to evaluate AI models, curating datasets and reproducible benchmarks, and releasing code to strengthen transparency and community engagement, while collaborating with engineers, product teams, and external researchers to push open research into production-ready workflows.
Location: San Francisco
Workplace: Remote
Employment Type: Full time
Job Function: Data Science & Machine Learning

Key Responsibilities

  • •Design and conduct experiments to evaluate AI model behavior across reasoning, style, robustness, and user preference dimensions
  • •Develop new metrics, methodologies, and evaluation protocols that go beyond traditional benchmarks
  • •Analyze large-scale human voting and interaction data to uncover insights into model performance and user preferences
  • •Communicate results with the broader research community via academic papers, educational content, conference talks
  • •Collaborate with engineers to implement and scale research findings into production systems

Key Requirements

  • •PhD or equivalent research experience in Machine Learning, Natural Language Processing, Statistics, or a related field
  • •Hands-on experience training large-scale models including reward models, preference models, and fine-tuning LLMs with RLHF, DPO, and contrastive learning
  • •Fluent in the full experimental stack from dataset design to large-batch training and rigorous evaluation
  • •Demonstrated ability to design and analyze experiments with statistical rigor
  • •Experience publishing research or contributing to open-source ML/NLP/AI evaluation projects
Experience:AIMLOpen source
Education:PhD / Doctorate
Skills:CollaborationCommunicationProblem-solvingExperimentationPresentation
Tech Stack:PythonPyTorchJAXTensorFlowLLMsRLHFDPO

Company Brief

Arena
LMArena operates a community-driven platform for evaluating and benchmarking large language models via crowdsourced pairwise comparisons and leaderboards, used by researchers and AI labs to measure real-world model performance. ([linkedin.com](https://www.linkedin.com/company/lmarena?utm_source=openai))
Industry: Developer Tools
Company Size: Small (11 to 50 employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series A
Headquarters: San Francisco, United States
Founded: 2025
WebsiteLinkedIn