Member of Technical Staff - ML Research

Arena
San Francisco
Workplace: RemoteFull timeFunction: Data Science & Machine LearningEducation: phdSkills: ["Communication","Collaboration","Problem-solving"]

Machine Learning Scientist role focused on designing and analyzing experiments to evaluate AI models, developing new evaluation protocols, and producing insights to improve model reliability and alignment. You’ll collaborate with engineers and product teams to productionize research findings, work on large-scale models and datasets, and contribute to open scientific publications and the Arena leaderboard.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Arena
Arena
8 months ago

Member of Technical Staff - ML Research

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 5 hours agoStatus: Live

Job Summary

Machine Learning Scientist role focused on designing and analyzing experiments to evaluate AI models, developing new evaluation protocols, and producing insights to improve model reliability and alignment. You’ll collaborate with engineers and product teams to productionize research findings, work on large-scale models and datasets, and contribute to open scientific publications and the Arena leaderboard.
Location: San Francisco
Workplace: Remote
Employment Type: Full time
Job Function: Data Science & Machine Learning

Key Responsibilities

  • •Design and conduct experiments to evaluate AI model behavior across reasoning, style, robustness, and user preference dimensions.
  • •Develop new metrics, methodologies, and evaluation protocols that go beyond traditional benchmarks.
  • •Analyze large-scale human voting and interaction data to uncover insights into model performance and user preferences.
  • •Collaborate with engineers to implement and scale research findings into production systems.
  • •Prototype and test research ideas rapidly, balancing rigor with iteration speed.

Pay and Benefits

Perks:Health InsuranceDentalVisionEquity

Key Requirements

  • •PhD or equivalent research experience in Machine Learning, Natural Language Processing, Statistics, or a related field.
  • •Hands-on experience training large-scale models, including reward models, preference models, and fine-tuning LLMs with methods like RLHF, DPO, and contrastive learning.
  • •Strong foundation in ML and statistics, with a track record of designing novel training objectives, evaluation schemes, or statistical frameworks to improve model reliability and alignment.
  • •Fluent in the full experimental stack, from dataset design and large-batch training to rigorous evaluation and ablation, with an eye for what scales to production.
  • •Deeply collaborative mindset, working closely with engineers to productionize research insights and iterating with product teams to align modeling goals with user needs.
Experience:AIMLNLPModel evaluationResearch
Education:PhD / Doctorate in Machine Learning / NLP / Statistics
Skills:CommunicationCollaborationProblem-solving
Languages:English
Tech Stack:PythonPyTorchJAXTensorFlowTransformers

Company Brief

Arena
LMArena operates a community-driven platform for evaluating and benchmarking large language models via crowdsourced pairwise comparisons and leaderboards, used by researchers and AI labs to measure real-world model performance. ([linkedin.com](https://www.linkedin.com/company/lmarena?utm_source=openai))
Industry: Developer Tools
Company Size: Small (11 to 50 employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series A
Headquarters: San Francisco, United States
Founded: 2025
WebsiteLinkedIn