Data Scientist - ML Research

Arena
San Francisco
Workplace: HybridFull timeFunction: Data Science & Machine LearningExperience: 6+ yearsSkills: ["Communication","Collaboration","Hypothesis development","Experiment design","Causal reasoning"]

Explore and analyze large-scale datasets that power Arena’s AI evaluations, uncovering patterns, biases, and causal relationships in model behavior. You’ll formulate hypotheses, design and validate experiments, and build statistical frameworks and reproducible analysis pipelines to improve reliability and interpretability. Partner with ML researchers and engineers on metrics and analyses across domains and tasks, and communicate insights to technical and non-technical stakeholders.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Arena
Arena
8 months ago

Data Scientist - ML Research

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 16 hours agoStatus: Live

Job Summary

Explore and analyze large-scale datasets that power Arena’s AI evaluations, uncovering patterns, biases, and causal relationships in model behavior. You’ll formulate hypotheses, design and validate experiments, and build statistical frameworks and reproducible analysis pipelines to improve reliability and interpretability. Partner with ML researchers and engineers on metrics and analyses across domains and tasks, and communicate insights to technical and non-technical stakeholders.
Location: San Francisco
Workplace: Hybrid
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Mid level

Key Responsibilities

  • •Explore and analyze large, complex datasets to uncover patterns, biases, and causal relationships in model behavior and system performance.
  • •Formulate hypotheses about data quality, evaluation outcomes, and model performance, and design experiments to validate or refute them.
  • •Build reproducible analysis pipelines using Python, Pandas, NumPy, and Spark to process and interrogate large-scale data.
  • •Collaborate with ML researchers and engineers to design metrics and analyses that evaluate model performance across domains, prompts, and tasks.
  • •Develop causal reasoning frameworks and statistical methods to explain why models behave as they do, and communicate insights to technical and non-technical partners.

Pay and Benefits

Equity and Bonus:Equity
Perks:Health InsuranceDentalVisionEquity

Key Requirements

  • •6+ years of experience in data science, ML analytics, or applied research, preferably in AI/ML or large-scale data environments.
  • •Strong proficiency in Python with deep experience in Pandas, NumPy, and distributed frameworks like Spark.
  • •Expertise in statistical modeling, causal inference, and experimental design.
  • •Experience reasoning about data distributions, sample quality, and distribution shifts.
  • •Strong communication skills and the ability to collaborate closely with ML researchers and engineers.
Experience:6+ yearsAIMachine learningLarge-scale data
Skills:CommunicationCollaborationHypothesis developmentExperiment designCausal reasoning
Tech Stack:PythonPandasNumPySparkA/B testingLLM outputsLLM-as-a-judgeEmbeddings

Company Brief

Arena
LMArena operates a community-driven platform for evaluating and benchmarking large language models via crowdsourced pairwise comparisons and leaderboards, used by researchers and AI labs to measure real-world model performance. ([linkedin.com](https://www.linkedin.com/company/lmarena?utm_source=openai))
Industry: Developer Tools
Company Size: Small (11 to 50 employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series A
Headquarters: San Francisco, United States
Founded: 2025
WebsiteLinkedIn