Software Engineer, Model Evaluation and Improvement

Benchling
San Francisco
Workplace: OnsiteFull timeUSD 136,435 - 166,754 annuallyFunction: Software EngineeringExperience: 2+ yearsSkills: ["Collaboration","Curiosity","Comfort with ambiguity"]

Build datasets, evaluations, and scalable systems that help improve frontier AI models for scientific tasks. Analyze model failure modes by running experiments across leading models, then design and implement pipelines to curate, transform, and validate structured scientific data. Partner with AI labs and work closely with scientists to translate domain judgment into rigorous evaluation criteria that distinguish strong model behavior, all in a fast-evolving area at the intersection of software engineering and biology.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Benchling
Benchling
4 days ago

Software Engineer, Model Evaluation and Improvement

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 6 hours agoStatus: Live

Job Summary

Build datasets, evaluations, and scalable systems that help improve frontier AI models for scientific tasks. Analyze model failure modes by running experiments across leading models, then design and implement pipelines to curate, transform, and validate structured scientific data. Partner with AI labs and work closely with scientists to translate domain judgment into rigorous evaluation criteria that distinguish strong model behavior, all in a fast-evolving area at the intersection of software engineering and biology.
Location: San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Build datasets for evaluating and improving frontier models by converting complex scientific data into high-quality tasks and environments for LLMs.
  • •Analyze model failure modes by running experiments across frontier models to understand where they struggle and identify improvement opportunities.
  • •Build scalable data infrastructure, creating pipelines that curate, transform, and validate large volumes of scientific data into tasks for model evaluation and improvement.
  • •Collaborate with frontier AI labs to develop and evaluate approaches for improving models on challenging scientific tasks.
  • •Work with scientists to translate expert judgment into problems and evaluation criteria that reliably distinguish strong model behavior.

Pay and Benefits

Salary: USD 136,435 - 166,754 annually
Equity and Bonus:Equity
Perks:Equity

Key Requirements

  • •2+ years at the intersection of biology and AI, with experience evaluating and improving scientific models or LLMs for biological applications.
  • •Experience building with LLMs and intuition for where models excel, struggle, and how to design systems around their capabilities.
  • •Curiosity and excitement about frontier AI, with a desire to push rapidly improving model capabilities.
  • •Comfort working on ambiguous problems as technical approaches evolve.
  • •Collaborative mindset working with engineers, scientists, and external research partners.
Experience:2+ yearsBiotechAILLMsBiological applications
Skills:CollaborationCuriosityComfort with ambiguity
Tech Stack:LLMsAI agents

Company Brief

Benchling
Provides a cloud-based R&D Cloud and laboratory informatics platform that centralizes scientific data, streamlines experiments, and accelerates biotechnology and life-science research and development for biotech, pharma, and academic teams.
Industry: Biotech
Company Size: Large (251 to 1,000 employees)
Revenue: USD 100M to 250M
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series E+
Headquarters: San Francisco, United States
Founded: 2012
Glassdoor
Glassdoor: 3.5
WebsiteLinkedInGlassdoor