Member of Technical Staff, LLM Evaluation Infra

Inception Labs
San Mateo
Workplace: OnsiteFull timeFunction: Software EngineeringExperience: 2+ yearsSkills: ["Communication"]

The role focuses on designing and building evaluation metrics and scalable evaluation pipelines for large language models, to quantify quality, safety, reliability, and regression performance. You’ll create robust evaluation frameworks, conduct statistical analyses to identify failure modes, and translate real-world use cases into meaningful evaluation criteria, collaborating with product and customer teams to shape model evaluation at production scale.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Inception Labs
Inception Labs
6 months ago

Member of Technical Staff, LLM Evaluation Infra

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 3 hours agoStatus: Live

Job Summary

The role focuses on designing and building evaluation metrics and scalable evaluation pipelines for large language models, to quantify quality, safety, reliability, and regression performance. You’ll create robust evaluation frameworks, conduct statistical analyses to identify failure modes, and translate real-world use cases into meaningful evaluation criteria, collaborating with product and customer teams to shape model evaluation at production scale.
Location: San Mateo
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Design, develop, and maintain robust evaluation frameworks and benchmarks for measuring LLM performance across diverse tasks and domains.
  • •Define and implement quantitative metrics that capture model quality, safety, reliability, and regression detection.
  • •Build scalable, automated evaluation pipelines that integrate into model training and deployment workflows.
  • •Conduct rigorous statistical analysis of model outputs to identify failure modes, biases, and performance gaps.
  • •Partner with product and customer-facing teams to translate real-world use cases into meaningful evaluation criteria.

Pay and Benefits

Equity and Bonus:Equity
Perks:Health InsuranceDentalVisionEquityCommuter BenefitsPaid Leave

Key Requirements

  • •BS/MS/PhD in Computer Science, Machine Learning, Statistics, or a related field (or equivalent experience).
  • •At least 2 years of experience in ML evaluation, applied ML research, or a related engineering role.
  • •Strong understanding of LLM fundamentals (autoregressive generation, instruction tuning, RLHF, in-context learning, decoding strategies).
  • •Proficiency in Python and ML frameworks such as PyTorch.
  • •Experience designing and implementing evaluation metrics and benchmarks for generative models.
Experience:2+ yearsAI/ML
Skills:Communication
Tech Stack:PythonPyTorchDockerGit

Company Brief

Inception Labs
Develops artificial intelligence solutions and research-driven products, focusing on machine learning models, AI tools, and enterprise AI integrations to help organizations automate workflows, extract insights from data, and build intelligent applications.
Industry: AI & Machine Learning
Website