AI Engineer, Model Quality and Performance

Cerebras
Sunnyvale
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningSkills: []

Own model quality and performance for inference offerings by defining what “good” means and measuring it at scale. Build AI-agent-driven eval suites per model release and customer use case, mining trajectories and synthesizing representative test sets. Automate end-to-end eval execution and release qualification using pipelines on Docker, Git, and CI. Forecast and benchmark production performance for top customers and create product-quality tooling that unifies quality and performance signals.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cerebras
Cerebras
4 months ago

AI Engineer, Model Quality and Performance

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 16 hours agoStatus: Live

Job Summary

Own model quality and performance for inference offerings by defining what “good” means and measuring it at scale. Build AI-agent-driven eval suites per model release and customer use case, mining trajectories and synthesizing representative test sets. Automate end-to-end eval execution and release qualification using pipelines on Docker, Git, and CI. Forecast and benchmark production performance for top customers and create product-quality tooling that unifies quality and performance signals.
Location: Sunnyvale
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning

Key Responsibilities

  • •Define model quality and performance criteria and translate measurement signals into artifacts customers and product can use.
  • •Design eval suites with AI agents in the loop for each model release (advanced, basic, long-context, and customer-specific evals).
  • •Build custom evals for target customers by orchestrating agents to mine trajectories and synthesize representative eval sets.
  • •Automate eval execution end-to-end using AI-driven pipelines built on Docker, Git, and CI so the system runs between releases.
  • •Build automations and benchmarking workflows to forecast model performance for top customers, including production runtime for customer workloads, and provide unified quality/performance tooling.

Key Requirements

  • •Experience building AI agents and shipping real systems using Claude (or equivalent).
  • •Strong math and statistics background.
  • •Comfort with Docker, Git, and standard automation tooling (including CI).
  • •Experience designing eval suites for agentic, coding, long-context, and/or multimodal use cases.
  • •Familiarity with open-source eval frameworks such as EvalScope and lm-eval-harness.
Experience:AIModel evaluationAI agentsPerformance benchmarkingOpen source
Tech Stack:ClaudeDockerGitCIEvalScopeLm-eval-harnessGPUsFPGAsCustom silicon

Company Brief

Cerebras
Designs and builds wafer-scale AI accelerators and systems for large-scale deep learning workloads, delivering specialized hardware and software to accelerate model training and inference for enterprises and research institutions.
Industry: Hardware Devices
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Headquarters: Sunnyvale, United States
Founded: 2016
WebsiteLinkedIn