EvalRoomshould we ship this?

Project / new run

Compare a candidate against a baseline

The run executes every case against both versions, scores each output, and ends with a release verdict — including hard gates that can block a high average.

Dataset

Baseline

Candidate

Evaluators

Repetitions per case (stability)