Project / new run
Compare a candidate against a baseline
The run executes every case against both versions, scores each output, and ends with a release verdict — including hard gates that can block a high average.
Dataset
Baseline
Candidate
Evaluators
Repetitions per case (stability)