Evaluation
Also known as: evals, eval set
What is evaluation?
Evaluation is the systematic scoring of AI output against a fixed set of test cases with known good answers. It is what turns 'the demo felt good' into a measurable quality baseline you can defend changes against.
Without an eval set, every prompt change is a guess and every model upgrade is a risk. Fifty well-chosen real cases beat a thousand synthetic ones.
Last reviewed by the gAIcko editorial team.