Skip to main content
    AI Engineering

    Evaluation

    Also known as: evals, eval set

    What is evaluation?

    Evaluation is the systematic scoring of AI output against a fixed set of test cases with known good answers. It is what turns 'the demo felt good' into a measurable quality baseline you can defend changes against.

    Without an eval set, every prompt change is a guess and every model upgrade is a risk. Fifty well-chosen real cases beat a thousand synthetic ones.

    Last reviewed by the gAIcko editorial team.