AI Evaluation

Systematic measurement of AI system behavior against defined expectations -- retrieval, generation, tool selection, workflow execution, and final task success evaluated separately where possible.

Quick help: AI evaluation measures whether an AI system actually performs its intended task. Evaluate components separately where possible: retrieval, generation, tool selection, workflow execution and final task success.

What It Is

AI evaluation is the systematic measurement of AI system behavior against defined expectations. Without evaluation, development often becomes: change prompt -> try example -> looks better. With evaluation: baseline -> defined dataset -> change -> same dataset -> measure -> compare.

Why It Matters

AI output is probabilistic. A change that improves one example may degrade twenty others. Evaluation turns anecdotal impressions into engineering evidence.

Evaluation Layers

Retrieval

Did the system retrieve appropriate evidence?

Generation

Did the model correctly use the evidence?

Agent Decision

Did the agent select appropriate tools/actions?

Workflow

Did execution follow an effective path?

Grounding

Are claims supported by evidence?

Outcome

Was the actual task successfully completed?

Efficiency

What resources were required? Measure: latency, tokens, cost, tool calls, iterations.

Golden Datasets

A golden evaluation dataset contains validated examples: question, expected answer, expected evidence, scoring criteria. The same dataset should be reusable across competing configurations.

Deterministic vs Model-Based Evaluation

Prefer deterministic evaluation where objective rules exist. Examples: expected source retrieved, JSON schema valid, required field present, tool called correctly, forbidden action absent. Some qualitative tasks require model-based or human evaluation. LLM-as-judge can be useful but introduces another probabilistic model into the evaluation system.

Failure Classification

An evaluation should ideally help determine whether a failure is: knowledge, retrieval, reasoning, workflow, tool, behavior, or deterministic rule. This classification guides remediation.

Governance Considerations

Evaluation results can become evidence for promotion decisions: Experimental -> Required evaluations pass -> Risk review -> Deployment candidate. Evaluation thresholds should reflect the use case and risk. A creative internal assistant and an operational agent affecting external systems should not necessarily have identical acceptance criteria.

Practical Experiment

Create 30 questions about the initial KB Sandbox Wiki. For every question define: expected article, expected evidence, expected answer characteristics. Compare two retrieval configurations. Do not change the test set between runs. Measure whether the new configuration actually improves performance.

Last Verified

August 2026