AI Evaluation
Systematic measurement of AI system behavior against defined expectations -- retrieval, generation, tool selection, workflow execution, and final task success evaluated separately where possible.
Quick help: AI evaluation measures whether an AI system actually performs its intended task. Evaluate components separately where possible: retrieval, generation, tool selection, workflow execution and final task success.
What It Is
AI evaluation is the systematic measurement of AI system behavior against defined expectations. Without evaluation, development often becomes: change prompt -> try example -> looks better. With evaluation: baseline -> defined dataset -> change -> same dataset -> measure -> compare.
Why It Matters
AI output is probabilistic. A change that improves one example may degrade twenty others. Evaluation turns anecdotal impressions into engineering evidence.
Evaluation Layers
Retrieval
Did the system retrieve appropriate evidence?
Generation
Did the model correctly use the evidence?
Agent Decision
Did the agent select appropriate tools/actions?
Workflow
Did execution follow an effective path?
Grounding
Are claims supported by evidence?
Outcome
Was the actual task successfully completed?
Efficiency
What resources were required? Measure: latency, tokens, cost, tool calls, iterations.
Golden Datasets
A golden evaluation dataset contains validated examples: question, expected answer, expected evidence, scoring criteria. The same dataset should be reusable across competing configurations.
Deterministic vs Model-Based Evaluation
Prefer deterministic evaluation where objective rules exist. Examples: expected source retrieved, JSON schema valid, required field present, tool called correctly, forbidden action absent. Some qualitative tasks require model-based or human evaluation. LLM-as-judge can be useful but introduces another probabilistic model into the evaluation system.
Failure Classification
An evaluation should ideally help determine whether a failure is: knowledge, retrieval, reasoning, workflow, tool, behavior, or deterministic rule. This classification guides remediation.
Governance Considerations
Evaluation results can become evidence for promotion decisions: Experimental -> Required evaluations pass -> Risk review -> Deployment candidate. Evaluation thresholds should reflect the use case and risk. A creative internal assistant and an operational agent affecting external systems should not necessarily have identical acceptance criteria.
Practical Experiment
Create 30 questions about the initial KB Sandbox Wiki. For every question define: expected article, expected evidence, expected answer characteristics. Compare two retrieval configurations. Do not change the test set between runs. Measure whether the new configuration actually improves performance.
Last Verified
August 2026