the case library

AI Agent Evaluation Cases

An agent can give a reasonable answer and still break a requirement. The mistake may be in what it changed, what it kept, or whether it had permission to act.

These synthetic cases each follow one situation: why the behavior first looks acceptable, what breaks when you examine it, and what the agent should do instead. Read them to practice spotting the difference between a convincing result and a requirement being met.

Each case has a permanent number and its own page. The examples are educational; they do not describe incidents in customer systems.

For the foundations, read what AI evals are.

Published cases