the case library
AI Agent Evaluation Cases
An agent can give a reasonable answer and still break a requirement. The mistake may be in what it changed, what it kept, or whether it had permission to act.
These synthetic cases each follow one situation: why the behavior first looks acceptable, what breaks when you examine it, and what the agent should do instead. Read them to practice spotting the difference between a convincing result and a requirement being met.
Each case has a permanent number and its own page. The examples are educational; they do not describe incidents in customer systems.
For the foundations, read what AI evals are.