The Test Still Passes. The Behavior Changed.
A stable agent score can hide behavior changes when the suite does not observe the part of behavior that moved.

The score stays the same.
It is easy to read that as stability.
A stable score, however, only describes the observations the suite was built to make.
If behavior changes somewhere outside that view, the test does not turn red. It has no reason to.
This is not a story about one model upgrade or one broken deployment. Treat it as a possible blind spot. The change could come from a model, prompt, tool, retrieval layer, policy rule, workflow condition, or surrounding product system. It may matter. It may not.
The score alone cannot tell you which one is true.
A Stable Score Can Hide Different Case Results
Aggregate scores compress evidence.
That is useful when a team needs a quick view. It is also lossy. A suite can keep the same overall number while the underlying cases move in opposite directions.
Imagine a suite where one group of cases improves and another group declines. The total score may not move much. From a distance, nothing changed. At case level, the behavior did.
That does not mean the suite is wrong. It means the aggregate is answering a narrow question: how did the measured set perform overall under this scoring rule?
It cannot answer a different question: which behaviors changed, in which conditions, and do those changes matter for release?
Those are case-level questions. A stable score can be the beginning of that inspection, not the end of it.
If the changed cases sit near a product boundary, the aggregate can be especially misleading. A small number of changed results may matter more than a large number of unchanged low-risk cases. The practical issue is not only how many cases passed. It is which cases moved and what those cases represent.
The Nugalaxy Evaluation Cases are built around that kind of distinction: a visible verdict is not enough unless the behavior and requirement underneath it are clear.
The Label Can Stay the Same While the Path Changes
There is another quieter version of the same problem.
The pass label can remain unchanged while the agent reaches the answer differently.
A grader might inspect the final answer, the schema, the presence of a citation, the refusal text, or the selected field. If the inspected output still satisfies the rule, the case passes.
But the route to that output may have changed. The agent may rely on a different assumption. It may retrieve from a different source. It may choose a different tool first. It may ask fewer clarifying questions. It may treat a missing field as irrelevant. It may produce the right answer for a reason the grader never checks.
All of those are hypothetical examples. The point is not that they happened. The point is that a pass label only covers the evidence the grader actually inspects.
Sometimes that is enough. If the product only needs the final field and the route has no consequence, the changed path may not matter.
Sometimes it matters a lot. A changed path can reveal that the agent is now depending on weaker context, skipping a useful check, or making an assumption the product does not want it to make.
The stable result does not settle that.
The team has to know what evidence the case preserves. Did the suite capture only the final output? Did it capture the tool call? Did it capture the retrieved source? Did it capture the condition that made the action allowed? Did it capture enough context to compare one run with another?
If not, the test may still pass while the behavior that matters has moved out of view.
Behavior Outside Included Conditions Is Invisible
The third blind spot is simpler.
The suite cannot measure conditions it does not include.
If every case is clean, the suite cannot tell you how the agent behaves when the request is messy. If every case has complete context, the suite cannot tell you what happens when context is missing. If every case assumes the tool is available, the suite cannot tell you what the agent does when the tool cannot answer.
That sounds obvious. It becomes less obvious when the score is stable.
A stable score can make the covered slice feel larger than it is. The team sees continuity in the metric and may infer continuity in the product. But the suite only speaks for the included conditions, not for all plausible conditions around the workflow.
Adding more cases is not enough.
More cases help only when they observe the behavior that matters. A large suite can still be blind if it repeats the same clean condition in slightly different clothing. A smaller suite can be more useful if it preserves the evidence needed to compare meaningful behavior across runs.
The useful question is: what would have to be observable for this score to support the release decision?
For harness-level thinking around what a suite observes and preserves, see the Nugalaxy harness-engineering guides.
What Has to Stay Identifiable
To compare behavior across runs, the team needs more than a final score.
The test input has to be identifiable. Otherwise, the team cannot tell whether the same behavior was tested again or a nearby behavior replaced it.
The grading rule has to be identifiable. Otherwise, a changed verdict might come from a changed standard, not a changed agent.
The relevant system conditions have to be identifiable. Otherwise, a changed result might come from the agent, the retrieval layer, the tool, the workflow state, or the environment around the agent.
This does not require exposing private methods or turning every test into a public artifact. It requires enough internal evidence to know what the result is comparing.
Without that, an unchanged score can become strangely quiet. The number looks stable, but nobody can say whether the same behavior was observed under the same relevant conditions.
That does not prove drift. It means the evidence is too thin to rule drift in or out.
A Release Decision Needs the Observed Delta
The release question is not only whether the suite passed.
It is what changed, what stayed stable, and whether the observed evidence covers the behavior the release depends on.
If the score is unchanged, the team still needs to know whether important cases flipped under the aggregate. It needs to know whether pass labels hide changed paths, assumptions, or dependencies. It needs to know which conditions were never included in the first place.
Sometimes the honest answer will be boring: nothing meaningful changed in the observed evidence.
Sometimes the honest answer will be narrower: the score is stable, but the suite did not observe the behavior we are worried about.
Those are different release positions.
The first can support confidence. The second should produce caution.
A passing suite is useful evidence. A stable score is useful evidence. But neither one proves behavioral stability unless the suite observed the behavior that might have changed.
The test can still pass.
The behavior can still move.
The only way to separate those two facts is to keep the observation boundary visible.