about

Evaluation Kept Failing Before the First Test Ran.

Nugalaxy did not start as an evaluation practice. It started with AI systems that had to work, and with the ordinary problem of showing that they did.

The failures were never where we went looking for them. The scoring was fine. The tooling was fine. What was missing sat further upstream: nobody had written down what correct behavior actually meant. So the evaluation measured whatever was easy to measure, the result came back green, and the requirement that mattered was never turned into a test at all.

Requirement

What the system was actually supposed to do, never written down.

Everything below inherits the gap
CriteriaTest casesExpected behaviorScore
The number comes back green. The requirement was never tested.

That gap is the whole of what this practice works on. We define what correct means first, then build the criteria, cases, harness, and evidence that make an evaluation result mean something to the person who has to sign off on it.


What We Actually Believe.


Small, On Purpose.

Nugalaxy is small, and the person writing the guides is the person doing the engineering. That is deliberate. The methodology stays specific because it comes out of real systems rather than a content calendar, and the work we take on feeds it directly. Nothing here is generalized from someone else's whitepaper.

It is also why harness engineering is written out in public rather than kept as a house method. A discipline nobody can read is not a discipline.


Computed, Not Drawn.

The same rule runs all the way down to the logo. Every planet in this universe is generated from math, not illustrated by hand. The cyan caltrop you see is a shape our own code produces. Form follows the engineering, with nothing drawn on top to hide it, which is the same reason we hand over evidence instead of assurances.

Tell Us What You Need to ProveRead the guides