01 / 06 · What counts as good, and how to check it
A trophy engraved BEST, followed by a blank nobody filled in. Better at what, exactly? Evaluation compares behaviour against an expectation written down before the run, and every faster, cheaper check approximates that written judgment.
A single quality score of 7.3 is crossed out: it moves without saying which part moved. Check properties instead. Does it parse, with the fields and types right? Is it supported by the source it cites? Does it decline what it is meant to? Does it hold the register the product committed to?
Four ways to check: deterministic assertions, exact, cheap and narrow; heuristics, a tripwire rather than a definition of correct; model-based graders, judgment at volume from the same probabilistic kind of system; and human review, the one the other three stand in for.
A jar of past failures: wrong fact, wrong tone, missed refusal, malformed, truncated, sure on thin evidence. Change something, rerun the same set, see what moved; only your change moved. Fresh samples are crossed out in favour of the same set. It is not a traffic estimate, just failures kept so their return gets noticed.
A retained metric should say what is measured, on what population, over what period, how it was produced, and which decision changes if it moves. If no decision would change, drop the metric.
A person writes a note: worse means a confident answer with no source. Write it before the change, not after, concretely enough that someone else could agree it happened. Everything else is machinery for checking that description.