01 / 06 · When is an AI system ready for real use?
A small bridge with a sign reading tested with: bicycles, and a lorry marked PROD driving onto it. Ready for what, exactly? Tested, approved and deployed are each worth something, and none of them is the answer. Readiness means the evidence and controls justify this exposure: these users, doing this job, with these consequences when it is wrong.
The same system is fine for a team reviewing its output internally, and not yet for customers acting on it. Same system, different exposure. Naming the exposure precisely, which users doing what with what riding on it, turns an unanswerable question into one with an answer.
The demo: one user, one happy path, little at stake. Real use: malformed input, a slow upstream system, a goal nobody anticipated, and failures that are expensive. Readiness asks about conditions the demo may never have exercised, and silence in the evidence isn't a pass.
A release bar set first, and a lower one moved after seeing the results, which is crossed out: that is fitting the rule to the results. Define the release criterion against an evaluation set that fits the use and its failure costs, before exposure.
Some effects don't come back: a sent message, a payment made. Ask what can be stopped before more happens, reversed, corrected after the fact, or contained to limit how far it travels, and at what cost.
Three answers, in writing: what evidence exists against the failures that matter, what limits the damage when that evidence is incomplete, and who decides and who answers, named rather than assumed. If they can't be answered, nobody can say yet, which is more fixable than unready.