Works On My Prompt•Post-Mortem, Pre-Written•Stage 06 · Evaluation & evidence

Better Than Last Week? Prove It.

After this, you can tell whether a change made your AI feature better or worse, with ten rows in a spreadsheet, before a user tells you.

WORKS ON MY PROMPT · POST-MORTEM, PRE-WRITTENBetter Than LastWeek? Prove It.PRUNINGMYPOTHOS.COMTASTED ONE SPOON. SERVED EVERYONE.SERVEDV2fine.

You changed the prompt, or the model, or what it gets to read. You tried it twice and it looked fine. Looking fine is a feeling about fluent text, and fluent text is exactly what these systems are good at. A short list of real inputs, each with what a passing answer must do, checked before and after the change, turns the feeling into something you can point at.

  1. Step 01

    Before changing anything, write one sentence saying what worse would look like.

    Check: Show the sentence and one output to someone else. They can say yes or no without asking what you meant.

    If it fails: You decide what good meant after reading the output. It agrees with you every time.

  2. Step 02

    Make ten rows of real inputs: ones that went wrong once, or ones you are nervous about. One input per row.

    Check: Every input could be pasted in as it is, and none of them were made up to pass.

    If it fails: Ten polite questions it was always going to get right.

  3. Step 03

    Next to each input, write what a passing answer must do, as something you can see: names the refund window, says no, comes back as valid JSON.

    Check: Two people mark the same output against the same row and agree.

    If it fails: "Should be helpful." Every answer is helpful, including the wrong ones.

  4. Step 04

    Run all ten on the current version and mark each pass or fail. Make the change, run the same ten, mark them again in the next column.

    Check: Two columns of marks on the same rows. You can point at the row that moved.

    If it fails: New examples each time. Something changed, and nobody can say what.

  5. Step 05

    When something breaks for real, add it as a new row the same day.

    Check: The newest row is the newest failure.

    If it fails: The same bug, fixed twice and discussed three times, never written down.

The tools, from their own docs

  • promptfoo

    For: An open-source CLI that runs your test cases against a model on your machine. It can read tests from a CSV file, with an __expected column holding checks such as contains:, is-json or llm-rubric:.

    Does not: Writing the expectations. It checks the ones you give it; llm-rubric hands the judgment to a model, which has to be checked too.

    As of 2026-10-01

  • CSV to Eval (this site)

    For: Turns rows with query, response and expected columns into JSONL, one line per row, in the browser.

    Does not: Running anything or marking pass or fail. It formats rows; you still read the outputs. It labels every expectation contains_phrase, whatever the expectation says. Its own page calls it a local prototype.

    As of 2026-10-01

Where this stops

Ten rows show what moved on those ten. They are no estimate of how often anything goes wrong in real use, and they cannot say whether you wrote the right expectation. That part stays a person's decision, and it is the part worth rereading when a row starts failing.

Why: What Counts as Good, and How to Check It

Sources