← Zahir Khan · Artifact · August 2026

The Agent Verification Checklist

Ten checks I run against any AI agent's claim of "done" before I believe it. Point it at a pull request, a status report, or an agent transcript. I write this from a software delivery seat, and most of it carries over to any report you are asked to believe.

Every one of these comes back to the same failure. An agent, or a person, or a vendor says a thing is finished, and the only evidence is the sentence itself. That is a claim wearing verification's clothes. So the test I apply to every item here, including the list itself, is whether I can name the exact thing that would prove it false. If I can't, it doesn't belong.

Two things this list refuses to pretend. Evidence costs time, and a standard nobody can afford gets quietly dropped, so I run this against the claims I would have to explain if they turned out to be wrong. And a receipt can lie. Attaching evidence is the start of the argument rather than the end of it, which is why half these items are about the quality of the receipt instead of its existence.

There is one reason agents need this more than people do. When a person overstates what they finished, you can hold them to it afterwards, and that accountability does some of the work for you. An agent gives you nobody to hold. The verification has to carry the whole load on its own.

Tick an item only when you can name the receipt, and write it in the box. An item with a tick and no receipt doesn't count, which is the whole argument of this page applied to the page itself.

0 of 10 verified

Everything you tick or type is saved to this browser only, using localStorage. Nothing is sent anywhere, and reloading the page won't lose it.

Where this comes from

I run a home server and a set of automated workflows for my own projects. Market screeners for my own use, side builds, small sites like this one. The failure mode I keep guarding against isn't the agent doing the wrong thing. It's the agent, or the report, or the pull request, confidently telling me it did the right thing with nothing behind the sentence. The stakes are lower on my own projects than they are in my day job, and the failure is identical. On my own projects I am allowed to talk about it.

About this page

An AI agent drafted this page, working from my practice and my edits, and I ran the list above against it before publishing. Item 2 held up, because the checkbox behaviour here was driven in a real browser with the console watched, rather than described to me in a summary. Item 5 failed on the first attempt. The review I asked for came back from the same model family as the agent that wrote the page, which is the same process agreeing with itself, and it took a second opinion from a different vendor to notice that two items on this list were saying nearly the same thing. Item 7 caught me as well. I had written down what must not appear on this page, then wrote one of those things anyway, and a reviewer found it before I did.

Tell me I'm wrong

If one of these ten is wrong in your experience, I want to know which one and why. Find me on LinkedIn. A list that argues every claim should be falsifiable ought to be willing to take the same treatment.

← Back to zahirkhan.com