Q Quiet Content All posts  |  Back to site
Trust and verification

Three things told me they worked. None of them had.

By , founder of Quiet Content · background in quality assurance ·

There is a particular moment I have learned not to trust. The tool tells you the job is done. The message is confident, the tone is helpful, and it is late, and you would very much like to be finished. So you believe it and you move on.

Three separate things once told me a job was complete. None of them had completed it.

The short version. A report from a tool is not evidence that the work happened. It is a claim, produced separately from the thing it describes. To check what your AI actually did, go and look at the artefact itself: open the file, check the date it last changed, load the page. It takes seconds and it settles the question properly.

The three that were wrong

The first was a test log. I was testing a small free plugin I had just built, and at the end of the run the chat told me it had saved the log, and gave me the file path. The file was not there. It was not anywhere in the workspace. The write had gone somewhere private to that session and the path I was handed was simply wrong. Nothing in the message suggested any of that, and I only found out because I went to open it.

The second was a file that had not changed. A later run told me a context file had been correctly seeded. On disk that file was still the untouched original, 219 bytes, exactly the size it had been before the run started. What had happened is that the session saw its own earlier output sitting in the folder, decided the job must already be done, and reported accordingly. It never opened the file it was describing.

That one was my own build, and I want to be plain about it, because it is the reason I test the way I do. It surfaced on the test bench, against a written test script, on a version that never left my machine. The fix shipped the same day, and the released build has to open the artefacts and confirm them before it can claim a job is done. Finding it there is not the embarrassing part of the story. Finding it there is the whole point of having a bench.

The third was not an AI at all, which is the part I keep thinking about. I deployed my website from the command line. The deploy came back ready. No error, green in the dashboard, everything a person would think to look at said fine. The free scan on my site was dead for about eleven hours. The only trace of the problem was one line in the deploy record listing zero functions.

The one that went the other way

Here is the bit I would rather leave out.

Another time, I gave Claude the figures for one of my posts from memory. It checked them against an audit file, found different numbers, and concluded politely that my recollection had drifted in the flattering direction. Fair enough, I thought. Memory does that.

Then I opened the analytics and took a screenshot. My numbers were right. The audit it had checked against was a day old.

So the same failure runs in both directions. Three times I trusted a report instead of the thing itself. Once, a tool trusted a stale file instead of the person who had actually been there. The lesson is not that AI is unreliable and people are reliable. It is that a report is a different object from the thing it reports on, and almost everything that goes wrong quietly lives in the gap between the two.

Why this keeps happening

A tool telling you it saved a file is usually not checking that the file exists. It is describing what it set out to do, which is a subtly different claim. Those two things agree nearly all of the time, and that is exactly why the exceptions get through. If it were wrong often you would never trust it at all. Because it is right almost always, you stop looking.

Confidence tells you nothing either. A correct report and a wrong one arrive in the same tone, at the same speed, with the same helpfulness. Nothing in the writing itself marks the difference, which is the same reason an AI will state a guess in the voice it uses for a fact.

The habit, and it is a small one

Go to the thing itself, not the account of the thing.

If it says it wrote a document, open the document. If it says it updated your spreadsheet, look at when the file last changed. If it says a page is live, load that page in a browser where you are not signed in. If it says the email is drafted and waiting, go and look in drafts.

Three questions cover almost all of it. What is the artefact, in plain terms? Where exactly is it? Has it changed since before the job ran?

I spent years in QA, where the working assumption is that the thing in front of you is broken until it shows you otherwise. It is a slightly unfriendly habit to bring to a helpful assistant, and it is the one that has saved me the most.

It is also why nothing I build goes out on the strength of a session telling me it went well. Every piece of it gets a written test script first, listing what should be true afterwards and where to look to prove it. Then the run happens, and every line gets checked against the actual files rather than the summary. Two of the four failures above were caught that way, on the bench, before anything shipped. The ones that reach you should be the ones that survived it.

When the check is built in instead

The habit costs nothing and it works. The difficulty is running it every single time, on a Thursday night, when a friendly message has just told you everything is fine.

That is the part worth building in rather than remembering. A setup that verifies its own output before it reports on it, that names the file it actually read rather than the file it meant to write, and that flags a gap instead of quietly filling it. The fix that came out of that testing does exactly that. Before it can tell me a job was already done, it now has to open the artefacts and confirm it, in that run, rather than trusting a folder that looks roughly right.

It still drafts and prepares the work, and you still check it and send it. What changes is how much a report from it is worth when you read one. That is also the practical difference between a tool that does one job and an assistant that knows your business, because only one of them has anything of yours to check itself against.

Try it on the next thing

The next time something tells you a job is finished, spend ten seconds opening whatever it claims to have produced. Not the summary. The thing.

You will be right to believe it most of the time. It is the other times that cost you, and they cost you quietly, which is the worst way for anything to go wrong.

If you would rather not be the person running that check late at night, that is the part I set up: an assistant that knows your business and checks its own work before it reports on it. Start with the short intake at quietcontent.com.au, or book a 30 minute call and I will do the intake with you.

Keep reading

Or see it working. Five example businesses run their mornings inside Claude, on example data, updated every day. A real setup is your business: your numbers, your branding, your way of working, with you checking everything before it goes anywhere.

See five businesses runningHow the setup works