Back to Case Studies
AIOwn system· Aug 2026

The Quality Gate I Said Never Fired

My AI content system had two automatic quality checks. I audited them, claimed one had never fired, and was wrong. The real lesson was about how I checked, not what I found.

Key Results

  • 38 of 38 scored drafts passed the 9 of 14 bar, with a mean self-score of 12.92
  • 1 voice-check block in 87 checked drafts, after I had first reported 0
  • 2 top-scoring posts (14 of 14) got 466 and 350 impressions, against 460 for an 11 of 14 post
  • 0 comments and 0 reposts for 2 weeks in a row while scores stayed near the top
  • 15 checks now re-measured every week, with a one-line result sent to my phone

What was broken

My content system writes social posts with AI every day. Two checks were meant to stop weak work.

The first was a voice check. It blocks banned phrases and dashes. If a draft fails, it is renamed as failed, set aside, and written again.

The second was a score out of 14. If a draft scored under 9, it was written again, up to twice.

On paper this was safe. In August I audited it. I looked at 143 drafts. 38 of them carried a score. Every one of those 38 scored 11 or more. The average was 12.92 of 14. Not one came near the bar of 9.

That looked too good. And there was a reason. The same model that wrote each draft also scored it. Nothing outside checked the score.

So I compared scores with real results. Only 4 posts could be matched. The two posts that scored 14 of 14 got 466 and 350 impressions. The post that scored 11 got 460. A higher score did not mean a better post. For two weeks in a row, posts got zero comments and zero reposts. Yet the part of the score that rates how much a post invites a reply was near the top.

I held this with care. Four posts is a small number. Reach had also fallen across the platform. But zero comments needs no adjustment.

What I built

First I wrote the finding up with a strong headline: the gates have never fired. My proof was one line. No failed draft had ever existed.

That same day, I found one. A draft from early June carried one banned phrase. The check had caught it. The draft was set aside, marked never to post, and written again. The clean version went out on three channels. The check worked exactly as designed.

So the true number was 1 in 87, not 0. That changes the job. A missing check needs building. A working check that fires once in 87 drafts needs its rules looked at.

The wrong claim had survived five audits. It was a claim that something did not exist, and nobody ran the search that would have shown it did. That search takes two seconds.

So I made a new rule. No check may record a zero without also recording the exact search that produced it.

Then I built a weekly job that re-measures the system from its own files. It runs every Sunday morning. Its last act is one line to my phone: did any number move the wrong way. The first run had 13 checks. By the next day it had 15. 14 of them measure something real. One, score against real results, is marked as not measured, because the reach data is not on disk yet. It is not shown as passing.

I also chose not to switch anything off. Some items were broken. I ruled that we build the record first and fix after. Without a live record, the next five problems stay invisible too.

What changed

The false claim is now corrected, and the old zero is kept on file as a known mistake.

The audit also showed that two of the five voice rules were never built at all. A check that reports two rules while naming the full set claims more than it checks. That is now a tracked item, not a hidden one.

Every number in the weekly job now carries the search behind it. When a number moves, I can see why.

What I would do differently

I would never let the writer grade its own work without an outside check. A score with nothing to test it against is only an opinion.

I would link every score to a real result from day one. The scoring sheet even said when it should be refit: after four more weeks of results. Seven weeks later it had not been, and nothing read the results log except to add to it.

And I would distrust my own zeros. A claim that something never happened is the easiest claim to get wrong, and the cheapest to test.

Want a similar result?

Let's talk about installing the right operating system for your business.

Book a Call
Talk to Anjani