Illustration — We Audited Our Own Catalog Against World-Class. Publicly, Here's What It Found.

Chart · Ideas · Article

We Audited Our Own Catalog Against World-Class. Publicly, Here's What It Found.

We pointed ten fresh-context AI auditors at our own report catalog and asked what world-class means, compared to what. The verdict was stranger than pass or fail.


October 1, 2026 · 8 min read
Share

There's a line in the instructions behind one of our reports that tells the model a data signal "does not exist yet."

For about two days in July, that line was false. The signal existed. We'd shipped a crawler that reads a brand's own published record (what its content actually covers, how often it actually publishes), and the report that most needed those numbers was still being told they weren't available. So it did what it was built to do in that situation: prescribed cold, hedged honestly, and left the measured numbers sitting one step away from the page.

No customer caught it. No engineering review caught it. The system wasn't wrong, it was stale, which is worse, because every automated check still passes. An auditor caught it. One of ten we had just pointed at our own catalog with instructions to compare every report we sell against the best thing money can currently buy.

This is the story of that audit, and I'm telling it in this much detail for a reason that has nothing to do with us: the method transfers. If you run a brand, you have your own version of the word this program was built to interrogate, and by the end you'll have the checklist for interrogating it yourself.

"World-class" is a checklist you run against yourself, not an adjective on your pricing page.


Turn the adjective into a question

We build a catalog of graded brand reports across strategy, creative, content, retention, web, PR, and retail: roughly forty report types, each one a scored, narrated diagnosis of some slice of a consumer brand. And for months, an uncomfortable word had been hovering around it in planning documents: world-class.

I distrust that word. It's on my own banned-phrases list, sitting right next to "game-changer" — because as an adjective it's unfalsifiable. Any pricing page can claim it. Yours can. Your competitors' already do. So we turned it into a question instead: world-class compared to what, exactly?

The program dispatched ten fresh-context AI auditors, one per report discipline, plus two cross-cutting sweeps, one on the competitive market and one on data tooling. Each auditor answered the same five questions for every report in its lane:

What is it

What is this report?

Real bar

What is the actual world-class bar — named competitors, real products, what they measure that we don't?

Gap

Where's the gap, stated plainly?

Fixes ranked

What fixes, ranked?

One input that raises the ceiling

What single brand-supplied input would raise the ceiling most?

The bar had to be real. The auditors benchmarked against the incumbents who own each signal: the consumer-panel firms, the digital-shelf trackers, the usability researchers with proprietary corpora, all the way down to the free 60-second AI audit tools, which are a genuine competitor when your buyer can't tell the difference at a glance.

Two rules made the whole thing worth running, and they're the two rules to steal. First: "already at bar" was a legal verdict. One audit report states its own policy: verdicts of "already at bar" are "stated plainly, not manufactured into findings." An auditor you commission is structurally tempted to justify the engagement by finding problems, so you have to grant permission for the truth to be boring, or every audit becomes a performance. Second: settled verdicts went into a considered-and-rejected ledger, with reasoning, so a settled call can't be re-litigated by the next enthusiastic pass. The program document literally lists its no-action verdicts under the heading "do not manufacture findings against these."


The verdict was stranger than pass or fail

Here is the opening of the resulting program document, verbatim:

"The catalog is already good. Several reports are genuinely world-class today, and the honesty architecture across the whole thing is at or above what any paid competitor ships."

If that were the whole verdict, the program would have been theater. An audit that returns "you're great" is a testimonial. But the second half is the part I'd frame:

"...the pipelines already compute or could cheaply fetch more truth than the rendered report shows the reader. The lift is data and wiring."

Not the prose or the models: wiring. The catalog's biggest quality problem was a class of bug I now think most data products have and few have named: the computed-then-discarded value. The system does the work, derives the truth, and then throws it away before the reader ever sees it.

The findings all had that shape.

What the system already computedWhat the page showed the reader
The growth report received a brand's own stated goals and never read them.The intake form asked what metric mattered most, and the answer arrived at the step that ranks recommendations and sat unused.In the program document's own words: "A brand that says 'my target metric is retention' gets the same generic acquisition-leaning order as one that said nothing."
The creative audit found we named the count, not the rival.The system already computed which competitors shared a brand's visual codes, and had everything it needed to say "2 of 3 rivals we read run the same warm-earth palette.""Your palette echoes the codes your rivals share."

We asked the question and ignored the answer, which is worse than not asking. The specific sentence was already sitting inside the system, unused.

And the sting: a free tool out-displayed us on data we had and they didn't. Our marketing scorecard graded from a real measured signal but collapsed three sub-signals into one number, while a free grader shows four sub-scores from a 30-second surface scan. The audit's verdict, verbatim: "a free competitor shows more granularity within the same footprint than our deeper, better-grounded report does." That sentence stung more than any failure would have.

One finding ran the other direction, a place we were accidentally dishonest in the brand's favor: the email report credited a technical DNS record toward its top deliverability band when, without a paid certificate most brands don't have, that record renders nothing anywhere. Honesty bugs don't only run in the pessimistic direction, and an audit that only hunts the flattering kind isn't an audit.

For symmetry, the auditors also found places the catalog was ahead of the humans: one report labels every channel recommendation as EVIDENCED or INFERRED, and the auditor — who went looking — reported that "no productized or agency deliverable in the research disclosed this distinction." Agencies routinely present a recommended channel mix indistinguishably from a description of what's already running. The audit gets to say that plainly too, because plain was the contract.

The system wasn't wrong, it was stale, which is worse, because every automated check still passes.


The auditors disagreed with each other. Good.

Ten independent passes over one catalog produced two direct contradictions, and the program document's handling of them is my favorite line in it: "flagged, not silently resolved."

The sharpest one: the marketing auditor recommended wiring a major platform's public ad library into our competitor scan, and called it the cheapest big win in the lane. The tooling sweep killed it with a regulation detail: that library only returns ordinary product ads for EU and UK targeting, which means a mostly-US client base's competitor ads are — the sweep's word — "structurally invisible" to it. A plausible, confident recommendation from a rigorous audit, and it would have shipped a near-zero-coverage feature. The only reason it didn't is that a second auditor with a different lens read the same ground. If you run any audit with a single reviewer, this is the finding to remember.

The ledger caught the other tempting mistake. Our capstone report deliberately refuses to emit a single composite brand score, because the famous industry composites blend opaque panel data into one number, and we'd be inventing certainty we don't have. An audit hunting for gaps could easily have filed "no overall score" as a miss versus the incumbents. Instead it filed the refusal under considered-and-rejected: keep it. Sometimes the world-class move is the thing you don't build, written down so nobody helpfully builds it later.


Then we shipped the fixes, and one refusal of our own

An audit that doesn't change the build is a compliance ritual. This one turned into waves of work, and every decision it surfaced got ruled the same day the program document landed. My favorite constraint from those rulings, on a buyer-visible methodology strip, is five words long: "sell the posture, never the machinery." Show the buyer that a rubric governs the grade and every claim is sourced; never expose enough internals that a competitor can reconstruct the grading system.

The first wave went in as one batch: thirteen quick wins across eight report disciplines, almost all of them carrying already-computed values onto the page. And the fourteenth item is the receipt I'd frame. The plan said: add a "what we're declining" section to the capstone, pulled from the refusal lists its constituent reports already produce, pending verification that those lists exist as structured data the system can read. They didn't. They existed only as sentences the model had written.

The audit-remediation batch refused to violate the principle it was remediating toward. That stop, more than any of the thirteen fixes, told me the honesty architecture is load-bearing — it held against us. The review of the batch was framed the same way the audits were, one adversarial pass asking a single question: "could any change let a report claim something the data doesn't back."

The system does the work, derives the truth, and then throws it away before the reader ever sees it.


The checklist you run against yourself

None of this requires AI auditors or forty report types. It requires the five moves.

01
Name the actual best alternatives

Name the actual best alternatives to your product (including the free ones, which your buyer sees first) and write down, per surface, what they show that you don't and what you show that they don't.

02
Grant the truth permission to be boring

An audit where "at bar, no action" is an illegal verdict will manufacture work, and an audit that can't find real gaps is flattery.

03
Keep a rejected ledger

Keep a rejected ledger so settled refusals stay settled.

04
Run more than one lens

Run more than one lens, and treat contradictions between your reviewers as signal to surface, never noise to smooth.

05
Look hardest for the computed-then-discarded class

And when the verdict comes back, look hardest for the computed-then-discarded class, the places your operation already knows something your customer never sees.

In my experience that's where the cheapest quality in the whole company is hiding: not in making the machine smarter, but in delivering what it already knows.

The word "world-class" is still banned in our copy. That was never the point of the exercise. The point is that the word is only honest as a question — compared to what, verified when, by whom? — and a question can be run, dated, and re-run.

"World-class" is a checklist you run against yourself, not an adjective on your pricing page. An adjective is a claim. A checklist is a practice. Only one of them survives an audit — including your own.

The fixes described here are built and reach customers only after a human has walked every page in a live browser — not yet live.


What does "computed-then-discarded" mean?

The system does the work, derives the truth, and then throws it away before the reader ever sees it. In my experience that's where the cheapest quality in the whole company is hiding: not in making the machine smarter, but in delivering what it already knows.

Why let an audit come back with "no action"?

An auditor you commission is structurally tempted to justify the engagement by finding problems, so you have to grant permission for the truth to be boring, or every audit becomes a performance. An audit where "at bar, no action" is an illegal verdict will manufacture work, and an audit that can't find real gaps is flattery.

Why run more than one auditor over the same ground?

A plausible, confident recommendation from a rigorous audit, and it would have shipped a near-zero-coverage feature. The only reason it didn't is that a second auditor with a different lens read the same ground. If you run any audit with a single reviewer, this is the finding to remember.

What is a considered-and-rejected ledger for?

Settled verdicts went into a considered-and-rejected ledger, with reasoning, so a settled call can't be re-litigated by the next enthusiastic pass. Sometimes the world-class move is the thing you don't build, written down so nobody helpfully builds it later.

See your brand

See what Chart measures

Free, across the six major AI engines — what they say about you, where they’re wrong, and where competitors show up instead. A Brand IQ score in a few minutes. No account.

See what Chart measures →