There's a comment at the top of the claim-verification service in our report pipeline, written for the engineers who maintain it: "a URL can be real while the specific claim is unsupported."
It's there because that is the precise way confident documents lie. A report can cite a source that exists and still assert something the source never said — and nothing about the prose tells you which sentences are which.
We built a machine to catch exactly that, which brings me to the question every buyer of AI-generated strategy should be asking, and the one this piece exists to take seriously:
Why would I pay for AI strategy when ChatGPT is free?
It's the right question. Most of the AI-report industry deserves it — a lot of what's sold as "AI strategy" today is a prompt in a trench coat, and it should die. (That takedown is its own essay.) But the question has a real answer, and the answer isn't "our AI is smarter." It's that a strategy answer and a strategy artifact are different things, and the difference is built out of machinery you can inspect.
I build in this category, so discount everything that follows accordingly. I'll try to earn back the discount with specifics.
What the free answer actually is
When you ask a general model for strategy, you get the statistical middle of your category. That's not a flaw — it's the design. The model was trained on everyone, so its default answer is the average of everyone. Ask it for a protein bar positioning and you'll get the positioning of the median imaginable protein bar: clean ingredients, busy lifestyles, fuel your day. Grammatically perfect. Strategically nothing — because your competitors' interns are getting the same paragraph from the same model this afternoon.
And there's a second problem underneath — the one we measured. We asked three frontier models one question about one small consumer brand: who founded it? We ran exactly this, on a real brand, and kept the dataset.
Across the full question set, the engines agreed with each other 28% of the time, and their fidelity scores — how close each engine's picture of the brand came to reality — were 36, 20, and 32 out of 100.
That's the second problem in numbers: the model's knowledge of you is inversely proportional to how much you need help. AI already knows the famous brands — they're saturated in the training data. Everyone else, it guesses about. The smaller you are, the more of "you" in the model's answer is invention. So the free strategy answer for a small brand is built on a foundation that is partly made up, and — this is the part that should bother you — made up fluently, with no signal distinguishing the recalled facts from the invented ones.
The invented founder didn't come with an asterisk.
That's the free answer: category-average reasoning applied to a partially fictional version of your company. For orientation, it's genuinely useful. As the input to a real decision — a repositioning, a channel bet, a pricing change — it's a confident guess about a company that only half-exists.
What you're paying for is not the words
Here's the reframe that took me too long to reach, even building these systems: when you buy strategy — from a consultancy, from an agency, from an AI — you were never really buying the document. You were buying the process that makes the document safe to act on. Research that actually happened. Claims that were checked. Somebody accountable for the difference between "true" and "plausible."
Consulting firms priced that process at six figures because humans doing it are slow. AI made the document nearly free — but the document was never the value. Research got cheap. Judgment didn't. The entire question of whether AI strategy is worth buying collapses to: did anyone build the judgment machinery, or did they just generate the deck?
Research got cheap. Judgment didn't.
So let me show you what judgment machinery physically is. These mechanisms are from the pipeline my team builds — not because ours is the only way, but because I can show receipts for it, and because each mechanism exists downstream of a specific failure.
Before a report generates anything, the brand exists as structured data — in our system, 346 signals of verified identity, positioning, product, and competitive evidence, extracted from what the brand has actually made. That number matters less than what it replaces: six adjectives in a prompt box. A general model reasons about the average brand; a grounded system reasons about your signals. This is the difference between "AI wrote about a protein bar" and "AI wrote about yours."
Every report in our pipeline derives its own research questions, fans them out to isolated research agents, retrieves sourced material, and persists every source it touched before a word of strategy gets written. The finished artifact closes with a Methodology and Sources section. If a strategy document can't tell you where its facts came from, it doesn't have facts. It has vibes with citations pending.
This is the mechanism I'd defend in a fight, and it's the one the comment I opened this piece with belongs to. It isn't enough to check that a cited link exists — a real URL can sit under an unsupported sentence. So a judge reads each individual claim against the actual text of its cited source and rules it supported or not. Unsupported claims go back to the writer. And the system is paranoid in the right direction — from the same file: an unparseable or errored judgment is treated as unsupported, so no claim slips through the gate by accident.
We learned this one the embarrassing way — an early report once invented an "18-year-old" on an alcohol brand; that story gets its full telling in The Case Against AI Slop Reports. The wall it left behind: age, health, legal, and financial claims on regulated products are emitted only if sourced and verified. A regulated claim that fails verification is, per the code, dropped entirely — "never softened-and-kept." Not hedged. Gone.
Multi-part AI output has a failure mode nobody talks about: each section is generated independently, so a bundle of reports can be individually defensible and collectively incoherent — the internal write-up on this names it plainly: each report is defensible alone; the bundle contradicts itself, and a client reading two documents that describe different companies loses trust. The newest wall we've built is an editorial pass that extracts every claim across a finished set, detects where reports assert incompatible things, and adjudicates each conflict against the verified fact base before delivery — resolving toward the best-supported reading, never fabrication. A single-pass generator ships its contradictions to you and lets you find them in front of your board.
None of this makes the words smarter. That's the point. The model was already eloquent. What it wasn't, was accountable — and every mechanism above is a piece of accountability implemented in code, each one built because generative systems fail in a specific, repeatable way that eloquence hides.
That's what you're buying. Not a smarter paragraph. A paragraph that survived a gauntlet.
What it costs, and when you shouldn't buy it
Honest section, because this piece is worthless without it.
It's slower, on purpose and by nature. Real research takes real time — sourced retrieval, claim verification, and reconciliation are actual compute doing actual work, not theater. I'll go further: I think instant delivery would be a lie about the work even if it were possible. If you're paying real money for a strategy document, an answer in eight seconds should make you suspicious — the wait is the sound of the machinery running. But that means if your situation needs an answer in the next ten minutes, this category is the wrong tool. Ask the free model. Genuinely.
AI already knows the famous brands. The inverse-fame finding cuts both ways.
| If AI is guessing about you | If you're a household name | |
|---|---|---|
| What grounding buys you | Grounding helps most exactly where the model knows least — small and mid-size brands, new categories, anything under-represented in training data. | The model's baseline picture of you is already decent, and the delta a grounded system buys you shrinks. |
The buyer this category serves best is the brand AI is currently guessing about — which, statistically, is almost everyone reading this, but isn't literally everyone.
A report you won't act on is the most expensive kind at any price. Strategy is only leverage if it changes a decision. If nobody on your team owns the next move, the most rigorous artifact in the world is a beautifully verified PDF in a drawer. The discipline I hold on my own side of this: a finding without a next action attached is decoration. Hold your vendors to that — and hold yourself to it before you spend.
And the category's honest state: much of what's sold in it today has none of the machinery above. The economics of slop are seductive — the deck is nearly free to generate, so the margin on not verifying is enormous. Which is why the burden of proof sits with anyone selling this stuff, me included.
How to buy it without getting slopped
You don't need to see a vendor's code to test whether the machinery exists. The questions do the work — they're cheap to ask and expensive to fake:
Ask "where did this specific claim come from?" — pointing at one sentence. A grounded pipeline answers with a source, because it kept the source. A prompt in a trench coat answers with confidence.
Ask what happens to a claim that can't be verified. The right answers are "it goes back to the writer" or "it gets cut." The wrong answer is a description of how rarely that happens.
Ask whether the system used your data or your category's. What of yours did it ingest — your site, your numbers, your actual positioning — and can it show you the encoding? If the input was your URL and a vibe, the output is the statistical middle wearing your logo.
Ask whether two sections of their output can contradict each other, and what catches it. Watch for whether the vendor has ever even thought about the question. Most haven't.
And ask what a run costs them. A vendor who knows their own compute cost per report has instrumented their pipeline; instrumented pipelines are the ones with gates in them. A vendor who doesn't know is selling you output they've never audited either.
None of these require technical depth. They require ten minutes and the willingness to make a salesperson uncomfortable — which, for a strategy purchase, is the job.
The founder who doesn't exist
Back to that invented founder — the confident, named, fictional human a frontier model placed at the head of a real company.
The unsettling part isn't that the model got it wrong. It's that the wrong answer and the right answer looked identical. Same fluency, same confidence, same formatting. Every mechanism in this piece — the verification judge, the regulated-claim wall, the coherence pass, the persisted sources — exists to do one thing: make truth and fabrication stop looking identical, so a human can finally act on the output without re-checking every line themselves.
That's the category. Generation was never the product; everyone has generation now. The product is the machinery between the model and your decision.
Pay for the decision architecture. The deck comes free with it.
Where did this specific claim come from?
A grounded pipeline answers with a source, because it kept the source.
See what verified meansIs the free AI answer useful for anything?
For orientation, it's genuinely useful. As the input to a real decision — a repositioning, a channel bet, a pricing change — it's a confident guess about a company that only half-exists.
Why would a paid strategy report take longer than a free answer?
Real research takes real time. Sourced retrieval, claim verification, and reconciliation are actual compute doing actual work, not theater. If you're paying real money for a strategy document, an answer in eight seconds should make you suspicious — the wait is the sound of the machinery running.
Should I trust an author who builds in the category he is writing about?
I build in this category, so discount everything that follows accordingly. The burden of proof sits with anyone selling this stuff, me included — which is why every mechanism here is described specifically enough to be checked.
What's the cheapest way to test whether a vendor's machinery is real?
You don't need to see a vendor's code. The questions do the work — they're cheap to ask and expensive to fake. None of them require technical depth; they require ten minutes and the willingness to make a salesperson uncomfortable.