There's a rule of evidence built into our report system that reads like courtroom procedure:
"a contradicts verdict requires non-empty grounding"
In plain terms: if the machine wants to rule that one report contradicts another, it has to cite the evidence for which side wins — or the ruling is thrown out as incomplete. Not softened, not waved through with a warning. Thrown out, automatically, before a human ever sees it.
We ended up needing a rule of evidence because we ended up needing a court. If you've ever bought strategy work in a bundle (from an agency, a consultancy, or increasingly from an AI), the problem that court exists to solve is one you have almost certainly experienced from the paying side.
The failure lives between the documents
A serious strategy engagement isn't one document. It's a bundle: positioning, competitive landscape, channel strategy, each piece produced with its own research and its own focus. Producing them independently is what makes each one sharp. It's also what makes the set dangerous.
Each of our reports starts from the same verified brand facts. But facts underdetermine interpretation, and AI models are interpretation machines. Give two independent AI drafts the same inputs and they will each commit to readings the other never sees: what the flagship product actually is, or whether a claim belongs to the brand or to the company that owns it. Every individual report can clear every individual quality check. Then a client sits down with two of them and notices they don't quite agree on what company they're describing.
Nobody wrote a false sentence, exactly. The failure lives between the documents, and the client is the first one to see it.
The machine's internal disagreement becomes the customer's confusion, at the exact moment they're deciding whether to trust you.
You've seen the human version. Four workstreams, four teams, one stapled deck, and a page-nine claim that quietly disagrees with page forty-one. Human agencies solved this a century ago with a role: the senior partner who reads the whole deck before it leaves the building. Our report system had no such person. So we built one: a pass that runs over the finished bundle after the last report is generated and before the email telling the client their work is ready goes out. If a report is later regenerated, the court sits again once the new version lands. It works less like an editor and more like a court, and each of its stages encodes a decision about honesty that most AI products skip.
First, depose every report
The court starts with discovery. Each finished section of each report gets broken down into individual claims: checkable assertions about what a company is, does, sells, or stands for. Tone and aspiration are ignored; you can't adjudicate a vibe.
Every claim gets tagged with the entity it's about and that entity's place in the brand's family tree (the brand itself, its parent company, a product line, a sibling brand). Two rules here do more work than everything else. Entity names are copied exactly, character for character, from the brand's verified record, because disputes are later matched by name, and a paraphrase would split one real dispute into two files that never meet. And when the machine isn't sure where an entity belongs, the instruction is blunt: it must answer "unknown" and never guess "brand." We call that the no-fabrication posture.
An uncertain tag stays uncertain. It never rounds up to a confident one. That posture — when unsure, say unsure — repeats at every layer of this system, and it's the difference between a court and a bluff.
Then build the docket without creativity
Next, potential disputes get paired up: two claims from two different reports about the same entity at the same level of the family tree become one contested point. This stage deliberately involves no AI at all. Its own design rule calls it "PURE and DETERMINISTIC": it uses no AI and reads no outside data, and the same inputs produce the same pairings every single time. It's plain matching. With no model in it, it can't hallucinate a dispute, and there is nothing for a lucky retry to find that the first pass missed. The stage that decides what gets heard cannot be creative.
The family-tree tag earns its keep here: a true statement about the parent company and a true statement about the brand can look contradictory while both being right, so the pairing logic refuses to put them in the same dispute. And the docket deliberately over-includes: two claims that merely emphasize differently get paired too. Pairing doesn't decide anything. Deciding is the next stage's job.
The judge gets three verdicts and a rule of evidence
Every contested point goes before a judge, a single AI call with a deliberately narrow job, ruling against a fixed body of evidence: the brand's verified facts on one side, and the sources the reports themselves cited on the other.
| What the evidence says | What happens next | |
|---|---|---|
| Consistent | The two claims are compatible, the same fact in different words. | Case dismissed. |
| Contradicts | The claims are incompatible and the evidence settles which one is right. | This is the only verdict that triggers a repair, and it's the verdict carrying the rule of evidence from the top of this essay: no grounding cited, no ruling. |
| Diverges | The claims are incompatible but the evidence is silent. Neither side can be proven right. | The system records the dispute, leans toward the better-supported side, and asserts nothing beyond that. |
Most points end here, by design: the mechanical stage pairs generously so the judging stage can dismiss precisely.
A court that ruled confidently on questions the evidence can't answer wouldn't be resolving contradictions. It would be manufacturing new ones with better posture.
A worked case shows the line between the last two. Suppose the positioning report describes a flagship product as the brand's own, and the competitive report attributes the same product to a sister brand under the same parent. Same name, same branch of the family tree, so the docket pairs them. If the brand's verified record carries a family fact confirmed by a person rather than inferred by a model, the judge treats that line as authoritative: the verdict is a contradiction, resolved to the side the family fact supports, with that fact cited as the grounding. If no confirmed family line bears on the question, the judge is told in so many words never to invent one, and the point is recorded as a divergence. One procedural rule sits underneath both: the ruling has to name the winning report, and it has to be one of the two actually before the court. A ruling that names neither is thrown out as malformed, because a resolution nobody can act on isn't a resolution.
And when the judge can't produce a valid ruling at all, after multiple attempts, the point is set aside as unresolved and written to the record, never resolved by guess. Every escape hatch in this system exits toward silence plus a note in the record, never toward a confident invention.
The repair touches only the loser, and never the numbers
Only a contradiction has a loser, and only the loser gets touched. The offending section (not the report, the section) is rewritten to remove the contradiction, under constraints that read like a probation order: same scope, same voice, same structure. And the blocks that carry measured data (charts, rankings, the numbers) are put back from the original untouched, because, in the words of the rule that governs it,
"a rewrite can only ever change prose."
The repair step is not allowed anywhere near the evidence.
The rewrite also has to survive two checks the original draft faced. Its shape is compared block by block with the section it replaces; the same kinds of blocks must appear in the same order, or the rewrite is rejected outright. Then it goes back through the same house-voice gate every first draft passes, so a repair that reintroduces a banned term or a phrase that reads as machine-written is retried rather than shipped. If no rewrite clears both checks within the allowed attempts, the section is left exactly as it was and the record says so. A failed repair is a documented contradiction. A forced one would be a new error in a fresh coat of paint.
The winning report is never touched either, and a diverges verdict is recorded and deliberately left alone. Forcing agreement would mean asserting something the facts don't hold. Then the paper trail: every verdict, with its evidence, is kept as a brief our team can read, and every upheld contradiction is stamped on the losing report's record with a reference to the winning one. The client sees a coherent bundle. We see every fight it took to get there, and who won each one, and why.
The verdicts are written down the moment the judge returns them, before any repair begins, so a pass interrupted partway leaves rulings with no repairs attached, which reads very differently from a clean run whose repairs all happened to fail. And the log can only grow. There is no path to edit or delete an entry after the fact, so the trail of what the court did is the trail a reader gets, not a version tidied up later.
What the court can't do
Three honest limits, because a system described without its failure modes is marketing.
The whole pass is built so that a breakdown inside it never holds up a paid deliverable. The rule is blunt: "a coherence failure can never block a paid deliverable or its completion email." That's a real trade; we decided a trust layer must never become an outage layer, and paid for it with the possibility that the court sits a session out, loudly recorded.
If every report confidently repeats the same wrong claim, there's no disagreement to docket, and the pass waves the bundle through. That is a known limit of the design, not a solved problem. Checking claims against sources is a different courtroom with a different job; this one only guarantees the bundle agrees with itself.
Every contested point is a paid AI call, so the docket has a ceiling and the judging is done in batches, a lesson learned the expensive way, which is its own essay.
That other courtroom does exist in our engine, and it sits earlier: every source a research-backed report draws on is stored before a word is written, each factual claim has to cite a stored source to be recorded at all, and a second model reads each claim against the source text, with an unreadable verdict counted as unsupported rather than waved through. Claims can each be true to their sources and still fail to agree with each other.
Someone is going to be the court
Here's the transferable part, whether you're buying strategy work or building AI products.
Multi-part output (report bundles, deck plus appendix, anything where piece two can disagree with piece five) does not get coherence for free, from AIs or from agencies. Each piece can be locally excellent while the set contradicts itself, because nothing in piece-by-piece production is responsible for the set. So when you evaluate a vendor, ask the senior-partner question directly: who reads the whole thing before it ships, what happens when two documents disagree, and can you show me a case where one was corrected? A real answer describes machinery or a named human. A vague answer means the adjudication happens anyway: undocumented and unpaid, by you, after delivery.
There is no option where nobody adjudicates. The wrong court is the reader.
See how a report gets built
Every source is stored before a word is written, and no claim is recorded unless it cites one.
Walk the buildWhat happens when two reports disagree and the evidence can't settle it?
The pass records the dispute as a divergence, leans toward the better-supported side, and asserts nothing further. Neither report is rewritten, because forcing agreement would mean claiming something the facts do not hold.
Can this catch a mistake that every report in the bundle makes?
No. If all the reports repeat the same wrong claim there is no disagreement to hear, and the bundle passes. Checking claims against their sources is a separate job that sits earlier: every source is stored before a word is written, each factual claim has to cite one to be recorded at all, and a second model reads the claim against the source text.
Can a repair change a chart or a number?
No. Only prose is rewritten. The blocks that carry measured data are put back from the original untouched, and the rewrite has to match the section it replaces block for block, or it is rejected.
What happens if the coherence pass itself fails?
The deliverable still ships and the completion email still goes out. We decided a trust layer must never become an outage layer, so a failure inside the pass is recorded loudly rather than allowed to hold up paid work.
