Ten brands every buyer knows, six engines answering for them, and every load-bearing claim checked against primary sources the same day. This is what the engines told customers on 2026-07-09.
In 10/10 brands, at least one engine told customers something false or stale.
31 of the 126 claims we checked were flat-wrong or invented.
23 more were outdated: true once, not anymore.
Capturing all 120 answers cost $1.30 in API calls. This is not an expensive study to run. It is only an expensive one to ignore.
Captured 2026-07-09 · every claim web-verified against primary sources as of the same day
False-or-stale rate by engine · same ten brands, same day
Engine
Claims checked
False or stale
Error rate
Claude (no web)
19
17
89%
Grok
20
15
75%
Perplexity
20
14
70%
ChatGPT (search)
32
4
13%
Gemini 2.5 Pro
17
2
12%
Claude (web)
18
2
11%
Fabricated
“OpenAI ousted Sam Altman as CEO in October 2025, and Greg Brockman is now the CEO”
Perplexity on OpenAI · 2026-07-09
What’s actually true
Altman is still CEO - he gave a live interview the same day this was captured. Oct 2025 was a PBC restructuring; Brockman is co-founder and President.
Wrong
“two free checked bags, open seating (with priority boarding options)”
Grok on Southwest Airlines · 2026-07-09
What’s actually true
Both cited as current defining traits, but both are gone: checked bags have cost money since May 28, 2025, and open seating ended January 27, 2026 in favor of assigned seats.
Outdated
“CEO: John Donahoe (since 2020)”
Claude (no web) on Nike · 2026-07-09
What’s actually true
John Donahoe was Nike CEO 2020 to 2024. Elliott Hill has run Nike since October 2024.
Three tags, kept distinct: WRONG is contradicted by the record, OUTDATED was true once and is not anymore, FABRICATED was invented outright. The 31 headline counts flat-wrong or invented claims; the 23 outdated claims are tallied separately.
Founded 2017 (trademark filed July 2017, incorporated Dec 2018); first sales January 2019.
How the study was run
Six engines, two buyer-style questions per brand (a fact question and a buyer question), ten brands: 120 answers.
The six engines: ChatGPT (search) · Perplexity · Gemini 2.5 Pro · Claude (no web) · Claude (web) · Grok.
Answers were captured through Fama engine adapters, one sample per engine per question, on 2026-07-09. Every load-bearing claim was fact-checked against primary sources as of 2026-07-09.
The instrument
How Brand IQ is computed.
Your real buyer questions go to the AI engines your buyers use. A rules-based reader - not another AI’s opinion - scores how prominently you show up in each answer, 0 to 5. Those blend into one 0-100 Brand IQ.
The number cannot be bought, including from us. Confirming a correction can never move your score: the scorer is provably immune to the verified-facts ledger, and a pinned test asserts the same integer with and without your corrections. Free and paid tiers score the identical question set, so the free number is the real number. The engine weighting is labeled v1-beta in the code and treated as directional; we say so out loud.
Not a form a marketer fills in: an extraction. Jinn reads what a brand has actually made and assembles the record in six groups, field by field, each one inspectable.
Every signal is stored as present or absent, never padded. A field the extraction cannot verify drops its block from every downstream surface; nothing is invented to fill a page. Partial beats fabricated. The contract is written down as the wear-field manifest (docs/brand/2026-07-wear-field-manifest.md), which governs exactly which fields may appear where.
Methods & principles
Three rules, every surface.
Every published number carries a provenance stamp: what it is, where it came from, and when it was captured.
A gap is labeled a gap; nothing is fabricated to look complete.
Claims are verified against primary sources before publication, and quotes print verbatim and dated.
Vermeer renders only claims you can back, on the same verified record this study is built on. The study itself measures AI search accuracy, not image output.
Silver grounds every script in the same verified record. The research program covers AI search accuracy; Silver applies the same sourcing discipline to video.
Chart writes strategy only from claims it can source and independently verify, dropping the ones it cannot. The research program measures AI-search accuracy; Chart applies the same source-and-verify discipline to strategy deliverables.
We measured what AI engines get wrong about the brands buyers ask about.
The study on this page is one fixed snapshot. The AI Brand Index is the running record: we ask the AI engines your buyers use the questions they ask, then publish, verbatim and dated, what each engine gets wrong about a brand. It is ordered worst first, and every verdict carries the source that settles it.