There's a small function in our reporting pipeline that exists because of how brands answer one question about themselves.
The question is simple: who are your competitors?
And a certain kind of answer shows up reliably enough that we had to write code for it. Not a list of rivals. A category. "Sustainable snack brands." "Premium bottled water." A shape where names should be.
The engineering note on that guard spells out what happens if the answer gets through. Fill in your competitors as a category instead of named companies and it "silently corrupts the share-of-conversation denominator" — the pipeline goes searching for press about a phrase instead of a company, comes back with thin, off-target results, and produces what the comment calls "an honest-but-WRONG zero." A number that looks like a measurement and isn't.
So the guard drops category phrases before they can touch the math, and the comment states its design philosophy in one line: it "errs toward the honest gap over a corrupted denominator." If we can't measure you against named rivals, you get a labeled gap. Never a wrong number.
There's even a carve-out in there that I find quietly funny: the guard can't just ban the word "brands," because real CPG companies have it in their legal names — the comment lists "Conagra Brands" and "Utz Brands" as exactly the rivals it must not delete. The line between a brand and its category is thin enough that separating them takes deliberate engineering, even at the level of a text filter.
I keep coming back to that guard because the failure it catches is not a data-entry mistake. It's a habit of mind. Asked who they compete with, operators reach for the category — because the category is how the whole industry talks. Distributors talk in categories. Buyers reset shelves by category. And every AI tool you've touched this year thinks in categories more deeply than any of them, because categories are what it was trained on.
Here's why that habit has gotten expensive. In CPG, the category average is not neutral ground. It's occupied territory. Two different machines are purpose-built to own it — one prints it in plastic, the other prints it in prose — and most of the AI tooling sold to brands right now is quietly coaching you onto their turf.
The middle already has two owners
The first machine is older than AI: private label.
Think about what a retailer's store brand actually is, mechanically. The retailer sits on the category's sales data at the item level — they can see exactly which format, which claims, which price points the category has already proven. When they commission a store-brand version, they don't copy the weird outlier. They copy the proven center: the pack format the category converged on, the front-of-pack claims the middle of the category makes, the shelf language every shopper already recognizes. Then they price it under the brands that did the proving. That's the whole strategy: compute the category's mean, manufacture it, and use structurally better economics to win it. It works because it's average — the average is the part of the category that's been de-risked.
The second machine is the one filling your feeds and, increasingly, your competitors' PDPs. A general AI model asked to write for a brand produces the most probable output for a brand of that shape. It was trained on your whole category — every rival's site, every review, every listicle — so its cheapest, most confident answer is the center of that category. What people call AI slop isn't broken output. It's the category mean, rendered fluently, at a marginal cost of roughly zero.
Notice these are the same computation.
| Private label | Generative AI | |
|---|---|---|
| How it runs | Manufacturing speed, with a real balance sheet behind it. | Typing speed, for free. |
| What it converges on | The category's average product. | The category's average language. |
If your own expression sits anywhere near that average — your packaging claims, your product page, your positioning deck — you are now competing against one opponent with structurally lower prices and another with a structurally zero cost of production, both optimized for exactly the position you're standing on.
I've made the general version of this argument before: a general model is a machine for finding the statistical middle of whatever it was trained on. What I want to do here is make it physical. In CPG the middle isn't an abstraction about model behavior. It's a facing on a shelf, two positions over from yours, at a lower price. And now it's also every fluent paragraph your category's tools generate for whoever asks.
To be clear, there's no villain in this. Private label is a rational business and frequently a good deal for the shopper; the models are doing precisely what they were built to do. The point isn't villainy. It's real estate. The middle is taken, by tenants you cannot out-spend and cannot out-volume.
Private label will match your format. The models will match your category's voice. The one thing neither machine can generate is whatever was never in the category data to begin with.
Which is what makes one particular sales pitch so dangerous right now.
“Works for any brand” is a confession
Nearly every AI tool aimed at consumer brands says some version of the same sentence: works for any brand, out of the box, no setup.
Decode that sentence mechanically. For a tool to work on any brand instantly, it has to run on inputs it can get without you. What can it get without you? The category prior — the model's trained sense of what brands shaped like yours say and do — plus maybe a crawl of your homepage. That's the entire ingredient list.
Every blank you don't fill gets completed with the most probable answer for a brand like yours. Your tone, if you didn't specify one, is the category's tone. Your customer, if you didn't define her, is the category's median customer. Your differentiation, if you didn't encode it, is whatever the model finds "differentiating" for your kind of company — which is, by construction, the same thing it finds for every company of your kind. The tool isn't lying when it says it works for any brand. That's exactly what it does: it produces the output that works for any brand. Output that works for any brand is, definitionally, output that distinguishes none.
Every blank you don't fill gets completed with the most probable answer for a brand like yours.
And there's a loop hiding in the "we personalize from your website" version of the pitch. If your site copy is itself category-average — and we'll test that in a minute — then the tool reads the average in, and confidently hands the average back, now with your name at the top. The machine didn't drag you to the middle. It just confirmed you were already there.
That's the trap in one sentence: a tool that works for any brand is optimizing you toward everyone else — and billing you for the trip.
What it looks like to actually write your brand down
The way out is not a better subscription. It's encoding: getting the specific, checkable substance of your brand out of founders' heads and scattered decks and into a structured record that your tools are forced to consult — so there are no silences for the category prior to fill.
I can tell you concretely what that looks like, because building this record is the unglamorous center of what my team does. Set the product aside; look at the shape of the record itself, because the shape is the argument. The canonical brand record we maintain doesn't ask for a personality paragraph. It asks for declared positions: a positioning wedge. A named brand enemy — not a competitor, the thing you're against. Banned words sitting next to safe words, with a slang policy between them. Audience tribes, each carrying its own motivation, not one blended "target consumer." Vulnerability windows — the actual moments your buyer is reachable. A differentiation matrix: what you can say that a rival can't truthfully say back.
None of those cells can be filled by a category. Every one of them forces a choice the middle never makes.
Two design details in that record matter more than the field list, and both echo the guard from the top of this piece.
In the contract's own words, "presence is tracked separately so a consumer never guesses whether '' means 'blank in DNA' vs 'field absent.'" Translation: when something about your brand is unknown, every downstream system sees a labeled gap, not a blank it's free to improvise over. A gap the machine can see is a gap the machine can't quietly fill with the category.
Generated, verified, or absent. The machine's inferences never silently graduate into facts about you; verified means a human said so. The record is honest about which parts of "you" are actually you.
That's the same philosophy as the competitor guard, applied everywhere: the system would rather show you a hole than hand you an average. And I'd argue that's the correct posture for any brand, not just any pipeline — because in the age of two mean-machines, being averaged without noticing is the failure mode that kills. You can fix a known gap. You can't fix a confident wrong number you never knew was wrong.
A brand described in a handful of adjectives can't do this job, because "bold, authentic, premium" describes the category, not you — the job takes verified signals, and it takes a lot of them. (The machinery for that is a story I've told elsewhere.) But you don't need our machinery to start. You need five minutes and a test.
The swap test
Or your About page, or the last caption your team shipped.
Read the whole thing again.
The share of your copy that survives the swap is a rough, brutal measurement of how much of your brand's expression currently is the statistical middle — which is to say, how much of it private label can print in plastic and a model can print in prose, today, without your permission.
You've just drafted your first verified signals, for free.
Every sentence that still works — every claim your rival could make without blushing — is a sentence the category wrote, not you. "Clean ingredients you can pronounce." Swaps fine. "Made for people on the go." Swaps fine.
Every sentence that still works — every claim your rival could make without blushing — is a sentence the category wrote, not you.
Then look hard at the sentences that broke. The origin story only you have. The formulation choice with a reason behind it. The enemy you named that your rival is friends with. That short list is your actual, defensible asset — and the seed of your encoding.
If you run this test and almost everything survives the swap, that's a painful result and a genuinely useful one — it's your own honest-but-wrong zero, caught before the machines monetize it.
What specificity costs
I'd be selling you something if I skipped this section, so: the bill.
How much work is this, really?
Specificity is work with no shortcut. Extracting it, verifying it, keeping it current as the brand evolves — that's sustained, unglamorous effort, and it never fully finishes. The category prior is free precisely because it's nobody's work.
Won't being this specific cost me customers?
Specificity narrows. A wedge that says something real excludes someone; a named enemy makes one. The middle feels safe because it offends no one — that's also why it defends nothing. If you encode who you are, some shoppers will correctly conclude you're not for them. That's the mechanism working, and it stings anyway.
Can I just write down a more flattering version of the brand?
Specificity has to be true. An encoded fiction is worse than the honest average — your customers will catch it faster than any model would, and unlike the category mean, a broken promise is distinctly, memorably yours.
What if my product genuinely isn't differentiated?
And the hardest one: if your product genuinely is undifferentiated, encoding won't invent a difference. It will document the absence — a wall of labeled gaps where a brand should be. But I'd want that diagnosis on paper before the copy machines deliver it to me at shelf, wouldn't you? A labeled gap is a to-do list. A corrupted denominator is a surprise.
The honest gap
Back to the guard one last time, because its one-line philosophy — err "toward the honest gap over a corrupted denominator" — has become how I evaluate every tool that wants access to a brand I care about.
Ask what the tool does where it doesn't know you. Does it show the gap — or fill it? Because "works for any brand" means it fills it, always, with the only thing it has: everyone else. The category average, fluently rendered, priced at zero, indistinguishable from insight until you notice your copy survives the swap test and your shelf has grown a cheaper twin.
Your brand isn't average. But average is the default output of every general machine now pointed at your category — and defaults win wherever nobody has written down the specifics.
Private label will match your format. The models will match your category's voice. The one thing neither machine can generate is whatever was never in the category data to begin with. Protect it the only way that works at machine speed: write it down, verify it, and make every tool you hire consult it — before your tools write it out.