Illustration — Baselines You Can Trust (Including the Restart)

Fama · Ideas · Article

Baselines You Can Trust (Including the Restart)

The question nobody asks a trend line: what happens when the measurement basis changes? Our own answer, including the honest sentence we shipped wrong first.


September 3, 2026 · 7 min read
Share

If you run a brand, you own a growing shelf of dashboards: retail velocity, search rankings, ad performance, and now tools that measure how AI engines talk about you. Every one of them draws trend lines. This piece is about the question almost nobody asks of a trend line, and how to ask it, using our own measurement product as the specimen, including the part we got wrong.

Start with a message in our AI-visibility product that I like more than any chart we've ever rendered. It fires when a brand's measurement has to start over, and the customer-facing copy reads, verbatim:

That's a product announcing its own discontinuity. Out loud. To the person paying for it.

Most analytics tools do the opposite. The ground shifts under a metric (a tracking change, a new sampling method, a redefined denominator) and the line just keeps drawing. Same chart, same axis, new arithmetic underneath. The trend you're reading spans two different measurements pretending to be one. Nobody lied, exactly. The chart just quietly became a different chart while wearing the old one's clothes.

You've probably been burned by this without knowing where the bruise came from: the "growth" that was actually a tracking migration, the dip that was actually a definition change. The defense is a single idea, and it's teachable in one sentence: the baseline's history is part of the metric. A baseline is the starting point a trend measures from. If the starting point changes and the chart doesn't say so, everything after the change is fiction with good production values.

Any metric that can't show you its baseline history is marketing, not measurement.

We build measurement for an environment where the ground shifts constantly — how often AI engines mention a brand when buyers ask them questions. Engines re-read the web on their own schedule. Question sets evolve. Competitor lists change. If any of that leaked silently into a trend line, the line would be worthless. So the rule above is machinery for us. This is what it looks like built, and what each piece teaches about any measurement tool you'll ever evaluate.

Five states, and only one shows a number

When a brand ships fixes and wants to know whether its standing in AI answers moved, our measurement can be in exactly one of five states. In the system's own labels:

Awaiting corrections

Nothing has been recorded as shipped yet, so there's nothing to attribute movement to.

Baseline pending

The before-picture doesn't exist yet. A true cold start.

Measuring

A baseline exists, but the follow-up window hasn't cleared. Not final.

Baseline restarted

The measurement basis changed. Old and new numbers don't compare.

Measured

Every wall cleared. This is your movement.

That's the whole thesis. A brand that shipped nothing and a brand whose fixes didn't work are in completely different situations, but a dashboard showing "0% change" for both has merged them into one lie. The rule we wrote is blunter than I would have been in a planning meeting: if nothing was recorded as shipped, the report says exactly that, and "NEVER 'correction failed'."

An honest tool refuses to measure early

The Measuring state has teeth. A result doesn't finalize until at least two follow-up captures have completed and ten days have passed since the fixes shipped. That's not bureaucracy. AI engines re-read the web over weeks, and the note we wrote beside that waiting rule names the exact failure it prevents: "The gate prevents the false 'no lift' failure mode — AI engines recrawl over weeks, so measuring on the very next audit is mostly noise."

Before a result can finalize
10days
Plus two completed follow-up captures since the fixes shipped.

Measure the day after you ship and you'll conclude your work did nothing, because the machines you're measuring haven't looked yet. A tool that shows you that number anyway is being wrong quickly.

There's a smaller discipline here to steal for any metrics program you run: the waiting rule lives in exactly one place. Two different parts of the product need the same 2-capture, 10-day rule, and the second one borrows it from the first instead of restating it. The note we left there: "A second hardcoded 2/10 gate would drift, so we import the one true gate." If the rule ever loosens, it loosens everywhere, visibly, at once. There is no version of the product where two screens quietly disagree about what "measured" means. Ask yourself whether your own reporting stack can say the same.

Measure the day after you ship and you'll conclude your work did nothing, because the machines you're measuring haven't looked yet.


We shipped the wrong honest sentence first

This is the confession that earns this essay.

The restart state existed from day one, but for one edge case, we shipped the wrong shape. When a version change wiped out a brand's comparable history entirely, the system found itself with no baseline at all and reported Baseline pending: "your starting picture isn't in place yet."

For that brand, this was false. The brand had a starting picture. It had history. The measurement basis changed and the history was honorably retired, but the receipt described a restart as a cold start, wrongly implying the brand had never been measured. A review caught the gap, and a follow-up fix shipped: a missing baseline plus a version change anywhere in the history now shows Baseline restarted, not Baseline pending. The two states carry different wording because they answer different customer questions: "when do I get numbers?" versus "what happened to my numbers?"

That distinction cost us an engineering change whose entire visible effect is that one honest sentence shows instead of another. I'd approve it again tomorrow.

States are cheap. Trust isn't.

The chart should wear the seam

One more piece of the machinery, because it's the one you can look for from the outside.

When we retired an old measurement method (brands whose question lists had been hand-curated by our team moved onto questions derived from their own brand data; same score, new basis), we added one field to every brand's record whose only job is honesty: a timestamp recording the moment a brand's measurement basis changed. The note we wrote beside it says what it's for: the trends page "renders a vertical 'measurement basis changed' marker at this point." The discontinuity gets drawn on the chart, at the seam, where you can see it. A scar instead of a cover-up.

And then the punchline, stamped right beside that change: "VERIFIED 2026-06-11: the cohort is EMPTY in prod." When we ran the change, zero brands were actually on the old method. We built the seam-marker machinery, shipped it, and it has fired for nobody. We kept it anyway. The next basis change is a matter of when, and the time to build the marker is before you need it, not while a customer is staring at an unexplained jump in their chart.

The words got the same treatment as the lines. When a result finally renders movement, every sentence states the change "alongside" the fixes the brand shipped — never because of them. Before-and-after observation can't prove causation, so the wording isn't allowed to claim it, and that's enforced: a banned list of sixteen causal phrases ("because of," "caused," "drove," "led to," "thanks to," "resulted in," "attributable to," and their cousins) with an automated check that reads every sentence the product can show, in every state, and refuses to let a build through if one appears. As the note above the wording puts it, "the copy here is written to pass it."

The one question to ask any vendor with a trend line

This discipline isn't free. A customer who ships fixes on Monday gets "not final yet" for a week and a half, minimum. A brand whose question set evolved gets a restart notice where a competitor's dashboard would show a thrilling 40-point jump, a jump that is actually two incomparable measurements stapled together, but that feels like progress in a way a restart notice never will. Every one of these choices trades a green arrow today for a defensible number next quarter.

The reason we pay it: a measurement company that blends versions, backfills baselines, or upgrades correlation to causation rots from the inside, and a customer who catches one blended chart doesn't discount that chart. They discount every chart you've ever shown them.

A chart that wears the seamA chart that hides it
When the ground shiftsThe discontinuity gets drawn on the chart, at the seam, where you can see it.The line just keeps drawing. Same chart, same axis, new arithmetic underneath.
What the trend spansThe next receipt measures movement within the new version — the two are never blended.Two different measurements pretending to be one.
What you're holdingA scar instead of a cover-up.Fiction with good production values.

The takeaway applies to every tool on your dashboard shelf, not just ours. Any metric that can't show you its baseline history is marketing, not measurement. Ask your vendor one question: what does the chart do when the measurement basis changes? If the answer is a seam marker or a restart notice, you can trust the line. If the answer is "nothing," you're not looking at a measurement. You're looking at a sales asset that renders on a schedule.

The baseline's history is part of the metric.

Every one of these choices trades a green arrow today for a defensible number next quarter.

See how the measurement works

What is a baseline, and why does its history matter?

A baseline is the starting point a trend measures from. If the starting point changes and the chart doesn't say so, everything after the change is fiction with good production values.

Why does a result take at least ten days to finalize?

A result doesn't finalize until at least two follow-up captures have completed and ten days have passed since the fixes shipped. A tool that shows you that number anyway is being wrong quickly.

What is the difference between a pending baseline and a restarted one?

The two states carry different wording because they answer different customer questions: "when do I get numbers?" versus "what happened to my numbers?"

Why can't the five states be collapsed into one number?

A brand that shipped nothing and a brand whose fixes didn't work are in completely different situations, but a dashboard showing "0% change" for both has merged them into one lie.

See your brand

Explore Fama

Free, across the six major AI engines — what they say about you, where they’re wrong, and where competitors show up instead. A Brand IQ score in a few minutes. No account.

Explore Fama