There's a sentence every consumer-brand operator has said in a meeting, usually with a sigh: "legal hasn't cleared it yet."
If your product goes in a mouth, on skin, or anywhere near a health outcome, you live inside a fence of what you're allowed to say. Words like "cures," "treats," "clinically proven." Implied claims that a regulator reads differently than your copywriter does. Platform ad policies layered on top, each with its own banned list. The review process this fence created was slow, and everyone resented it, and it existed because the alternative was worse.
Then AI content tools arrived, and something strange happened that almost nobody said out loud: the review stage didn't get automated. It got deleted.
Volume broke the old gate, and most tools shipped the break
The old system worked because volume was low. Three videos a quarter, a person reads every script, the bottleneck is annoying but survivable.
AI video makes volume the whole point. Dozens of scripts, variants per script, hooks by the batch. And a human review process designed for three videos a quarter meets that volume in one of two ways: it collapses into rubber-stamping, where a person "reviews" forty scripts the way you accept terms and conditions — or it gets quietly bypassed, because the tool posts straight to the ad account and the review was never built in at all.
Most of the AI video category sold the second failure as speed. Type, generate, post. There is no stage in a prompt box where anyone (machine or human) asks whether the claim in the third second of your clip is one your brand can legally make. The tool doesn't know your category is regulated. It doesn't know what a structure/function claim is. It will write "clinically proven" with the same cheerful confidence it writes anything else, because to a general model, that's just a phrase that performs well in ads.
Autopilot isn't a feature in a regulated category. It's a liability with a subscription fee.
The fix is a sequence, and the order is the design
Returning to a human reading everything raw is off the table; volume already killed it. So is trusting the machine, because accountability killed that before AI existed: a regulator does not accept "the model wrote it" and neither does your co-packer's counsel.
The design that actually works is a gated sequence: machines filter in order, and a person decides at the end, on the filtered set.
We built this into Silver, our video product. Every script that survives the idea stage passes four checks before production:
The claims fence, checked explicitly, as its own stage rather than a hope. Every script is read against FTC and FDA guidelines, and what comes back is specific: the claims that got flagged, the disclaimers the script now needs, the risks a lawyer would circle.
Is this something we would say, in the register we say things in. The check runs against your own brand record, the same one the research stage studied, so the lines it holds the script inside are your tone, your vocabulary, your positioning, and your banned-terms list. If it doesn't sound like you, it gets rewritten.
The rules of the venue where the ad will actually run, which are their own moving target. Ad rules, format requirements, and content restrictions for TikTok, Instagram, YouTube. A script a platform won't run is not a finished script, however clean it reads.
Not a human option. A gate. Nothing reaches production without a person deciding it should — our own product copy states the design intent in a phrase I'll borrow back: the fourth stage is a human, by design.
Two details in that sequence carry most of the value, and they're the two most tools skip.
And the machine shows its work. Scripts arrive ranked with the reasoning attached. A reviewer who can see why something was flagged or favored can exercise judgment; a reviewer staring at an unexplained pass/fail learns to click approve.
A gate is only real when the person standing at it has enough context to disagree.
Why the claims check goes first
The order inside the machine stages isn't arbitrary either; it's the choice I'd argue hardest for.
Regulatory runs first. Before brand, before platform. The reason is where the cost of a mistake lands.
| Caught at the script stage | Caught after a render | |
|---|---|---|
| The cost of the same regulated claim | A rewrite: a few sentences change, the script goes back through the checks, and nothing downstream has happened yet. | The render, and everything that was built on top of the script (the shots, the character references, the timing) gets thrown away with it. |
Put the most expensive failure at the top of the funnel, where a failure is cheapest to fix.
There's a second, quieter reason. Brand voice and platform fit are matters of degree; a script can be a little off-voice and still be worth a person's time. A regulated claim is binary. You can say it or you can't, and no amount of tone work changes which side of the line it's on. Binary checks belong first because they clear the slate of things no later stage can rescue.
"Fail" here doesn't mean the script dies. A script that fails a check is rewritten and re-checked, so what reaches the person is the fixed version, with the fix visible, rather than a graveyard of rejections. This matters more than it sounds. A gate that only says no trains people to route around it. A gate that says "no, and here's the version that passes" is one they'll keep using.
What the person at the gate actually sees
The whole argument rests on the human stage being real, so it matters what the person is looking at. Not raw output. The slate is the set of scripts that cleared all three automated checks, ordered by a ranking pass that scores each one and attaches a rationale line: the kind of short justification that says the hook is the strongest of the batch, or the platform fit is highest, or the angle is the one most likely to read as authentic. The score is a number; the rationale is a sentence a person can argue with.
The person's options are approve, edit, or regenerate. That third option is the tell that the gate is a decision and not a checkbox. If the best script on the slate is still not the one, the reviewer can send the whole batch back rather than settling for the least-bad option in front of them. And when they approve, the approved script is theirs; you keep full rights to what you sign off on.
One more thing about that reviewer, drawn from a mistake I nearly shipped elsewhere at Jinn. We had built a "brand-voice consistency guardian" that audited every output against the brand's DNA and flagged anything that drifted. It looked responsible. It was wrong at the premise: the brief is the user's declared intent, so the moment someone deliberately asked for something off-voice, the guardian would have flagged their own instruction as drift. I killed it as soon as I saw that. The guardrails in Silver check the script against your record, and the person at the gate can override them. The human's decision is the source of truth. Tooling serves it; it doesn't second-guess it.
The deal underneath all of it
I've held one conviction through every AI system we've built at Jinn, and it predates all of them: AI generates. Humans decide. That's the deal.
It's a design philosophy rather than a safety disclaimer bolted on for the enterprise slide, and in regulated consumer categories it happens to also be the only philosophy that survives contact with reality. The machine's job is to make the decision cheap: run the checks, kill the weak options, surface the reasoning. The person's job is the decision itself — the yes that puts your brand's name on the thing. Every tool that tries to automate the yes is misunderstanding which part was expensive.
AI generates. Humans decide. That's the deal.
The deal has a second clause that I hold just as firmly: autonomy is earned, never assumed. A system starts by doing the work and showing it. Then it earns the right to surface recommendations. Only once it has proven itself on real usage, and only with a way to revert what it did, does it earn the right to act on its own. Off, then surface, then auto, and you don't skip to the end. In a regulated category the end may never be the right place to be, and that's fine. The point of the sequence is that a person is in the seat for as long as being in the seat is the thing that protects you.
What this means for your operation
If you're evaluating AI video vendors (or building your own operation with general tools), the audit takes five minutes.
Ask where the claims check happens, and whether it's a stage or a suggestion.
Ask what a reviewer sees: raw output, or a filtered slate with reasoning.
And ask what, mechanically, stands between a generated script and a live ad.
If the honest answer is "nothing," you've found where the liability lives.
Then run the same audit on yourself, because most brands doing this with general tools have the same gap and no vendor to blame. Write down, in order, what happens between "the model wrote a script" and "the ad is live." If the list is one item long and that item is a person skimming forty scripts on a Friday afternoon, you have the rubber-stamp failure. If the list is empty, you have the bypass. Either way the fix is the same shape: put the binary checks first, make every rejection come with a fix, and give the person at the end a short, ranked, explained slate so their attention lands where it counts.
The camera work will keep getting more impressive. The part of the process that protects you doesn't produce a single frame.
The fourth stage is a human, by design.
The design that actually works is a gated sequence: machines filter in order, and a person decides at the end, on the filtered set.
How it worksIf a machine runs the claims check, aren't I still trusting the machine?
You're trusting it to filter, not to decide. The automated checks reduce the slate and attach their reasoning; the decision, and the accountability for it, stays with the person who signs. A regulator asks who approved the claim, and the answer is a name, not a model.
Doesn't a four-stage gate slow everything down?
It slows down the part that should be slow: the yes. Everything before it (research, hooks, scripts, the three automated checks) runs in hours rather than weeks, so the person's time is spent on a ranked, explained slate instead of on reading raw output. The gate is where the hours go, and it's the one place they're worth spending.
Silver is in private beta. The compliance pipeline described here runs today; self-serve access isn't open yet, and nothing shown on our pages is a customer's work.