There is a field in our character system called Ref_Hands_Detail.
Its job is to store close-up photographs of the hands of a person who does not exist. The slot's own description reads "Detailed hand reference for gestures," and the generation instruction attached to it is byte-for-byte: "close-up of hands, detailed fingers, natural pose, studio lighting."
It's one slot of twenty on what we call a character's reference card. Headshot facing front. Headshot at three-quarter. Full profile. An extreme close-up of the eyes — "detailed iris texture, macro photography." A genuine smile with teeth visible. Full body from the front, the side, the back. A neutral resting pose with hands visible. An expressive gesture. An engaged expression and a concerned one, because a character who can only make one face isn't a character. The character standing in their typical environment.
Twenty named angles of nobody. Kept on file so that every future render of somebody comes back the same.
That reference card is the visible tip of something I think most brand teams haven't priced in yet: a brand character is now a data product. Not a drawing, not a vibe, not the thing your agency's one great illustrator "just gets." A structured record with about a hundred fields, a provenance trail back to your brand's encoded identity, an approval lifecycle, and a mechanical enforcement layer that holds the face together across renders. This piece is a tour of how that works in our system — and why the old way of keeping a character consistent was never going to survive contact with generative volume.
How characters used to stay consistent
Before the machines, character consistency was a human memory problem, and the industry solved it with humans. A mascot lived in a style guide — front view, side view, approved poses, forbidden poses — and, more honestly, it lived in the hands of the small set of people who had drawn it long enough to feel when it was off. The guide was reference; the judgment was tribal. If the illustrator retired, you found out exactly how much of the character had been living in their head.
That system worked because output was scarce. A character appeared in a few dozen executions a year, each one passing through people who knew the character personally.
Generative pipelines break the economics on both sides at once. Renders are no longer scarce — a single campaign can want fifty of them — and the renderer no longer has a memory. An image model does not remember what your character looked like yesterday. Ask it twice and you get two strangers wearing similar clothes; the model happily regresses your character toward the statistical middle of every face it was trained on. The one thing the old system relied on — a mind that accumulates familiarity — is precisely what the new tools don't have.
So you have two options. Hope, or structure.
An image model does not remember what your character looked like yesterday.
What a character is made of now
In our system a character is a capsule: one structured record, roughly a hundred fields deep, that contains everything a machine needs to render the same person again. The type file's own header states the design goal plainly: each capsule "contains enough detail to generate consistent characters across any video platform… or train a LoRA model" — enough, in other words, to take the character to whatever renderer exists this quarter, or to bake a dedicated model of them.
The fields run from casting sheet to psychology file. Physical identity down to the level a machine can't misremember — not just "auburn hair" but the hex value of it, because an adjective is an interpretation and a hex code is a fact. Distinguishing features, scars and birthmarks, the best camera angle and the lighting the character reads well under. Wardrobe: default outfit style, clothing colors, what they'd never wear. Then the parts that make a character more than a face: backstory, core traits, values, motivations, fears, quirks. Voice type, speech tempo, vocabulary level, catchphrases. Posture, walk style, signature gestures, energy level.
That psychological half isn't decoration. When a character shows up in a script, a video, a piece of social content, those fields are what keeps the behavior consistent, the same way the hex codes keep the pixels consistent. A mascot with a persistent fear and a signature gesture is a character. A mascot with neither is a costume.
And here's the part that gives this piece its title: the capsule knows where it came from. Characters in our pipeline aren't invented from a moodboard. They're cast from the brand's DNA — the structured, verified encoding of who the brand is that everything else in our stack runs on. Each capsule links back to the source brand's record, carries the audience tribe it was built to represent, and stores three provenance fields I'd point to before any of the glamour fields: the tonal attributes it was sourced from, the tribes it was sourced from, and a written generation rationale. The capsule doesn't just describe the character. It documents why this character exists — which segment of the brand's actual audience they were cast to embody.
That was the design intent from the first spec, which framed the whole system as a casting operation: characters are "reusable assets aligned with customer archetypes." A brand doesn't get one mascot by committee; it hires a small cast, each member standing for a real segment of its audience. The same spec was already fighting the failure mode everyone now recognizes as AI slop, in writing, in its production standards: "CRITICAL: avoid waxy AI look, include subtle pores and natural imperfections."
One more thing about the capsule, and it's the same discipline that runs through everything we build: a character has a lifecycle — draft, generating, review, approved, archived — and nothing gets cast into real work until a human has moved it to approved. The machine proposes a person. A person decides whether they join the roster.
The reference card is the enforcement layer
Fields describe the character. The reference card enforces them.
Each of the twenty image slots carries its own generation instruction, and together the approved images become the character's stable visual anchor. When the character is cast into a shoot in our ad studio, the casting contract is explicit — a shoot's cast is one of exactly three things: no character at all; an ad-hoc character, described in plain English, whose "reference [is] generated first + anchored" before anything else renders; or an existing capsule from the brand's roster.
Then the anchor mechanism takes over, and this is the sentence I'd frame from the render planner's own documentation: when a pre-approved reference exists, every shot featuring the character references it directly, "so the character stays consistent independent of render order and a Gate-2 redo of any one shot can't drift it."
Unpack that, because it's the whole thesis in one line of code comment. Every people-shot in the set is conditioned on the same approved reference image — the render call literally carries the reference alongside the product photography.
| Without a pre-approved card | Anchored to the card | |
|---|---|---|
| Where each new shot gets the character | Shot four inherits it from shot three. | Shot four doesn't inherit the character from shot three; both inherit from the card. |
| What a redo is anchored to | Every regeneration is a fresh casting call, and your character drifts one approved-looking render at a time until nobody can say when they stopped being themselves. | When someone rejects one shot and regenerates it next Tuesday, the new render is anchored to the same bytes as the original set. |
Which means order doesn't matter, parallel renders don't matter, and — the case that actually bites — a redo doesn't matter.
Two details in the same planner show how much of consistency is knowing where it doesn't apply.
First: product shots deliberately never receive the character anchor. Early spike testing found that feeding an establishing lifestyle render into a product close-up "fabricates the product" — the wide shot lacks the label-level detail a close-up needs, so the model invents it. So there are two consistency lanes that never cross: people hold via the character anchor, products hold via real product reference photos. Consistency isn't one blanket mechanism; it's the discipline of matching each element to the reference that can actually vouch for it.
Second: when there's no pre-approved reference and the system falls back to deriving an anchor from within the shoot — first character shot establishes, the rest follow it — that mode only engages when automated render QA is on for the brand, because an "UNSCORED weak hero" can "drag down a whole collection." An anchor is a multiplier, and a multiplier applied to a bad render multiplies the bad. (Every render here is judged by machines before a human sees it — that pipeline is its own story, told elsewhere on this site.)
Consistency isn't one blanket mechanism; it's the discipline of matching each element to the reference that can actually vouch for it.
And the card compounds. Several reference slots are auto-populated from the character's own approved shoot renders — the body shot, the lifestyle shot, the environment shot flow back into the card as they're accepted. The more work the character does, the thicker their file gets, and the more anchored the next job is. That's the property no artist's memory ever had: it appreciates on the record, not in someone's head.
What happens when the brand changes
A fair question for any "persistent character" pitch: your brand DNA isn't frozen — positioning shifts, audiences get re-segmented. What happens to the cast?
Honest answer: nothing automatic, and that's deliberate. An approved character never silently mutates because an upstream field moved. What the capsule gives you instead is the audit trail — those provenance fields say exactly which tonal attributes and which audience tribe each character was cast from, so when the DNA moves, you know precisely which characters were built on assumptions that no longer hold, and re-casting is a decision a human makes with the rationale in front of them. A face your customers recognize changing without anyone deciding it should is not consistency; it's a different failure with better paperwork.
A face your customers recognize changing without anyone deciding it should is not consistency; it's a different failure with better paperwork.
What this costs, honestly
A capsule is real work. A hundred fields don't fill themselves usefully; reference cards take generation and review cycles; the approval gate means a human spends real attention before a character earns a roster spot. The system degrades gracefully for thin capsules — a character with a name and a headshot can still be cast — but the consistency you get is proportional to the structure you've banked.
The machinery also declines work it can't do honestly. An ad-hoc character can't even be generated for a shoot unless there's a shared scene contract to compose them from and at least one people-shot for them to appear in — in the code's own words, "otherwise ad-hoc is pointless and we degrade to product-only." And "no character" is the default cast, not a failure state: most product shots need product fidelity, not a face. A system that put your mascot in everything would be optimizing for the demo, not the brand.
The question to ask any vendor — including us
If someone pitches you AI-generated brand characters, ask one question: where does the character live?
The model has no memory of you; you're describing a coincidence with good marketing.
Ask what happens when the platform those prompts were tuned for gets replaced — this year, that's not hypothetical.
You've reinvented the retiring illustrator, at generative volume.
The answer you want is boring: a record you can inspect, with provenance you can trace to your own brand data, reference images a human approved, and an enforcement mechanism that threads those references into every render — one you could take with you, to any renderer, including the ones that don't exist yet.
Your mascot used to survive on scarcity and one artist's memory. At a thousand renders a year, memory doesn't scale — structure does. The brands whose characters stay themselves through the next five model generations won't be the ones with the best prompts. They'll be the ones who treated the character as what it now is: a data product, cast from the brand's own DNA, with a file thick enough that no machine ever has to guess.
The whole pipeline described here currently runs in operator mode. This is how we make things, not yet a self-serve button. I'd rather tell you that than let a screenshot imply otherwise.
See how the studio holds a character together
The reference card, the anchor mechanism and the approval gate, on the product pages.
How the studio worksWhat actually keeps an AI-generated brand character consistent?
Not the model. An image model does not remember what your character looked like yesterday. Consistency comes from a pre-approved reference card that every render is conditioned on, so shot four doesn't inherit the character from shot three — both inherit from the card.
What happens to a character when the brand's positioning changes?
Nothing automatic, and that's deliberate. An approved character never silently mutates because an upstream field moved. The capsule's provenance fields say exactly which tonal attributes and which audience tribe each character was cast from, so re-casting is a decision a human makes with the rationale in front of them.
Does every render need a character in it?
"No character" is the default cast, not a failure state: most product shots need product fidelity, not a face. Product shots deliberately never receive the character anchor, because feeding an establishing lifestyle render into a product close-up fabricates the product.