Nobody actually wants an autonomous marketing AI. They want one that earns it, the way a new hire does.
I've come to that conclusion slowly, watching both sides of the question — building AI that does marketing work, and being the human who has to decide how much of it to trust. This piece is the closest thing we have to a philosophy of the whole company, and it ends with a framework you can use on any AI tool you run, whether or not it's ours.
A staff, not a copilot
I don't think in copilot-versus-replacement terms anymore. Both framings miss what's actually being built. A copilot assumes a human doing the work with help; a replacement assumes the human gone. What we're building at Jinn is a staff:
Drafts overnight and has the queue full by morning.
Shoots on-brand product creative.
Writes the board report.
Gates every claim before it ships.
The staff framing changes what you build, because the moment you call them roles, a question appears that the copilot framing never has to answer: what are the terms of employment? A copilot has a user. A staff has a manager, and management is a real discipline with real machinery — reviews, sign-offs, probation, promotion. The part almost nobody builds is the employment terms.
Autonomy is a ladder, and the human does the promoting
Ours goes like this, and we wrote it down early as doctrine rather than improvising it one product at a time.
Every role starts with its autonomy off.
It does the work, files it in an approval queue, and a human signs off on every piece.
Only after it proves itself does the human promote it, one capability at a time, never across the board.
Surface is easier to picture than to define, so picture the publicist at that rung. Overnight it reads the market and picks the opening worth taking. By morning the draft is written, in the brand's voice, waiting in a queue. Someone reads it, tweaks a line, and approves it or sends it back. Only then does it publish, on the cadence that person set. The publicist did all of the work and none of the deciding. That gap is the whole rung.
The queue is also where the evidence starts. Every draft carries a stamp of where it came from: written by hand, promoted from an idea, pulled from a listening feed, or drafted overnight, fixed at the moment of creation rather than guessed later. Every state change is written to a trail that only grows, and the moment someone approves, the system freezes a snapshot of exactly what was on screen, including the compliance verdict showing when they pressed the button. Months later you can still answer what was approved and what actually went out without taking anyone's word for it.
Two details in that ladder carry most of the safety, and they're the details to check in anything you evaluate.
Per capability, never across the board. A role that has earned unreviewed drafting has earned nothing about publishing. Blanket autonomy is how one good month of drafts turns into an unsupervised system doing things it was never observed doing. Granular promotion means the trust you extend is exactly the trust that was earned, no more.
The photographer shows what a single promoted capability looks like in practice. On ad creative, its checks are advisory: a legibility score, a fidelity flag, a warning that text has landed over a face, all surfaced next to the frame for a person to weigh. Nothing about that has been promoted.
| A cosmetic glitch | A detected competitor logo | |
|---|---|---|
| What the photographer is allowed to do | Exactly one automatic retry. | No retry at all. It stops and waits for a human. |
| Why | A reroll fixes a glitch. | A reroll cannot fix an intellectual-property problem, and retrying silently would hide the one failure a person most needs to see. |
Product shoots go one rung further, and only if you arm it: an accuracy lock scores every frame against your real product photos, and a frame under your floor re-rolls on its own. You switched it on, so nothing about it is silent either. That is what per-capability promotion means at ground level. The retry earned auto. The judgment did not.
The human does the promoting. Not a confidence score or a streak counter, and never the system grading its own work. A machine that promotes itself has an obvious conflict of interest, and a buyer who accepts self-promotion has outsourced the one decision that was theirs.
The same rule holds inside the staff. A writer role hands drafts up the chain instead of marking its own homework; an editor role owns the standard and approves anything that ships; an owner sees everything and signs anything. The compliance officer sits across all of it, and its gate is not a suggestion the model can talk its way past. A draft that fails a check waits for a fix. It never slips into the queue unchecked, and it never gets to argue that it was close enough.
That is how you'd treat a person. You don't hand a new hire the company card in week one. You read their drafts until you notice you've stopped changing anything, and then you read fewer of them. Earned autonomy is the basis of every employment relationship that works. We wrote it into software, and the approval queue turned out to be the probation period.
We wrote it into software, and the approval queue turned out to be the probation period.
Most of our own staff sits at surface today. Auto comes slowly, on purpose. I'd rather concede that than sell you autonomy the staff hasn't earned.
Trust is the actual product in AI labor
This is the part I think decides who wins the category.
"AI employees" is the most crowded frame of the year, and the models underneath are drifting toward interchangeable. Whatever generation quality distinguishes one vendor's staff this quarter gets matched next quarter. What buyers can't get anywhere is a reason to leave one unsupervised.
That's the actual scarce good. Any vendor can sell you output. A staff you can graduate one capability at a time, with a record of what it did and who approved it, is a different purchase from an intern with amnesia and admin access. The record matters as much as the ladder: trust isn't a feeling you develop about a tool, it's a conclusion you draw from evidence, and a system that keeps no evidence of its own track record is asking you to promote on vibes.
A staff you can graduate one capability at a time, with a record of what it did and who approved it, is a different purchase from an intern with amnesia and admin access.
So when I say trust is the actual product in AI labor, I mean it structurally. The generation is the commodity. The machinery that lets a cautious, accountable human extend real autonomy safely — that's what's being bought, whether the buyer has words for it yet or not.
Evidence is also what makes the ladder honest in the other direction. A demotion is only possible if you can see what a role did after you stopped watching. The snapshot at each approval, the origin stamp on each draft, the trail that can't be quietly edited: none of that is there to make the staff look trustworthy. It's there so you can check, and a staff that invites checking is the only kind that deserves to be promoted.
The ladder is a career map, not a threat
If you make your living doing this work and the staff reads like a threat, look at where the ladder ends. Every promotion on it is a human decision. Somebody has to catch the strategist's bad take before the board sees it. Somebody decides when the publicist has earned unreviewed. The job shifts from producing the work to running the staff.
That shift is a real change. But notice it's a promotion in the org-chart sense too: the human moves from doing the drafts to holding the judgment that governs them. The skills that make someone good at that — knowing what good looks like, catching the subtle miss, deciding what deserves trust — are exactly the skills experienced marketers and agency people already have. The approval queue needs a discerning human at the top of it more than the old workflow ever needed another pair of producing hands.
I run the rest of my company the same way — the engineering staff that built Jinn earned its autonomy rung by rung, but that's another post.
Run the ladder on the tools you already have
You can apply this framework this week, on any AI in your stack, ours or anyone's.
The obvious objection is that surface sounds like more work, not less, and for a week it is. A queue full of drafts still has to be read. What changes is the shape of the reading. You stop opening a blank page and start opening a finished draft, and reading a finished draft in your own voice is a different job from producing one. When there are several brands or clients in the picture, the unit of approval moves up a level too: you approve a queue, not a post. The staff doesn't reduce the number of decisions. It reduces the number of decisions that aren't yours.
Then use the promotion signal that works on people: you've stopped changing anything. When your edit rate on a capability's output hits zero and stays there, that's evidence it has earned the next rung. When you can't remember the last time you reviewed something that publishes in your brand's name, supervision has lapsed, the same end state with none of the evidence.
The web is being rebuilt for readers that aren't human, and brand work is already outgrowing what human hands can make alone. The staff is coming either way. When someone sells you one, ask who does the promoting. If the answer is the machine, walk.
Every role works from one brand record
Before any of the staff drafts a line, Jinn reads what the brand is made of and assembles it into one working record that every product works from. This is how that record gets built.
How it worksWhat does the surface rung actually mean?
It is the rung where a role does all of the work and none of the deciding. Output lands in an approval queue rather than in public, and a human signs off on every piece before it goes anywhere. The production is automated; the judgment is not, and that gap is the entire point of the rung.
Why promote one capability at a time instead of trusting the whole role?
Because trust earned on one job says nothing about another. A role that has earned unreviewed drafting has earned nothing about publishing. Granular promotion keeps the autonomy you extend equal to the autonomy you actually observed, and it means a capability can be pulled back a rung without switching the whole role off.
Can a system promote itself once its output is consistently good?
It should not, and a vendor offering it has moved the one decision that was yours onto the machine. Confidence scores and streak counters have the thing doing the work grade the work. The promotion is a human act, taken against a record of what the role did and who approved it.
How do I tell when a capability has earned the next rung?
Watch your own edit rate. When you have stopped changing anything on a capability’s output and it stays that way, that is evidence rather than a feeling. The inverse is the warning sign: if you cannot remember the last time you reviewed something publishing in your brand’s name, that capability is already running at auto and nobody promoted it.