There's a table in our system called Scraped Content, and one row of it holds more usable brand intelligence than most Slack channels produce in a month.
The row starts life as a TikTok video that went viral in a category we care about. By the time the machine is done with it, the row holds forty-nine columns. The raw layer: views, likes, comments, shares, bookmarks, follower count, duration, publish date, caption, the full transcript. Then the analysis layers stack on top, and the column names tell you exactly what we decided matters: Visual Hook. Undeniable Proof. Sound Strategy. Product Intro Timing. Creative Archetype. Emotional Tone. Why It Works. Risks or Weaknesses. From the comment section: Common Questions. Objections. Purchase Intent Signals. And my favorite column in the entire system: Language To Steal.
Now compare that to how your team probably processes the same event. Someone sees the video. They screenshot it. They drop it in Slack with "have we seen this??" Three people react with the eyes emoji. A month later nobody can find it, nobody remembers the account, and the one thing everyone agreed on, "we should do something like this," never specified what this was.
| The screenshot in Slack | The row in the table | |
|---|---|---|
| What it records | That something happened. | Why, in fields you can query, sort, aggregate, and feed into the next thing you make. |
| What you have afterwards | Three eyes-emoji reactions, and a month later nobody remembers the account. | Forty-nine columns: the metrics, the full transcript, and the analysis layers on top. |
| What can read it next | Nothing. A screenshot has no fields for a machine to read. | The next machine, which mines the library for the patterns that have already landed in that niche. |
That difference is the whole subject of this piece.
How a video becomes a row
The pipeline is a set of automated workflows, and the discipline is visible in the plumbing, not just the output.
A keyword or an account. On a keyword run: the last thirty days only, fifty videos per keyword, captions rather than video files.
Engagement adjusted for account size and for how long the video has been up, resolved into one virality number by the spreadsheet itself.
Only the rows that clear the bar, and then three passes: one on the video, one on the comment section, one on the transcript.
A structured pattern summary lands on the row, the run reconciles its own batch against the clock, and the machine goes quiet.
A scrape run starts from either a keyword or an account. The fork in the workflow is literally labeled "Keyword vs. Account." A keyword run is deliberately narrow: it asks the platform for videos from the last thirty days only, fifty per keyword, and it asks for the captions but not the video files. Old virality is trivia, and a video file is expensive to move around, so the machine doesn't fetch one until a video has earned it. Then the scraper tries something most social tools skip: it checks whether the video has native English captions, downloads the subtitle file if so, and strips it to plain text, timestamps and cue numbers gone. If there are no captions, it doesn't guess. It records "no TikTok captions" and moves on. The transcript column is either real or empty. Never inferred.
Every scraped video lands as a row immediately, metrics and transcript included. Then the spreadsheet itself does the first round of judgment: Base Engagement Score, Follower Tier Multiplier, and Time Factor Score are computed columns that combine into a Virality Score — engagement, adjusted for account size, adjusted for how long the video has been up. A million views from an account with ten million followers is a Tuesday. A million views from an account with four thousand followers, posted last week, is a signal. The formula knows the difference. A screenshot doesn't.
For the rows that do clear it, the machine downloads the actual video, hands it to a vision-capable model, and lands the results: hook type, sound strategy, when the product first appears on screen, what the call to action was, why it worked. Then it pulls the comment section and runs a second pass. That's where Objections and Language To Steal come from, because a comment section on a viral product video is a free focus group that doesn't know it's being observed.
Your buyers are in there right now, telling your competitors' videos exactly what they doubt, what they wish existed, and what words they use for both.
Then a third pass on the transcript: the problem the video names, the desired outcome it promises, the angle, the tone. A final pass distills all of it into a structured pattern summary, stamps the row "done", and the machine goes quiet.
Three analysis passes and a pattern build, stacking columns onto one row. Nothing lands anywhere else. There is no PDF. There is no deck. The deliverable is the row.
Who reads it? Nobody, mostly, and that's the design. The rows are a library, and the thing that reads the library is the next machine. When our brand video pipeline starts work on a brand, its research stage studies the brand's own record and then mines this library of analyzed in-category videos for the hooks, beats, and patterns that have already landed in that niche. A hook advances into a script because the pattern has worked in your category, not because it sounded good in a prompt. So Objections from a competitor's comment section becomes the doubt your script answers in its opening line, and Language To Steal becomes the phrasing that answers it. The Slack screenshot could never do that, because a screenshot has no fields for a machine to read.
And the part that earns the "while you sleep" framing: nobody babysits this. When a batch fires, the run itself gets its own tracking row (started time, status "running") and the pipeline checks its own progress every thirty seconds, counting finished analyses against expected ones. A video whose analysis fails counts as finished, because the point is to reconcile the batch, not to pretend every video will cooperate. And there's a five-minute ceiling, so a stuck video can't hold the batch forever. When the count reconciles, or the clock runs out, the run stamps itself "done" with a finished time. I can start a scrape of a category at breakfast and read structured results whenever I get back. My presence is not an input.
One worker in a fleet
The viral-video analyzer is one worker in a larger fleet, and the same shape repeats across it: go look at something, come back with fields.
There's a version for videos without captions. There's an Instagram sibling that does the same transcript-and-comments treatment for Reels. There's an image analyzer, a product-page analyzer, a landing-page analyzer, and a full-site brand analyzer that crawls an entire website and lands the brand's actual guidelines as structured data. And there's a social research sweep that runs competitive queries through a research service and writes back fields for up to five named competitor types, a differentiation matrix, and up to three vulnerability windows, competitive intelligence as columns on the brand record, not as a strategy doc nobody opens twice.
As of our July 5 inventory:
A node is one step in a workflow: fetch this, score that, write it here. Every active one exists to turn something unstructured — a video, a comment section, a website, a search result — into something a query can reach.
What the honest version costs
I'll tell you what the screenshot-in-Slack crowd never has to admit: this approach has a janitorial bill.
Those inventory numbers cut both ways. If 40 workflows are active, 54 are not. The retired ones are experiments, dead ends, and testing scaffolds, including a workflow literally named "My workflow 6" that somehow accumulated 393 nodes. The fleet looks like a machine from the outside and like a workshop from the inside. Signal engineering is engineering, and engineering leaves shavings.
Better: one of the analysis columns is spelled "Hook Tyle". A typo, made once, in a column name, and now effectively permanent, because that misspelled column has data behind it and workflows writing to it, and renaming it means touching everything that touches it. Social listening tools never show you their typos. Systems you actually build wear theirs forever.
And the deepest cost is the one that looks like a feature: the taxonomy is the judgment. Language To Steal exists as a column because at some point we decided comment sections are copywriting mines. Product Intro Timing exists because we believe when the product appears is a decision worth tracking.
The machine extracts exactly what we thought to ask for, at scale, forever — which means a wrong column list industrializes the miss.
The fleet watches while you sleep. It does not think while you sleep. Choosing the columns is the thinking, and that part never automates.
Vibes versus fields
For years as an operator I watched videos hit two million views with no structured way to understand why. This system is what the other side of that looks like: the "why" has columns now.
That's the actual difference between social listening and signal engineering, and it isn't the tooling. Social listening asks what are people saying? and accepts a mood board as an answer. Signal engineering asks what fields does the answer live in? and then builds the thing that fills those fields every time, whether or not anyone is watching.
You can start this without any automation at all. The next time a video in your category takes off, resist the screenshot. Open a spreadsheet instead and answer the columns by hand: what was the hook, when did the product first appear, what are the top three objections in the comments, which comments read like someone reaching for their wallet, what phrasing did commenters use that you'd never have written yourself. That last column is the one to be strict about. Our definition of it is "high-impact phrases worth borrowing," and the bar is that you wouldn't have written the phrase yourself; if you would have, it's your language, not theirs. Ten videos in, you'll have opinions backed by fields instead of memory. The automation only matters once you know which columns your category's answers live in, because the machine can only industrialize judgment you've already made.
A screenshot tells you something worked. A row tells you why. A thousand rows tell you what to make next, and they'll still be there, queryable, when you're ready to ask.
The brand video pipeline mentioned in this piece is in private beta.
The same move, one layer up
This fleet turns a video into fields. Jinn does the same thing with a brand — it reads your site, your products and what you've published, and assembles it into one working record you can open and read.
See how it worksWhat happens when a video has no captions?
The scraper checks for native English captions and downloads the subtitle file when it finds one, stripped to plain text. When there are none, it records that fact and moves on. The transcript column is either real or empty, and never inferred — which matters, because two of the three analysis passes read from it.
Doesn't analyzing every video get expensive?
It would, which is why the machine does not. Videos are scored before anything expensive happens — engagement adjusted for account size and for how long the video has been up — and only the rows that clear the bar earn the video download and the analysis passes. When nothing in a batch clears it, the run marks itself finished on the spot and no analysis is spent at all.
Who actually reads all these rows?
Mostly nothing human, and that is the design. The rows are a library, and the thing that reads the library is the next machine: our brand video pipeline mines it for the hooks, beats and patterns that have already landed in a niche, so a hook advances into a script because it worked in that category rather than because it sounded good in a prompt.
Do I need automation to start?
No, and starting without it is the better order. Open a spreadsheet and answer the columns by hand for the next ten videos that take off in your category: the hook, when the product first appeared, the top objections in the comments, the phrasing you would never have written yourself. The automation only matters once you know which columns your category's answers live in.