Summary:
- The bottleneck isn’t search, it’s the absence of signal. Most archives aren’t hard to search. They’re empty of anything to search for, until enrichment builds it.
- Judgment-shaped output beats classification-shaped output. “This moment is worth clipping” is a different and harder thing to produce than “this moment contains a crowd.”
- An answer that disappears helps one person once. A field that sticks helps everyone after. The value isn’t the conversation, it’s what gets left behind on the asset.
- The best demo is the customer’s own footage. Nothing sells enrichment like watching your own unsearchable archive become legible in front of you.
Before anything else happens, the starting point for most of our customers is a file. For the sake of this conversation, let’s say that file is three hours long, has a title, calculated duration, and maybe a project code if that team is organized, that’s it. They don’t have a knowledge base or searchable archive. Just a folder of long-form footage sitting exactly where it landed at ingest.
We’ve been running VERI, a conversational AI tool inside our products, against that kind of file inside the Digital Media Hub. We looked at hours of long-form sports content, all of which has not been enriched with data. The question we had wasn’t “can it tag this?” It was “if we ask it to actually understand this file, what comes back?”
We found that the gap wasn’t search capability; it’s that the data to get there doesn’t exist.
I’ve written before about how conversational discovery changes the relationship between a person and an archive. Ask a question in plain language, get an answer grounded in the file. That’s true, but it assumes the signals are already there to be queried (transcripts, visual detections, structured metadata sitting quietly underneath the product, waiting to be surfaced).
For most customers, that assumption doesn’t hold. The signals don’t exist yet. In the case of our three-hour file example, no one has watched it all. And nobody has tagged the moments that matter. Because of that, the archive isn’t simply hiding value behind a bad search box. It’s genuinely opaque, all the way down, until something goes through it and builds the layer that a search capability depends on—that’s the layer VERI is built to add.
What actually comes back
Point VERI at a long-form sports file and ask two things: summarize it and find the moments worth pulling for social. An everyday task that a digital producer has to do ten times a day when they’re staring down hours of footage and a deadline.
The summarization was expected. It was accurate, had good compression, and delivered no surprises. The moment-finding was the interesting part. It wasn’t returning “crowd noise detected” or “high motion segment” the way a generic model tags action. It was closer to surfacing ”why” a moment would work as a clip. That’s the kind of editorial signal that helps a producer move faster, not just a vision model output.
I’ve written about that same distinction between generic and domain-specific AI: one produces metadata, the other produces something closer to intelligence. Results will vary depending on the content, but seeing it happen on a real long-form file, unprompted by us pointing to a specific timestamp, was the part that actually impressed me.
It doesn’t disappear after you ask
Here’s the part that matters more than the answer itself. What VERI finds doesn’t evaporate the moment you read it. It writes back to the asset as a standalone intelligent field attached to the clip, not just to the conversation you had about the clip.
That distinction sounds small but it isn’t. A chat answer is disposable. Ask the question, get the response, move on, and the next person has to ask it again from scratch. A field on the asset is durable. Once VERI flags a moment as worth clipping, that signal becomes part of the file’s record, making it searchable, filterable, reusable by anyone else who touches that asset later without ever having to re-ask the question or re-watch the footage.
That’s what turns a single clever answer into actual infrastructure. The producer who asked for social moments today built something the ad-sales team benefits from next quarter, without either of them coordinating on it. The archive gets smarter once and retains that intelligence, instead of getting smart for the length of one conversation.
Why this matters more than the chat interface
It’s tempting to talk about this as a search or UX improvement: “now you can ask your archive questions!” That undersells it. The real product isn’t the interface, and it isn’t even the answer. It’s the field that gets written back to the asset. It becomes permanent, structured, and remains there for the next person who touches that file, whether that’s next week or three years from now.
For a rights holder sitting on years of footage, that’s the entire ballgame. An asset that’s technically stored but practically un-understood isn’t an asset, it’s a liability with storage costs attached. The moment enrichment runs against it, that calculus flips. Content that nobody could describe becomes content you can search, clip, license, and pitch.
That’s also the honest pitch, if I’m being direct about it. Most customers don’t know what they’re sitting on. They know they have footage. They don’t know it has moments in it worth monetizing until they see it enriched once. That first look, three hours of undifferentiated footage turning into a set of flagged, explainable moments, is the most convincing thing we can show anyone, more convincing than any feature list.
We’re still early in rolling this out across a wider range of content types, but running it against real long-form sports footage told me something the roadmap slides couldn’t. The gap between what a file is and what a file could be is bigger than most customers realize, and closing it is the actual product.
Further Reading:
From Sticky Note to Shipped: How AI is Changing Product Management
Stop Managing Your Media and Start Asking It Questions
One Pipeline, Infinite Possibilities: The Case for Domain-Specific AI in Media





