by Brian Campanotti, Next Generation Archive Visionary at Cloudfirst
Last month I sat on a panel with TwelveLabs, AWS, and Iconik to discuss something that I think a lot of people in this industry privately know but rarely say out loud – the archive problem isn’t really a storage problem, and it never was. Storage is just where the conversation usually starts and, for most organisations, where it also tends to stop.
What we were actually talking about was value. Specifically, how much of it is sitting completely untouched in archives that were built to preserve content rather than to use it, and what it would actually take to change that.
Legacy systems weren’t built for what we’re trying to do now
I truly believe that the legacy MAM and archive systems most organisations are still running were never designed to do what we’re now asking of them. They were built to store content and retrieve a specific file when you knew exactly what you were looking for. That was the brief, and for a long time it was perfectly adequate. The problem is that the world around those systems has changed completely and the systems themselves largely haven’t.
Archives have grown at a pace nobody really anticipated. Higher resolutions, more formats, multi-geography distribution, content being produced across more channels than ever, and the infrastructure managing all of it hasn’t kept pace. So organisations end up sitting on petabytes of content they can’t properly search, can’t accurately value, and genuinely struggle to monetise. Work in progress footage, raw rushes, regional cuts, decades of material that was never properly indexed and has never really had a chance to earn. And because these systems are proprietary by design, changing that feels like an enormous undertaking. That feeling is exactly why most organisations don’t start, not because the problem isn’t real, but because the path forward seems harder than it actually is.
The first thing I always want to know
When I sit down with a new client, the first question I ask isn’t about what technology they want to move to or what their budget looks like. It’s a much simpler one; do you actually know what you have? And more often than not, the honest answer is not really. They might know roughly how much storage they’re consuming or have a sense of their top-level folder structure, but the real picture, the formats, the condition, the utilisation patterns, what’s been accessed and when, what’s duplicated, what’s genuinely unique and irreplaceable, that level of detail is almost always missing.
This is where we always begin, and I can’t overstate how important it is to start here rather than jumping straight to solutions. Before a single asset moves anywhere, our data sciences team goes deep on the existing infrastructure to build that picture properly. The goal is to get every assumption out of the room so that when an organisation makes a decision about their future archive strategy, it’s grounded in what’s actually happening rather than what someone thinks is probably happening. That distinction sounds obvious but it changes everything, including where the real costs are, where the risk actually sits, and what the genuine opportunity looks like on the other side. It also reframes what the whole exercise is for, because this should never just be a data lift and shift. Discoverability is the point, and you can’t build a credible discoverability strategy without first knowing what you have to work with.
On the question of downtime
The objection I hear most often is some version of “we can’t afford to touch the archive while everything is running.” I understand why. There are live operations to consider, content moving in and out every day, and the idea of introducing a migration project into that environment is genuinely anxiety-inducing. But it is a solvable problem.
When the NHL came to us needing to migrate 15 petabytes of archive content housing over a century of hockey history to a hybrid AWS architecture, zero production impact wasn’t a preference or a target. It was the only outcome that was acceptable to them. We built the entire programme around that constraint and we delivered it. Fully managed, end to end, with their existing workflows completely untouched from start to finish. Trece, which is part of the COPE Group, had a different priority. They were done with being beholden to a legacy vendor and wanted direct, unconditional control of their assets back. Again, fully managed migration, no disruption to operations, and no new lock-in waiting for them on the other side.
I share these not to make a pitch, but because I genuinely think a lot of organisations are sitting on the sidelines based on a fear that careful planning and the right partner can address. The downtime concern is legitimate, but it doesn’t have to be the reason nothing happens.
What actually changes when your content is properly indexed
Once content is out of proprietary silos and indexed properly, what becomes possible is genuinely different from anything these organisations have been able to do before. AI models trained specifically on video, not general-purpose models but ones built to understand video at a deep level, can move through an archive at scale and enable something that simply wasn’t on the table before. The ability to search across content that’s never been manually tagged, without needing to know in advance what you’re looking for. Entity search that can surface a specific person, object, action, or theme across decades of footage without any prior indexing work. That capability alone changes the conversation about what an archive is actually worth.
What that opens up on the monetisation side often surprises people. The obvious buyers are obvious: broadcasters, distributors, the usual suspects. But properly indexed archive content consistently attracts buyers that organisations have never thought to approach. Once you add AI-powered dubbing and subtitling into the picture, content that was only ever relevant to one market can suddenly have an audience in five. The content flywheel that becomes possible when rich metadata is driving automated workflows, and those workflows are driving discovery, and that discovery is generating revenue, that’s a fundamentally different way of operating than what most archive owners are doing today. It’s not incremental improvement, it’s actually a different model entirely.
What the programme makes possible
One of the things I was glad we could talk about openly on the panel was the joint Video Understanding Programme AWS has put together with Cloudfirst, TwelveLabs, Iconik, and a few other partners. It exists specifically to address the two things that most consistently stop organisations from taking the first step: cost and complexity. For qualified organisations it includes up to $25,000 in migration assessment funding, promotional indexing pricing, free re-indexing as AI models continue to improve, and MAP funding eligibility that can be stacked alongside the other benefits. Migration support comes through qualified partners and Cloudfirst is one of them.
The whole thing is built on open standards and API-first architecture, because one of the principles we’ve always operated on is that when we’re done with a migration, we’re genuinely done. Clients leave with full ownership and control of their assets, with no ongoing dependency on us or on any other vendor. That’s not a selling point. It’s just the right way to do this work. We’d rather have organisations come back to us because we did a good job than because we’ve made ourselves hard to leave.
Why now matters
Demand for licensable, AI-ready video content is growing, and the organisations that are building the infrastructure to meet that demand today are the ones who will have the relationships and the capabilities that everyone else will be scrambling to catch up to in a few years. The window is open right now in a way that it genuinely might not stay. The funding programmes, the pricing, the partnerships we’ve built to make this accessible, none of that is permanent.
If you’ve been having the internal conversation about your archive and haven’t found the right way to start, my honest advice is to begin with the data. Understand what you actually have. Model what it costs today and what it could generate tomorrow. Everything else gets a lot clearer from there, and you might be surprised by what’s already sitting in your vaults waiting to earn.