Footage Metadata Consistency: Tagging Video Series and Collections

How to keep metadata consistent across a shoot so related clips surface together, and which technical fields editors actually filter on.

# Footage Metadata Consistency: Tagging Video Series and Collections Most footage metadata advice treats each clip as an independent asset. For video that is the wrong unit of analysis, because video editors rarely want one clip. They are cutting a sequence, and they need several shots that go together. If your related clips carry inconsistent metadata, the buyer finds one, uses it, and never discovers the four others from the same shoot that would have completed their edit. You license one clip instead of five. ## The core set and the differentiator The technique is straightforward: split your keywords into two layers. **The core set** is identical across every clip from the shoot. It covers what does not change: subject, location, setting, time of day, season, mood, and the broad concept. If you shot a morning market in Lisbon, every clip carries the same terms for market, Lisbon, Portugal, morning, street food, vendors, local commerce. **The differentiator** is what makes each clip distinct: shot type, camera movement, specific action, framing. One clip is a wide establishing shot; another is a close-up of hands exchanging money; another is a slow tracking shot past stalls. Applied consistently, this means a buyer searching for your subject finds the whole set, then chooses between clips based on the differentiating terms. That is exactly how an editor works — find the scene, then pick the shots. The failure mode is writing metadata clip by clip from scratch, which produces drift: the first clip says "market," the fifth says "bazaar," the eighth says "street vendors." All accurate, none consistent, and the set never surfaces together. ## Technical fields are filters, not description Before an editor looks at a single frame, they filter. Resolution, frame rate, duration, sometimes codec and colour profile. If those fields are missing or wrong, you are excluded from the result set before your work is judged on its merits. **Resolution.** 4K is the practical commercial baseline in 2026, and editors filter for it routinely. Higher resolutions are a genuine selling point because they give buyers reframing headroom in a 4K or HD timeline. **Frame rate.** Matters for both matching a project's timeline and for slow-motion use. A 60fps or 120fps clip has value a 24fps clip does not, but only if it is declared. **Duration.** Editors need enough handles to cut. Very short clips are frequently unusable regardless of how good they look. **Log or graded, alpha channel.** Both change how a clip fits into a workflow, and both are worth declaring explicitly. These are not metadata polish. They determine whether your clip is in the running. ## Similar content versus a series There is a tension worth naming. Agencies reject for similar content, but a series is by definition multiple clips of the same subject. The distinction reviewers apply is whether the clips are *coverage* or *takes*. Coverage means genuinely different framing, movement or action — a wide, a medium, a detail, a moving shot. Takes means several near-identical versions of the same shot. Submit coverage. Cut the takes. A well-constructed series of five distinct shots is valuable to a buyer; five variations of the same push-in is redundant catalogue and invites a rejection across the batch. ## Doing this at scale The core-set-plus-differentiator approach is easy to describe and tedious to execute by hand, because the whole point is that the core set has to be *identical* across dozens of clips. Retyping it is exactly where drift creeps in, and a single inconsistent term breaks the grouping you were trying to build. This is a natural fit for batch processing: define the shared metadata once at the shoot level, generate per-clip differentiators, and apply both consistently. Rastock AI [generates footage metadata in batches and delivers over FTP](/features), which keeps the shared terms genuinely shared rather than approximately shared. Our [Pond5 contributor guide](/blog/pond5-video-metadata-automation) covers how editors search for individual clips, and the [Storyblocks guide](/blog/storyblocks-video-collection-ftp) covers how subscription payouts change which footage is worth producing. Plans by upload volume are on the [pricing page](/pricing). ## Summary Tag the shoot, not just the clip. Keep a core keyword set identical across the series so it surfaces together, and differentiate on shot type, movement and action. Declare technical fields accurately, because they filter before anything else. And submit distinct coverage rather than multiple takes. Done properly, one shoot licenses as a set rather than as a single lucky clip. This guide is part of our overview of [where to sell stock photos](https://rastock.ai/blog/where-to-sell-stock-photos), which compares every major agency on royalties, review standards and AI policy. For the full picture, see our guide to the [stock photography workflow](https://rastock.ai/blog/stock-photography-workflow) — every stage from selection through metadata to delivery. ## Related reading - [Pond5 Contributor Guide: Stock Footage Metadata That Sells](https://rastock.ai/blog/pond5-video-metadata-automation) - [VectorStock Contributor Guide: EPS, JPG and Vector Metadata](https://rastock.ai/blog/vectorstock-metadata-for-illustrators) - [MotionArray Contributor Guide: Tagging Templates and Motion Assets](https://rastock.ai/blog/motionarray-technical-tagging)

Frequently asked questions

Why does metadata consistency matter for video?

Editors often need several clips from the same setup to cut a sequence. If related clips carry inconsistent keywords, a buyer finds one and never discovers the others, so you license one clip instead of five from the same shoot.

How do I tag a video series consistently?

Define a shared core keyword set for the shoot — subject, location, setting, mood, project identifier — and apply it to every clip. Then add per-clip terms for what makes each one different: shot type, movement, action.

What technical fields do video buyers filter on?

Resolution, frame rate, duration and codec, plus whether footage is log or graded and whether it has an alpha channel. These are filters applied before browsing, so inaccurate or missing values remove you from consideration entirely.

Is 4K required for stock footage in 2026?

It is the practical baseline for commercial work. Editors routinely filter for it, and delivering below 4K narrows your buyer pool before content is evaluated. Higher resolutions also give buyers reframing headroom, which is a selling point.

Should every clip in a series have identical keywords?

No. Keep the core set identical so the series holds together in search, then differentiate with shot-specific terms. Identical keywords across visually different clips make them compete with each other rather than complement each other.

How do I avoid similar-content rejections with a series?

Submit clips that are genuinely distinct in framing, movement or action rather than near-identical takes. A series should read as coverage of a subject from different angles, not as multiple versions of the same shot.