Bulk Metadata Generation: How to Tag Thousands of Files Without Losing Quality

Tagging at volume without quality collapsing. The batch workflow, where automation holds up, where human review is still required, and what to verify.

# Bulk Metadata Generation: How to Tag Thousands of Files Without Losing Quality The reason most stock portfolios stall is not a shortage of images. It is a folder of finished work that never got tagged, because tagging is the step where enthusiasm goes to die. At five to twenty minutes per image, a 500-file batch is several weeks of work. Most contributors do not do it, which is why so many portfolios are a fraction of the size they could be. ## Quality degrades with manual volume The problem is not only speed. Manual keywording gets measurably worse as a session goes on. By roughly the twentieth image, most people are recycling the same safe terms, skipping conceptual keywords, and writing thinner descriptions. The last hundred files in a batch receive noticeably weaker metadata than the first ten — and since metadata determines whether a file is ever found, that fatigue converts directly into lost licences. This matters more than it sounds. It means the tail of every batch you have ever uploaded is underperforming, and you cannot see it because the files look fine. Generated metadata does not have this property. File one thousand receives the same treatment as file one. That consistency, rather than raw speed, is the strongest argument for automating the first pass. ## The batch workflow The structural change is moving from *authoring per file* to *reviewing per file*. **1. Work in batches.** Group by shoot or theme. A coherent batch produces more consistent metadata and is faster to review, because you are checking against one mental model rather than switching context every file. **2. Generate the first pass for the whole set.** Titles, descriptions, keywords across every file at once. This is the step that used to be weeks. **3. Review as a grid, not a form.** Looking at fifty results together surfaces systematic problems — a location consistently wrong, a repeated misidentification — that reviewing one at a time hides. **4. Correct what needs outside knowledge.** This is the irreducible human step, covered below. **5. Embed IPTC.** So metadata travels inside the file to every destination. **6. Export per-agency formats.** CSV where the platform uses one, with the correct fields and flags. **7. Deliver, then verify before the next batch.** Confirm ingest and metadata landed correctly. Finding a systemic error across three hundred files is recoverable; finding it across eight thousand is not. ## What automation cannot know Being specific about this matters, because the failure mode of automated keywording is confident wrongness rather than blank fields. A vision model reads pixels. It cannot know: **Specific places.** It can identify "coastal town with pastel buildings." It cannot know it is Burano rather than somewhere similar. **Editorial facts.** Dates, event names, identified individuals. Editorial content lives or dies on these, and they are not in the image. **Precise species.** "Bird" is reliable; the exact subspecies frequently is not. **Brands and products.** Often identifiable, often subtly wrong, and wrong here has commercial-release consequences. **Intended meaning.** For abstract or conceptual work, the concept you were commissioned to express is in your head, not in the frame. The workflow that holds up is AI-first, human-reviewed: let the tool do the exhaustive, fatiguing work, then spend thirty seconds per file on the things only you know. That is a fundamentally different cost structure from authoring from scratch. ## Embedded versus CSV Use both, because they do different jobs. **Embedded IPTC** lives inside the file. It travels to every agency automatically and is the only route Adobe Stock offers — Adobe has no CSV import for metadata. **CSV** handles platform-specific fields that IPTC cannot carry: Shutterstock's category assignments and editorial flags, Freepik's mandatory `_ai_generated` tag. Embed for portability, layer CSV for platform-specific fields. Our [Shutterstock CSV template guide](/blog/shutterstock-csv-template) covers one platform's format in detail. ## One keyword set does not fit all agencies This is where multi-agency contributors lose files: - **Adobe Stock**: 49 keywords, IPTC only, no CSV route - **Shutterstock**: 50 keywords, categories from a fixed list, CSV after upload - **Getty and iStock**: controlled vocabulary, no externally generated AI at all - **Freepik**: mandatory `_ai_generated` flag on generative work - **Alamy**: no keyword cap, AI banned by contract from 1 September 2026 Copying one list everywhere produces rejections in some places and invisibility in others. Per-agency metadata is not an optimisation — it is the minimum for multi-agency distribution to work. ## What this changes If manual tagging caps you at 200 uploads a month, automation plausibly takes you to 2,000. Stock income scales close to linearly with the number of well-tagged files live across agencies, so that is not a 10x time saving — it is a 10x change in the size of the asset base you are building. Rastock AI [generates metadata for whole batches, embeds IPTC, exports per-agency CSVs and delivers over FTP](/features) to 10+ agencies in one pass. Our guide to [AI metadata generation](/blog/ai-metadata-generator-for-stock-photos) covers accuracy in more depth, and plans by upload volume are on the [pricing page](/pricing). ## Summary Batch rather than iterate. Generate first, review second. Accept that automation cannot know place names, dates, species or your intent — and build the thirty seconds of human review into the process rather than hoping it is unnecessary. Embed IPTC for portability, add CSV for platform-specific fields, and never ship one keyword list to six agencies. For the full picture, see our guide to the [stock photography workflow](https://rastock.ai/blog/stock-photography-workflow) — every stage from selection through metadata to delivery. ## Related reading - [AI Metadata Generator for Stock Photos: Why Manual Tagging Is Dead](https://rastock.ai/blog/ai-metadata-generator-for-stock-photos) - [Stock Photography Workflow: From Shoot to Sold](https://rastock.ai/blog/stock-photography-workflow) - [Bulk Uploading to Vecteezy: FTP Delivery and Batch Metadata](https://rastock.ai/blog/vecteezy-bulk-ftp-strategy)

Frequently asked questions

How do I keyword thousands of stock photos efficiently?

Process in batches rather than file by file: generate metadata for the whole set, review as a grid to catch systematic errors, correct the files needing outside knowledge, then embed and deliver. The per-file step should be review, not authoring.

Does keyword quality drop when working at volume?

With manual tagging, measurably — thoroughness declines after roughly twenty images in a session, with thinner descriptions and recycled terms. Generated metadata does not degrade with volume, which is the main argument for automating the first pass.

What still needs human review in automated keywording?

Anything requiring knowledge not visible in the pixels: specific landmarks and place names, editorial dates and events, named species, brand and product identification, and the intended concept behind abstract or conceptual work.

Should metadata be embedded or supplied by CSV?

Both, where the platform supports it. Embedded IPTC travels inside the file to every agency, while CSV handles platform-specific fields like Shutterstock categories and Freepik's AI flag that IPTC cannot carry.

How large should a metadata batch be?

Large enough that automation pays off, small enough to verify before committing the next one. A few hundred files is manageable — you can confirm the output is correct before scaling, rather than discovering a systematic error across thousands.

Can one keyword set be used across all agencies?

No. Keyword limits, category systems and required flags differ — Adobe caps at 49, Shutterstock at 50 with fixed categories, Freepik needs an _ai_generated flag, Getty uses controlled vocabulary. A shared set causes rejections in some places and poor discovery in others.