Stockflow
“Eliminating mechanical cataloging friction for high-volume photography through strict multimodal schema pipelines.”
An end-to-end multimodal automation suite for batch photo renaming and marketplace metadata generation using Gemini Vision.
- Role
- Sole Creator & AI Engineer
- Context
- 3 Weeks (Production Tool)
- Team
- Solo Project
- Core Stack
- Python 3.11, Google Gemini Vision, API Key Pooling, CSV Engine
Fig 1.0 — Architecture execution snapshot (Stockflow)
The Friction
Why build an AI pipeline for stock photography metadata?
For photographers working in high-volume macro and landscape photography, capturing images is creative, but preparing them for marketplace ingestion is agonizing mechanical labor. Stock agencies like Shutterstock require strict 20–40 keyword tags, two categorized taxonomies, and clear editorial descriptions for every single photograph.
Existing commercial tagging software either uses outdated non-contextual computer vision or requires expensive per-image subscription tokens. I needed an automated CLI pipeline that could inspect high-resolution RAW exports, understand nuanced micro-details (like insect anatomy or botanical lighting), and output 100% compliant marketplace CSV sheets without human intervention.
Deliberate Constraints
The system architecture was not chosen in an unconstrained vacuum. Each structural decision emerged directly from four non-negotiable technical boundaries.
Free-tier and standard multimodal API endpoints enforce tight RPM (requests per minute) caps on large photo folders.
Engineered a round-robin API key pool (GEMINI_API_KEYS=k1,k2,k3) coupled with exponential backoff retry logic to maintain continuous batch throughput.
Stock agencies instantly reject commercial submissions containing visible or hallucinated trademark names (Sony, Canon, Nike, Apple).
Built a post-generation regex inspection layer that scrubs all suggested keywords and descriptions against an extensive commercial brand denylist.
Shutterstock accepts only exact matches from their 26 official category strings (e.g. 'Animals/Wildlife', 'Nature', 'Signs/Symbols').
Engineered a normalization lookup dictionary that maps arbitrary AI category predictions to verified agency taxonomies.
High-resolution camera files must be analyzed without modifying pixel data or stripping EXIF camera metadata.
Separated visual analysis from file modification: Step 1 renames filenames on disk via semantic slugs; Step 2 creates an external CSV mapping without touching raw file streams.
System Architecture & Data Pipeline
A decoupled two-stage CLI pipeline: Step 1 (Visual Renamer) examines image composition via Gemini Vision to generate standardized snake_case slugs; Step 2 (Metadata Generator) batches images, queries the model with structured prompt contracts, validates brand safety, and outputs ingest-ready CSVs.
Alto, Cultus, Corolla. Standard tiered rental base.
Audi A6, BMW 7, Land Cruiser. Chauffeur insurance rate.
Sportage, Tucson, Fortuner. All-terrain security deposit.
Bolan, Hiace, Coaster. High-capacity commercial rate.
Dynamic Polymorphism at Runtime: The orchestrator holds a single container std::vector<Vehicle*> fleet. When executing reservations or computing quotes, method calls to v->calculateCost(days) dynamically dispatch to the concrete subclass implementation through each instance's vtable pointer.
Subsystem Decomposition
AI Visual Renamer (Step 1)
image-file-rename-scriptScans raw directory folders and generates concise, descriptive, SEO-rich file names based on visual geometry.
API Key Pooling & Rate Handler
gemini_client.pyRotates through an array of API credentials to maximize throughput while respecting Google AI Studio rate limits.
Multimodal Metadata Generator (Step 2)
stockflow-scriptProduces descriptive titles, 20–40 ranked keywords, and dual category classifications per photo.
Validation & Ingestion Exporter
validation.py & CSV WriterEnforces agency guidelines, scrubs brand names, and generates the final CSV file.
The Hard Part: Multimodal Hallucinations & Commercial Trademark Violations
Preventing AI vision models from inventing brand names or violating strict marketplace metadata schemas.
Stock agencies impose severe penalties—including account bans—for commercial submissions containing copyrighted brand names in keywords or descriptions. Multimodal vision models frequently over-attribute equipment names (like guessing 'Canon EOS' or 'Sony G Master') based on macro depth-of-field cues, even when no logos are present.
Because foundation models associate high-end photography terms with camera brands in their training weights, they instinctively inject manufacturer names into keyword lists, instantly invalidating the submission.
# Multi-tier brand denylist validation
BRAND_DENYLIST = load_brand_denylist() # Sony, Canon, Nikon, Apple, Nike...
def validate_row(filename: str, description: str, keywords: list[str]) -> tuple[bool, list[str]]:
warnings = []
# 1. Inspect Description for trademark mentions
for brand in BRAND_DENYLIST:
if re.search(r'\b' + re.escape(brand) + r'\b', description, re.IGNORECASE):
warnings.append(f"Description contains forbidden trademark: '{brand}'")
# 2. Filter Keyword Array
clean_keywords = [
kw for kw in keywords
if kw.lower() not in BRAND_DENYLIST and len(kw) >= 3
]
# 3. Enforce Agency Quota (20-40 keywords)
if len(clean_keywords) < 20:
warnings.append(f"Insufficient keywords ({len(clean_keywords)}) after filtering")
return len(warnings) == 0, clean_keywordsWe implemented a multi-stage validation layer: first, the prompt explicitly instructs the vision model to describe optical attributes without equipment names; second, post-processing filters every keyword and description against a comprehensive trademark denylist, flagging any violations for manual inspection.
Prompt engineering in production is not about creative phrasing; it is about building defensive type-checking and schema contracts around nondeterministic probabilistic engines.
Batch Ingestion & Shutterstock CSV Compilation
Console execution of Stockflow processing an archive of 24 macro photography captures with key pooling and brand safety checks.
Engineering Reflection
“AI is most valuable not when it generates synthetic content, but when it removes the mechanical friction that suffocates human creative work.”
As a photographer, spending hours typing descriptive tags and double-checking Shutterstock category dropdowns felt like a waste of human energy. Building Stockflow gave that time back.
The project taught me that the hardest part of building AI software is rarely the model itself—it is the defensive engineering around the model: rate limits, schema validation, brand safety, and edge-case handling.
When an automated pipeline works properly, the technology disappears and you are simply left with the photographs.
Interested in discussing this architecture?
I'm always open to technical dialogue, code reviews, and exploring system constraints.