← All posts

Making Product Images and Video AI-Readable

Text markup alone isn't enough — a practical guide to describing product images and video so AI shopping agents can actually use them.

TL;DR

  • Most AI-readability work so far has focused on text: JSON-LD, llms.txt, product descriptions. But a meaningful share of what makes a shopper choose one product over another — material, color accuracy, fit, scale, how something actually looks in use — lives in images and video, not text.
  • AI agents can't reliably "see" an image the way a human shopper does unless the surrounding data tells them what's in it. A photo with no alt text, no structured image schema, and no descriptive caption is close to invisible to most AI retrieval, even if a human finds it obviously informative.
  • Video is even further behind. Very few Shopify product pages provide any machine-readable summary of what a product video actually shows, so that information is effectively unavailable to an AI agent forming a recommendation.
  • This is a natural extension of the same JSON-LD/llms.txt foundation already covered — it's the same underlying discipline (structured, explicit, machine-readable data) applied to a different content type.

Why isn't a good product photo automatically AI-readable?

An AI agent parsing a product page doesn't "look" at an image the way a person does — it relies on the surrounding markup and text to know what the image contains: alt attributes, Product schema image fields, captions, and any structured additionalProperty values that describe visual attributes explicitly (color, material, pattern). A page can have excellent photography and still be functionally silent to an AI agent if none of that supporting text exists, because the image itself doesn't carry meaning an AI retrieval process can extract without it.

This is the same "content structure vs. content quality" gap covered in the ACCC framework, just applied specifically to visual content rather than body text.

What does an AI-readable product image actually require?

At minimum, three things need to line up: a descriptive alt attribute that states what the image shows in plain language (not just the filename or a generic label), a Product schema image field pointing to the actual asset, and — where the visual detail matters to a buying decision — an explicit text callout of what the image demonstrates (e.g., "shown in Walnut finish" or "true-to-color under natural light"). Without the third piece, an AI agent may know an image exists but still can't tell what specific claim it's supporting.

What about product video — does the same logic apply?

Largely yes, but video has an added gap: even a well-tagged video file rarely comes with any text summary of what it actually shows minute-by-minute. A VideoObject schema entry with a description field, plus a short text summary of what the video demonstrates (assembly steps, scale comparison, a 360° view, a specific use case), gives an AI agent something concrete to retrieve — otherwise the video is effectively unindexed content, no matter how useful it would be to a human watching it.

What does the gap between a "nice photo" and an "AI-readable photo" actually look like?

ElementPresent on most product pages?What makes it AI-readable
Descriptive alt textOften missing or generic ("product photo 1")States the specific visual claim: material, color, angle, context of use
Product schema image fieldUsually present for the primary imageShould point to all meaningful images, not just the hero shot
Explicit visual-claim calloutsRarely presentA short text statement of what the image proves (e.g. "actual size shown next to a standard coffee mug")
Video schema + summaryAlmost never presentVideoObject schema plus a plain-text description of what the video shows

How should a Shopify merchant actually fix this?

  1. Audit existing product images for alt text quality, replacing generic or missing descriptions with specific, factual statements of what each image shows.
  2. Make sure Product schema covers all meaningful images, not just the primary hero shot — secondary angles, scale references, and material close-ups all carry real buying-decision information.
  3. Add explicit visual-claim callouts in text next to images that are doing real persuasive work (fit, scale, color accuracy) rather than assuming the image speaks for itself.
  4. Tag product videos with VideoObject schema and a plain-text summary of what the video demonstrates, so an AI agent has something to retrieve even if it can't process the video content directly.

What does this actually look like once a merchant runs a scan?

A catalog scan is what turns "images and video probably have gaps" into a specific, fixable list. On a typical Shopify catalog, a structured scan of image and video markup usually surfaces the same handful of patterns: hero images with generic or missing alt text ("product photo 1" instead of a description of what's shown), secondary angle and material-detail shots that were never added to the Product schema's image field at all, and product videos sitting on the page with no VideoObject entry or text summary — meaning an AI agent has no way to know what those videos demonstrate, no matter how useful they'd be to a human watching them.

Once those gaps are fixed — descriptive alt text added, every meaningful image registered in schema, video tagged with a VideoObject description — the practical change is that an AI agent parsing the page now has explicit, machine-readable answers to questions it previously had no way to resolve from a photo or clip alone: what material or color a product actually is, what scale it is relative to a familiar object, what a video actually shows step by step. That's the difference between a page that reads as visually rich to a human and one that's also legible to the AI agent evaluating it on a shopper's behalf.

The reason this is worth checking even if a merchant thinks their photography is already strong: whether an AI agent can use an image has very little to do with how good the photo is and almost everything to do with whether the surrounding markup exists. A merchant can't tell this by looking at their own site, since the gap only shows up in what the underlying data says, not in what the storefront looks like.

Where DeepLumen fits

DeepLumen positions Agentic Page as the layer that extends AI-readable structured data beyond text — generating the same kind of explicit, machine-readable markup for a Shopify catalog's images and video that it already provides for product facts, so a store's visual content isn't left out of what an AI agent can actually retrieve and reason about. This is the same underlying mechanism behind DeepLumen's broader AI-visibility results across its 700+ installed Shopify merchants (structured-data fixes → AI citation and conversion lift) — applied here specifically to the image/video layer that most catalogs haven't touched yet.


Check whether your product images and video are actually contributing to your AI visibility, or sitting invisible next to markup that only covers text. Learn more about Agentic Page or book a demo for a full catalog scan.