How to run an e-commerce AI visibility audit

An e-commerce AI visibility audit measures whether AI answer engines name and recommend your products for the questions shoppers actually ask. You map shopper moments, design prompts across the buying journey, test them across engines and regions, record inclusion and recommendation position, audit the data engines draw on, diagnose why you were left out, and prioritize fixes. The result is a baseline, never a guaranteed placement.
Who this is for: e-commerce and merchandising leaders, DTC founders, and SEO or content teams who want to measure and diagnose how AI engines recommend their products — before committing budget to fixes.
E-commerce AI visibility audit: the short answer
- Start from shopper questions, not keywords. An e-commerce AI visibility audit measures the conversational buying questions shoppers pose to engines, so the prompt taxonomy is the foundation the rest of the audit rests on.
- Separate mentions, recommendations, and citations. Being named in an answer, being recommended as a pick, and having your page cited as a source are three different results, and a useful audit records all three.
- Diagnose the failure, then prioritize by business value. Once you see where you are absent, classify the reason — access, understanding, trust, category fit, authority, freshness, or conversion friction — and fix the products that matter most by margin, inventory, and demand first.
What the audit measures
Traditional SEO audits examine crawlability, indexation, and ranking positions on a search results page you can see. An e-commerce AI visibility audit asks something different: when a shopper asks Google AI Overviews, Gemini, ChatGPT, Perplexity, or Claude "what should I buy for X," does the engine name your product, does it recommend it as a pick, and where does it place you relative to competitors? The generated answer is the surface, not a leaderboard of blue links.
Set expectations before you begin. An audit is a measurement exercise. It tells you what engines are doing right now for a defined shopper-prompt set, and it does not — and cannot — promise that a given fix will earn a recommendation or move a product into an answer. The engines do not document how they select which products to name, and their outputs vary between sessions, models, regions, and dates. Treat every reading as a bounded snapshot: accurate for the prompt, engine, and moment you captured it, and nothing more. That discipline is what makes a baseline worth building, because you are comparing like-for-like snapshots over time rather than chasing a number no engine publishes. This audit is the diagnostic companion to the broader e-commerce brand visibility on AI strategy; if you are still deciding whether measurement belongs in your stack at all, the AI visibility strategy overview frames where it fits before you commit.
Audit inputs: what to gather first
Before you run a single prompt, assemble the material the audit will test against. An e-commerce AI visibility audit reconciles what engines say with what your own systems contain, so the inputs list determines how far the diagnosis can reach. Gather:
- Catalog and SKU list — the products in scope, grouped so you can measure by category and SKU group rather than one item at a time.
- Merchant feeds — the product feeds syndicated to marketplaces, ads, and comparison surfaces, with their refresh cadence.
- Product pages — the PDP content and structured data for the SKUs in scope.
- Category and collection pages — how the catalog is framed for discovery and comparison.
- Reviews — first-party review content and the third-party review surfaces that carry weight in your category.
- Policy pages — shipping, returns, and warranty content that gets pulled into price and post-purchase answers.
- Inventory and availability — current stock and availability state, so you can spot stale answers.
- Locations — where availability, delivery, or store data varies by region.
- Analytics — referral and assisted-conversion data, including any AI-assistant referrers you can already attribute.
Log each input's owner and last-updated date as you collect it; that record is what turns "the engine is wrong" into "the engine is reading a feed we updated three weeks after the page." This is the diagnostic side of the broader AI visibility engine and services picture — the audit produces the observations that operational systems then act on.
The E-Commerce Recommendation Readiness Audit
This is a nine-step framework. Work it in order; each step feeds the next.
- Map shopper moments and categories. List the moments a shopper reaches for an engine and the categories those map to, so your prompts cover real demand rather than your internal taxonomy.
- Design the prompt taxonomy. Separate discovery, use-case, comparison, price, feature, constraint, and post-purchase prompts — each surfaces a different slice of engine behavior.
- Test fixed questions across models, regions, and dates. Hold the wording constant and vary the engine, region, and date so differences you record are engine behavior, not prompt drift.
- Record the result structure. For each answer capture inclusion, recommendation position, competing products named, and the framing used to describe you and rivals.
- Audit the cited sources. Note whether engines lean on your product pages, category pages, editorial content, marketplace listings, review sites, or third-party evidence.
- Validate the underlying data. Check product facts, structured data, merchant feeds, availability, and policy pages against what the engine seems to "know."
- Diagnose the failure type. Classify each gap as access, understanding, trust, category fit, authority, freshness, or conversion friction.
- Prioritize by business value. Rank fixes by margin, inventory depth, demand, and buyer value — not by which gap is easiest to close.
- Remeasure on a cadence. Re-run the same prompts on a schedule so you can tell whether anything actually changed after you acted.
Mentions versus recommendations versus citations
The single most common mistake in an e-commerce AI visibility audit is collapsing three distinct results into one. They come apart constantly, and each carries a different implication.
A mention means the engine named your brand or product somewhere in the answer — perhaps in a list of "options people consider," perhaps as an also-ran. A recommendation means the engine actively put you forward as a pick, often near the top of an ordered answer, sometimes with a reason. A citation means the engine linked to or attributed one of your pages as a source, regardless of whether it recommended you. An engine can recommend a competitor warmly while citing your specification page as the evidence for a claim; it can mention you in passing while recommending a rival at the top. Recording only one of the three gives you a partial and often misleading read on recommendation readiness.
Shopper-question taxonomy (anonymized methodology example)
Prompt design is where audits succeed or fail, so it deserves a worked example. The following is an anonymized methodology example that illustrates how to design prompts — it is a demonstration of the technique, not a performance case study, and no results are implied.
Imagine an anonymized garment-care brand. Instead of testing one broad prompt, you decompose the category into the questions a real shopper asks at each moment:
- Discovery: "best product for travel" — surfaces whether the engine even considers the brand in a broad, unconstrained ask.
- Constraint: "best for delicate clothing" — tests whether the engine matches a specific requirement to the right product.
- Feature/comparison: "which handheld option heats quickly" — tests whether the engine understands a feature-level distinction between models.
Running fixed prompts like these across engines, regions, and dates lets you see not just whether the brand appears, but whether it appears for the right reasons — matched to the constraint the shopper actually stated. When a brand shows up for "best product for travel" but vanishes for "best for delicate clothing," that is a diagnostic signal, not a mystery: the engine likely lacks a clear, structured association between the product and the "delicate" use case. The taxonomy turns a vague "are we visible?" into a set of testable, decision-useful questions. For a catalogue of the framing errors this exercise tends to expose, see the companion piece on common AI brand visibility mistakes.
The shopper-question matrix
Capture your prompt tests in a single matrix so patterns are visible at a glance. Use word ratings and source-type labels rather than invented figures. The row below uses the anonymized garment-care example.
| Category | Use case | Constraint | Comparison | Price | Evidence needed | Likely source type | Result | Priority |
|---|---|---|---|---|---|---|---|---|
| Garment care | Travel packing | Compact, dual-voltage | vs. two rivals | Under a stated budget | Spec + use-case content | Product page, editorial | Partial (mentioned, not recommended) | High |
| Garment care | Delicate fabrics | Low-heat, fabric-safe | vs. one rival | Not stated | Fabric-safety claims, reviews | Review site, third-party | None (absent) | High |
| Garment care | Fast heat-up | Handheld, quick-ready | Feature-level | Mid-range | Feature spec, comparisons | Product page, marketplace | Strong (recommended) | Medium |
Each row is a hypothesis about why a result landed where it did. A "None" in a high-margin, high-demand row is where you focus; a "Strong" you simply protect.
Prompt segmentation by dimension
The taxonomy above is organized by shopper moment; segment it a second way, by the dimension each prompt stresses, so your set has no blind spots. Each dimension exercises a different slice of your catalog data:
- Audience — who is asking ("for a beginner," "for a professional," "as a gift for a teenager"). Tests whether the assistant matches your product to the right buyer.
- Use case — the job to be done ("for travel," "for daily commuting"). Tests use-case content on the page.
- Constraint — a hard limit or requirement ("machine-washable," "under a stated budget," "fits a specific space"). Tests whether the constraint is stated in extractable prose.
- Price — budget-bounded questions ("cheapest," "best value under X"). Tests offer accuracy and feed freshness.
- Compatibility — fit and interoperability ("compatible with model Y," "works with X"). Tests whether compatibility facts are present and correct.
- Availability — stock and delivery ("in stock now," "ships to my region," "available for pickup"). Tests inventory and location data.
- Comparison — head-to-head or shortlist ("X vs Y," "alternatives to Z"). Tests how the assistant frames you against rivals.
A robust set crosses moments with dimensions: a comparison prompt with an availability constraint behaves differently from a comparison prompt with a price constraint, and the difference is diagnostic. Missing whole dimensions — most commonly compatibility and availability — is why an audit that looks thorough can still overstate coverage.
Brand, category, and SKU-level measurement
Because catalogs are large, measure at three levels rather than trying to test every SKU. Keep the levels separate in the results so a pattern at one level does not mask a gap at another.
- Brand level. A small set of brand and reputation prompts — "is [brand] reputable," "who makes X" — testing whether the assistant names you accurately and frames you well. A healthy brand read can coexist with weak product results.
- Category level. For each priority category, discovery, constraint, and comparison prompts that reveal whether the category surfaces you at all. This is where you find categories you have exited from the consideration set entirely.
- SKU-group level. For your highest-value clusters — grouped by margin, inventory depth, and demand — the specific constraint, compatibility, and comparison prompts a real shopper would use. This is the most granular and most actionable level.
Report each level with the mention / recommendation / citation split intact, because a SKU can be cited (its spec page used as evidence) while a rival is recommended. That mention-based logic is the same one behind share of citation — the percentage of relevant answers that name your brand at least once — applied to a commerce prompt set.
Auditing product and category pages
Engines cannot recommend a product they cannot understand. Your product and category pages are the primary evidence surface, so audit them against what the engines seem to know.
- Product pages: Are the core facts — what it is, who it is for, the constraints it satisfies — stated in plain, extractable prose, not buried in imagery or a spec table an engine may not parse well? Does the page answer the use-case and constraint questions from your taxonomy explicitly?
- Category pages: Do they frame the category the way a shopper asks about it ("best for delicate fabrics") rather than only by internal merchandising labels? Category pages often carry the comparison and discovery load.
- Policy pages: Shipping, returns, and warranty pages frequently get cited when shoppers ask constraint and post-purchase questions. Missing or ambiguous policy content is a real gap.
Merchant feeds and structured product information
Structured data helps machines read your catalogue reliably. Google's Product structured-data documentation defines the properties — name, description, availability, price, review data — that make a product listing machine-readable, and Google's structured-data introduction explains how the markup is consumed. Validate that your Product markup is present, valid, and consistent with what appears on the page and in your merchant feeds; mismatches between feed, markup, and page are a common source of "understanding" failures.
State the boundary plainly: valid markup does not guarantee an AI recommendation. Structured data makes your product legible and eligible for structured features, but engines do not document a rule that says correct markup earns a pick. Google's own AI features and eligibility guidance frames these as eligibility signals, not guarantees. Treat structured data as removing a barrier, not as buying a result — verify behaviour against your own measured baseline, and against each vendor's current documentation.
Product facts and the stale-information log
A recurring, high-severity finding is that an engine describes a product with facts that were once true. Availability, price, variants, bundle contents, and policy terms drift, and an assistant reading a cached page or a stale feed will happily repeat last quarter's price. Keep a stale-information log as a first-class audit deliverable: for each engine answer that contains a wrong or outdated fact, record the claim, the current truth, the surface the engine appears to have drawn it from (page, feed, marketplace, or third-party), and the owner who can correct it.
| Wrong/outdated claim | Current truth | Apparent source surface | Owner | Severity |
|---|---|---|---|---|
| Price shown below current | Repriced two weeks ago | Merchant feed (stale) | Feed/merch ops | Strong |
| "Out of stock" for a live SKU | In stock | Cached PDP | Web/platform | Strong |
| Old variant list | Two variants added | Product page + feed mismatch | Merch + web | Partial |
The log matters because stale facts are the failures most likely to cost a sale at the moment of purchase and the ones with the clearest, most defensible fix. It also feeds directly into the routing side of the work — turning each row into a correction task is exactly what AI visibility CRM and content workflows describe operationalizing.
Reviews, publisher authority, and third-party evidence
Step five of the framework — auditing cited sources — matters because engines frequently recommend products by leaning on evidence you do not own. When the source audit shows an engine citing review sites, editorial roundups, or marketplace listings rather than your pages, that is a trust-and-authority signal, not a content-volume one. Ask which third-party surfaces carry weight for your category and whether your product is represented accurately there. Google's guidance on creating helpful, people-first content describes the experience, expertise, authoritativeness, and trust signals that make content dependable — the same qualities engines tend to reward in the sources they draw on. You cannot manufacture authority, but you can find and fix inaccurate third-party representations of your products.
Competitor recommendation analysis
Because the audit records every product named in every answer, it doubles as a competitive map. For each high-value prompt, record which rivals the engine recommends, in what position, and with what framing. This reframes the core question from "are we recommended?" to "how often are we recommended compared with the alternatives shoppers are shown, and for which constraints do rivals consistently win?" A competitor who owns the "delicate fabrics" constraint across engines is telling you exactly where your product data or third-party evidence is thinner than theirs. Competitor recommendation analysis is often the fastest route from a vague sense of underperformance to a specific, testable diagnosis. When a single engine dominates in your category, a deeper single-engine protocol like the Claude AI visibility audit shows how to run priority prompts repeatedly and trace sources under disclosed conditions.
Platform-specific audit paths
The framework is the same across stacks, but where you look for the underlying data differs. Rate each path against your setup.
| Platform | Where product data lives | Feed / markup control | Typical friction |
|---|---|---|---|
| Shopify | Theme templates, metafields | App-managed feeds, theme schema | Markup completeness varies by theme and app |
| WooCommerce | WordPress posts, product plugins | Plugin-managed schema and feeds | Plugin sprawl and inconsistent markup |
| Magento / Adobe Commerce | Catalogue attributes, PDP templates | Native feeds, custom schema | Attribute mapping and template drift |
| Marketplace-first | The marketplace listing itself | Marketplace-controlled fields | Limited control of your own framing and citations |
| Custom build | Wherever your team put it | Fully custom, fully your responsibility | No default markup — everything is on you |
The marketplace-first row deserves a flag: when your primary presence is a marketplace listing, the engine often cites the marketplace rather than you, and your control over framing is limited. That is a category-fit and authority question you should surface early in the audit.
Prioritization scorecard
Not every gap is worth closing, and not in the order they surfaced. Score each candidate fix with word ratings across the dimensions that determine business impact, then work the strongest rows first.
| Product / gap | Margin | Inventory depth | Demand | Buyer value | Fix priority |
|---|---|---|---|---|---|
| Delicate-fabrics gap | Strong | Strong | Strong | Strong | Strong |
| Travel discovery gap | Partial | Strong | Strong | Partial | Partial |
| Niche feature gap | Strong | None | Partial | Partial | Partial |
| Clearance line gap | None | None | None | None | None |
Read the scorecard as a filter. A "Strong" across margin, inventory, and demand is where diagnosis and effort belong; a "None" line is one you note and move past, however easy the fix looks.
Diagnosing the failure type
The audit's value is the diagnosis, not the list of absences. For each priority gap, classify the likely failure:
- Access: engines or their crawlers cannot reach the page (robots rules, rendering, blocked bots).
- Understanding: the page does not state the facts or use case in extractable prose.
- Trust: claims are unsupported or inconsistent across feed, markup, and page.
- Category fit: the engine does not associate the product with the shopper's stated constraint.
- Authority: the third-party evidence engines lean on does not represent you well.
- Freshness: availability, price, or content is stale relative to the answer.
- Conversion friction: the shopper reaches you but the post-click experience or policy content undermines the match.
Each failure type points to a different owner and a different fix, which is why classification precedes action.
Sample action plan
A realistic first plan, drawn from the anonymized example, might read:
- Delicate-fabrics constraint (understanding + category fit): rewrite the relevant product and category pages to state the fabric-safe use case in plain prose; validate Product markup against the page.
- Travel discovery (authority): identify the editorial and review surfaces engines cite for the category and correct any inaccurate representation of the product there.
- Feed consistency (trust): reconcile merchant feed, on-page facts, and structured data so all three agree.
- Remeasure in four weeks: re-run the same prompt matrix and compare snapshots — treat any movement as a bounded observation, not proof of cause.
For the ongoing cadence side of this — how teams operationalize repeated measurement rather than one-off audits — see how teams build workflows around citation monitoring. If you conclude you want managed content, SEO, and GEO execution rather than in-house diagnosis, our partner Percepture offers commerce technical and content remediation. Prime AI Visibility and Percepture have a commercial relationship. Prime provides visibility intelligence and diagnosis; Percepture provides managed implementation. Recommendations and comparisons use the criteria shown on this page.
The 30-day baseline and retest
An audit is only useful if it becomes a baseline you can compare against. Run it as a bounded 30-day sequence rather than an open-ended project, so you finish with a defensible before-and-after rather than a one-off snapshot.
- Days 1–3 — freeze the inputs and the prompt set. Lock the catalog scope, the input list above, and the exact prompt wording. Once frozen, the wording does not change, because a changed prompt is a new experiment, not a remeasure.
- Days 4–10 — capture the baseline. Run the frozen prompts across the engines your buyers use, recording engine, model or product, search state, location, and date for every run. Repeat priority prompts more than once so you can see one-run variability rather than mistaking a single answer for the truth.
- Days 11–14 — diagnose and prioritize. Classify each gap by failure type, populate the stale-information log, and score the prioritization matrix by margin, inventory, and demand.
- Days 15–27 — implement the highest-value corrections with the assigned owners, leaving the prompt set untouched so the retest stays like-for-like.
- Days 28–30 — retest under disclosed conditions. Re-run the identical prompt set and compare snapshots. Read any movement as a bounded observation, not proof of cause: AI answers can vary by platform, model or product, search state, location, prompt wording, time, and repeated run. Results describe a defined observation method, not a permanent universal rank.
The output is a repeatable cadence you can run every 30, 60, or 90 days. Agencies delivering this for clients can fold it into the reusable agency AI visibility audit template; regulated sellers should also read how a healthcare AI visibility audit handles evidence capture under stricter constraints.
What not to do
- Do not test one broad prompt and call it an audit. A single "best garment steamer" query is an anecdote; the taxonomy exists because constraint and comparison prompts reveal different behaviour.
- Do not treat a mention as a recommendation. Being named in a list is not being recommended as a pick — keep the three results separate or you will misread your position.
- Do not assume valid markup buys a placement. Structured data makes you legible and eligible; it does not guarantee a recommendation, and any tool implying otherwise is overreaching.
- Do not chase the easiest gaps first. Prioritize by margin, inventory, demand, and buyer value — closing a low-value gap because it is simple is motion, not progress.
- Do not read a single snapshot as cause and effect. Engine outputs vary by session, region, and date; only a like-for-like remeasure over time supports any claim that something changed. If you want a rubric for choosing measurement tooling, weigh vendors with the AI search visibility services guide, and see how AI shopping optimization platforms fit the execution side.
Methodology and sources
The garment-care brand in this article is an anonymized methodology example used to demonstrate prompt design and diagnosis technique. It is not a performance case study, and no results, rankings, or recommendations are implied. AI engine outputs vary by session, model, region, and date, so every reading described here is a bounded snapshot rather than a stable fact. Claims about engine and structured-data behaviour are bounded to the primary sources cited below; verify vendor behaviour against each vendor's current documentation. This article was authored by Alex Mannine; the methodology was reviewed by Bob Generale. Prime AI Visibility and Percepture have a commercial relationship. Prime provides visibility intelligence and diagnosis; Percepture provides managed implementation. Recommendations and comparisons use the criteria shown on this page.
References
- Google Search Central, Product (structured data). https://developers.google.com/search/docs/appearance/structured-data/product
- Google Search Central, Intro to structured data markup. https://developers.google.com/search/docs/appearance/structured-data/intro-structured-data
- Google Search Central, AI features and your website. https://developers.google.com/search/docs/appearance/ai-features
- Google Search Central, Creating helpful, reliable, people-first content. https://developers.google.com/search/docs/fundamentals/creating-helpful-content
- OpenAI, Product discovery in ChatGPT search. https://openai.com/chatgpt/search-product-discovery/
Next steps
- Start from the e-commerce brand visibility strategy so the audit sits inside a plan for the whole shopper journey rather than as a one-off report.
- Compare AI shopping optimization platforms if you want to see how execution tools map to the gaps your audit surfaces.
- When you are ready, create a Prime AI Visibility workspace and bring your shopper prompts to see how the engines recommend your products today.
Frequently asked questions
What is an e-commerce AI visibility audit?
It is a structured pass that measures whether AI answer engines name and recommend your products for the questions shoppers actually ask, and diagnoses why you are or are not included. It records inclusion, recommendation position, competing products, and the sources engines cite, then classifies the failure type so you can prioritize fixes by business value.
Does correct product structured data guarantee an AI recommendation?
No. Valid markup, as defined in Google's Product structured-data documentation, makes your product machine-readable and eligible for structured features, but the engines do not document a rule that correct markup earns a recommendation. Treat structured data as removing a barrier, not as buying a placement, and verify behaviour against your own measured baseline.
What is the difference between a mention, a recommendation, and a citation?
A mention means the engine named your product somewhere in the answer; a recommendation means it actively put you forward as a pick; a citation means it attributed one of your pages as a source. They come apart routinely — an engine can cite your page while recommending a rival — so a useful audit records all three separately.
How is this different from a normal SEO audit?
An SEO audit checks crawlability, indexation, and ranking positions on a search results page you can see. An e-commerce AI visibility audit measures whether your product is named and recommended inside a generated answer, where there is no ranked list. The objects being measured are different, so one cannot substitute for the other.
How many shopper prompts should I test?
Enough to cover the discovery, use-case, comparison, price, feature, constraint, and post-purchase moments for your priority categories, tested across the engines your buyers use. Depth matters more than volume: a well-decomposed set of constraint and comparison prompts reveals more than a large pile of broad, unconstrained queries.
Can this audit tell me why an engine recommends a competitor?
Not directly. Engines do not document how they select which products to name, so the audit observes the output — who was recommended, cited, and how they were framed — and infers a likely failure type from the pattern. It gives you a testable hypothesis to act on, not a reading of the model's internal reasoning.
How often should I remeasure?
Run the same prompt matrix on a fixed cadence so you are comparing like-for-like snapshots. Because engine outputs vary by session, region, and date, only repeated, controlled measurement supports any claim that something changed after you acted — a single follow-up read is not proof of cause.
Run an e-commerce AI recommendation audit
Create a workspace, bring the shopper questions your buyers ask, and watch which products get named, which get cited, and which competitors show up in their place across every engine.
Start your recommendation audit
