Free run
NO EMAIL- One scenario × two agents
- Live view with screenshots and the agents’ thoughts
- Stage scores, stuck points and fixes
- A 33-item checklist of what tires the agent, with evidence
- A link you can share
Two AI agents shop your store live in a real browser: they search, open a product, choose a size or colour, add it to the cart and stop at checkout. You watch every step and get a report on where they got stuck and how to fix it.
Like a mystery shopper, our agents walk into your store the way a customer’s AI assistant would and report what got in their way. Unlike a person, they show you every step, and they stop before paying.
A real run of our free simulator on a live fashion store, with its name removed. First what the agent needs to finish the task, then the run step by step, then what the report told the store.
Real run · October 8, 2026 · one agent · store name and product names removed
Task: Find a named polo T-shirt in purple, size M, add it to the cart and go to checkout.
| Requirement | Result | Evidence from the run |
|---|---|---|
| The store opens for the agent’s browser | Met | HTTP 200 on the home page. |
| Site search works | Met | Step 1: typed into the search box. Step 2: a results page. |
| The product page carries price and stock data | Met | Product JSON-LD on every product page opened: 24.90 EUR, in stock. |
| Size and colour can be chosen | Partly | Size M was a labelled button (step 9). The colour swatches had no name, so the agent clicked them by screen position (steps 5 and 17). |
| The agent sees the store your customer sees | Not met | The browser asked for Turkish. The store served its German version with prices in EUR until the agent switched the language itself (step 16). |
| “Add to cart” appears after choosing a size | Not met | Not found in 22 steps: the agent scrolled, reopened the page and changed the size (steps 10-13, 21, 22). |
| Guest checkout | Not reached | The run never got this far. |
The agent’s own note at each step, translated from Turkish, with product names removed. The badge shows the language of the page the store served.
Stopped on the product page after 22 steps. It never reached the cart.
BeforeRECORDED
After a fixEXPECTED · NOT RECORDED
The same task wasn’t run again after a fix, so this side is what we would expect, not a measurement. On your own store, run the simulator before and after a fix to see both sides for real.
Serve the language the browser asks for instead of deciding by IP address alone, or show a visible country and language choice that doesn’t hide add-to-cart.
Give every swatch a text name (visible, or an aria-label) so an agent can pick a colour without guessing a screen position.
Is this a quirk of our test? No. Shopping agents’ browsers also run in data centres, which may be in another country. If a store decides by location alone, an agent can see a version of it your customer never would.
Source: Specoria agent simulator, recorded run · · 1 run, 22 steps
This is what you see live in your report: every page the agent opens, what it was thinking, where it hit friction and where it stopped.
Concept demo: the store, figures and steps are illustrative, not a real customer result.
The agents behave like a shopper using an AI assistant. The rules below are enforced by our harness on every step, not left to the model.
Search box, categories, product pages, size and colour options, the cart: the same interface your customers use.
Reaching the address or payment form, or the “log in / continue as guest” screen, counts as success. They never press the order or payment button.
They never type a name, email, phone, address or card, never log in or create an account, never accept cookie banners and never leave your domain.
Our browser runs on Cloudflare and is identified as Cloudflare’s signed (verified) bot; your store sees that, not ChatGPT. We don’t disguise it, solve CAPTCHAs or use residential proxies.
If your bot protection stops the agents, the report says so plainly. Real shopping agents can hit the same wall, so it’s worth knowing.
The agents are Specoria’s own agent loop, each driven by a different AI model (for example Gemini). They are not the ChatGPT or Gemini consumer agents themselves.
We check that the store is reachable, whether the first response is a bot wall, what your robots.txt says to AI agents, and whether product pages carry Product/Offer data. A real product from your sitemap makes the task concrete.
Two agents shop side by side. Every step shows the screenshot the agent saw and what it was thinking.
Stage scores from finding the product to checkout, a friction map (steps, backtracks, retries), stuck points with screenshots and fixes, the flows that work, and confidence based on how far the agents agree.
Getting stuck is normal: shopping agents are still young and stall on many stores. The report separates what is on your side from what is on ours.
Figures published by the organizations named, mostly for the US market. Read them as a leading indicator.
Growth in AI-referred visits to US retail sites in May 2026 versus a year earlier
Source: Adobe Analytics (Digital Commerce 360) · Jun 17, 2026
Higher conversion rate for AI-referred visits than for visits from other sources
Source: Adobe Analytics (Digital Commerce 360) · Jun 17, 2026
Share of US consumers who used at least one AI tool to shop last month; the first time above 50%
Source: NIQ Agentic Commerce Tracker (The Next Web) · Sep 24, 2026
The figures are published by the organizations named, not by Specoria.
No. They stop at checkout and never press the order or payment button. A product may sit in a cart session, the same as when a visitor leaves without buying.
No. Our browser runs on Cloudflare and is identified as Cloudflare’s signed bot. We don’t imitate any assistant, solve CAPTCHAs or use residential proxies.
No, that is a finding. Real shopping agents can be stopped the same way. If you want verified agents in, allow them in your WAF or bot settings and run the test again.
Different models find different paths. If every agent gets stuck at the same stage, we report a repeated block and show the technical evidence: each step and screenshot up to where they stopped. If only some do, it’s friction. One run is not a final verdict, so fix the cause and run it again.
Yes. The pre-check recognises from your home page whether it is an e-commerce, accommodation or services site. On a hotel site the agent chooses dates, the number of guests and a room and goes as far as the guest details form; on a services site it finds the service and reaches the contact or quote form, reading which details the form asks for. In both cases nothing is typed into the form and no form is sent. If your booking opens in an engine on another domain, the agent doesn’t follow it there; the report marks that part as “not measured”.
Your own store’s domain. Marketplaces and a short list of large sites whose terms forbid automated agents are excluded. Free runs are limited to 3 a day per visitor.
We keep the run and its screenshots for 30 days so the shared link works, then delete them unless you ask for the full report. We store a salted hash of your IP address for the daily limit, never the address itself.
Starter re-measures every month with real agent simulations and turns each finding into a task with ready-made fixes. First month $1, then $149/month.