← All insights
Field notes6 min read

We sent AI agents shopping in a real browser 40 times. Here's where they got stuck.

Four model families, five live stores, two shopping tasks per store. The strongest agents reached checkout every time, but a hidden checkout button, a hard-to-find cart link and missing guest checkout show where stores still lose them.

Most talk about shopping agents is about protocols and payment standards. But today, many agents still do something much simpler: they open a browser, look at the page and click, just like a person would. So we asked a plain question. If an agent shops on an ordinary online store today, how far does it get, and where does it get stuck?

On October 7, 2026, we sent AI agents shopping on five live stores, 40 times in total. In this article, we share what we saw and what store owners can change.

How we ran the test

  • Five stores, five setups. A Shopify store, an ikas store, a Ticimax store, a WooCommerce store and a large custom-built store. We don't name them; the point is the patterns, not the stores.
  • Two shopping tasks per store. For example: "Find X under a given price, add it to the cart and go to checkout."
  • Four model families. Each task was run by four different large language models. 5 stores × 2 tasks × 4 models = 40 runs.
  • A real browser. Each agent drove a real Chromium browser. At every step it saw a numbered list of the page's interactive elements plus a screenshot. It could click, type into the site's search box, choose options, scroll, go back or open a URL on the same site.

The safety rules were strict. Agents never typed personal data, never logged in, never accepted cookie consent, never left the store's domain and never pressed the final order or payment button. A run ended as soon as the agent reached the checkout step: the address or payment form, or the "log in or continue as guest" screen.

The results at a glance

Of the ten tasks, nine could be completed. On those nine:

  • The two strongest models reached checkout 9 out of 9 times.
  • Two lighter, faster models reached checkout 7 out of 9 times.
  • Median run time was between 57 and 86 seconds depending on the model, with an average of 8 to 13 steps.
  • Cost per run ranged from under one US cent on the lightest model to about 14 cents.

In other words, shopping by browser is no longer a lab trick. A capable agent can find a product, choose a variant and reach checkout on an ordinary store in about a minute. The gaps between models appeared at exactly the points where the store made the path harder.

The tenth task: knowing when to stop

One task was impossible: the store doesn't sell that product category at all. The ideal behavior was to say so quickly. One model concluded "this isn't sold here" in 10 steps. The others spent 24 to 25 steps searching before they gave up.

An agent can only stop early if the store gives it a clear signal. Clear category navigation and a site search that honestly says "no results" helped agents reach the right conclusion. For a store, that's not a loss: an agent that leaves quickly is better than one that wanders, then reports something wrong to the user.

Where agents got stuck

1. A checkout button the agent couldn't see

On one store, the "Checkout" button in the cart drawer sat inside a third-party widget built with a closed shadow DOM. The button was visible on screen, but agents that read the page's structure (the DOM and the accessibility tree) couldn't see it, so it never appeared among the elements they could act on. Agents reached checkout only by guessing the /checkout address.

Lesson: Keep a plain, standard link or button to checkout in the main page. A widget can look great to a person and still be invisible to an agent.

On another store, the link to the cart was hard to discover. Several agents tried to guess the cart address (for example /sepet or /cart) and landed on 404 pages. One of the lighter models ended its run on a browser error page.

Lesson: Offer a standard, crawlable cart link, and use predictable cart and checkout URLs. When an agent has to guess, some guesses will be wrong.

3. Log in or continue as guest

Many Turkish stores show a "log in or continue as guest" screen before checkout. Agents reached this screen reliably; that's where our runs ended. But if the store had no guest option, an agent acting for a user would have had to stop there, because it can't create an account or log in on the user's behalf.

Lesson: Offer guest checkout, and make it as easy to spot as the login button.

4. Clear variant choices made things fast

Not every finding was a problem. On stores whose product pages offered size and color as real controls (a select box, radio buttons or plain buttons), agents got through the task in just 3 to 6 steps. When options are real form controls, an agent can recognize them and choose correctly on the first try.

What store owners can do

  1. Keep a plain checkout link in the main page. Even if a third-party cart widget shows its own button, put a standard link or button to checkout in the page itself, where an agent reading the page can find it.
  2. Make cart and checkout URLs predictable. Use a cart link in the header that crawlers can follow, and make sure common addresses like /cart and /checkout lead somewhere sensible instead of a 404.
  3. Offer guest checkout. Agents can't create accounts or log in for the user. A required login is where an agent's journey ends.
  4. Show variants as real controls. Offer size, color and other options as a select box, radio buttons or plain buttons, not as images or custom elements that only look clickable.
  5. Help agents say "no." Make your category navigation clear, and let site search return an honest "no results" when you don't carry something.

The bottom line

Today's agents can shop on an ordinary online store, and the strongest ones do it reliably. What separates a smooth run from a failed one is rarely the agent's intelligence. It's usually one hidden button, one unpredictable address or one missing guest option. These are small, fixable details, and most of them also help human shoppers.

Specoria tests stores with real buyer tasks, shows where an agent gets stuck on the way to checkout, and re-runs the same tasks after each fix. To see how your store looks to a shopping agent today, request a free readiness report.


About this test

  • Measurement date: October 7, 2026. 5 stores, 2 tasks per store, 4 model families, 40 runs in total.
  • Stores and model vendors are not named. This is not a ranking of e-commerce platforms or of AI models.
  • 40 runs is a small sample. Read the numbers above as field observations, not as success rates you can generalize to every store.