Our crawler: how Specoria visits your site

Specoria measures how ready sites are for AI shopping agents. This page explains openly which requests we make to your site, how we identify ourselves and how to opt out.

When do we visit?

Only when someone starts a free test or an agent simulation for an address on specoria.com, and when we measure a project in our panel. We don’t run a general background crawl of the web.

Which requests does the free test make?

  • Public pages: robots.txt, the home page, sitemap, llms.txt, llms-full.txt, AGENTS.md, /.well-known/ucp, a few product and policy pages, and one random address that doesn’t exist, to see whether it returns a real 404. One test is limited to a 20-second time budget.
  • Agent protocol discovery files (only if the home page opened; each small and short-lived): /.well-known/mcp.json, /.well-known/mcp, /.well-known/agent-card.json, /.well-known/ai-catalog.json, /.well-known/api-catalog, /.well-known/http-message-signatures-directory, /checkout_sessions. All are GET requests; we look at /checkout_sessions with GET only, never POST, and never open a session.
  • Policy links (terms, privacy, returns, shipping, contact, about, FAQ, support): one GET request for each one found on the home page that the test hasn’t already opened.
  • Your store platform’s public, read-only product and search endpoints: /products/<handle>.js and /search/suggest.json on Shopify, the Store API product search on WooCommerce. We never send requests to cart endpoints.
  • Up to 3 on-site links in llms.txt are checked with small (capped at 4 KB) GET requests to see whether they open.
  • We read robots.txt rule by rule for each AI agent’s name; we never send requests to your site under those agents’ identities.
  • If the site reports a rate limit (HTTP 429) we honour Retry-After and retry at most once; if the requested wait is longer than the test allows we don’t retry in that test. We don’t count a rate limit as a block.
  • These requests carry ordinary browser headers, because we measure how a shopping agent sees your page.
  • One request identifies itself openly as a bot: “Mozilla/5.0 (compatible; SpecoriaBot/1.0; AI agent readiness check; +https://specoria.com/bot/)”. It shows how your bot protection treats self-identified bots (AI agents included). We never impersonate another company’s bot.
  • If the same address is tested again within 6 hours we don’t visit your site; we show the stored result.
  • For paying projects, a light daily sentinel check runs once a day: robots.txt, the home page and the product page from the last scan (3 GET requests in total). It only exists to alert that project’s owner about changes.
  • We don’t submit forms, log in or add products to a cart.

What does the agent simulation do?

  • In a real browser (Cloudflare Browser Run) an AI model browses your site like a customer: it searches, opens products, chooses options and adds to cart.
  • It stops at checkout. It never places an order, pays, or types personal or card details; on hotel and service sites it types nothing into guest or contact forms and never submits them.
  • It stays on the domain it was given and doesn’t follow links to other domains.
  • On the public simulator, one IP address can start at most 3 runs per day.

Catalog reading in the panel

For customers who connect their site to our panel, we read the product catalog from public sources: /products.json and /meta.json on Shopify, the product feed the store publishes (Merchant XML or CSV), or up to 50 product pages from the sitemap. To compare price, stock and GTIN between the feed and the page, we open up to 8 product pages linked from the catalog, one at a time with a pause. If the customer has turned on a stable URL for their product feed, the same read repeats by itself at most once a day (only for Shopify product data and feed URL sources; the same requests, the same limits). We only read; we never write anything to the store or send the product feed anywhere.

AI visibility measurement

Visibility measurement doesn’t visit your site: we ask ChatGPT, Claude, Gemini and Perplexity questions through their official APIs and record whether the brand appears in the answers.

Opting out

If you don’t want your site tested or simulated, email your domain to contact@specoria.com; we add it to our exclusion list and no new test or simulation can be started for it. Government, education and military domains are never tried.

Security and contact

If you saw a problem with a request, write to contact@specoria.com with the time, address and identity; we’ll look into it and reply.