Mystery shopping, by AI agents

Can an AI agent actually use your website?

Buyers now send AI agents to research and shop for them. We send leading agents to mystery-shop your site, check whether they finish the tasks that make you money, and show you exactly where they fail.

Agents your buyers already use

ChatGPT Dots Claude Gemini Grok Bot Muse Instinct

Get your free report

We run AI agents on your site and email you what passed, what failed, and why.

If the report is useful, which plan would you pick?
Monthly Pro pricenorthstar.example · journey #2 · 4 Oct 2026
Confirmed incident
Agent
Claude agent · desktop Chrome · US
Baseline
correct 3 of 3 (20 Sep)
Now
incorrect 2 of 2 (scheduled + confirmation)
Outcome
Task failed · counts against the site
  1. 1open northstar.exampleok
  2. 2click "Pricing" in the main navok
  3. 3read Pro plan price · billing tab left on "Annual" · extracted "$24/mo"ok
  4. 4assert price == $29/month · observed $24failed
ObservedThe agent reported the annual-billing equivalent as the monthly price.
Likely contributor · MisreadBoth prices appear without their billing-period label in the accessible tree.
Suggested changeTie the billing-period label to each amount (visible text, not only the tab state).
view replay · compare with last pass · re-test after fix

Evidence attached: extracted answer, pricing-page screenshot, selected billing tab.

Sample report

The problem

Agents are already on your site. Your analytics won't tell you when they fail.

Demand

Buyers delegate the research

"Find me a project tool under $30 a month with SSO" is now a job for Dots, Grok Bot, Muse or Instinct. The agent opens your site in its own browser, reads your pricing, picks a plan, and reports back. If it can't, your competitor gets the recommendation.

Silent failure

Agents fail without telling you

They read the annual price as monthly, pick the wrong variant, stall behind a cookie wall, or click a destructive button. The person hears "I couldn't find it" and moves on. No error, no support ticket, nothing in your dashboards.

Blind spot

Existing tools test the wrong thing

Uptime and scripted synthetic monitors replay a path you wrote. Readiness checkers look for files like robots.txt and llms.txt. Neither tells you whether an agent, finding its own route, completes the task.

Aug to Sep 2026OpenAI's Dots, Meta's Muse, xAI's Grok Bot and Instinct all launched always-on agents that browse the web on their own. OpenAI · Meta
Oct 2026Merj observed a leading agent click a destructive button first in 25 of 25 runs on a real site. The owner never saw it. Research
2026Google published guidance for agent-friendly websites and shipped an agentic-readiness audit in Lighthouse. Chrome blog
Live nowAgent checkout is live at Walmart, Target, Etsy and Shopify merchants, and behind ChatGPT Shopping.

How it works

Give an agent a goal in plain English. We verify whether it got there, and tell you why not.

A journey is one task a real buyer would hand to their agent. Three journeys, two agents, one site:

"Find the monthly price of the Pro plan."
OpenAI agent · passed 3 of 3Claude agent · passed 1 of 3
Home
Pricing
Pro plan
Monthly pricemisread
"Add a medium black T-shirt to the cart."
OpenAI agent · passed 3 of 3Claude agent · passed 3 of 3
Home
Men's tees
Product page
M · Black
Cartreached
"Find the refund policy."
OpenAI agent · passed 0 of 3Claude agent · passed 0 of 3
Home
Footer
Helpnot discoverable: hover-only menu
Refund policy
1

Describe the journeys that make you money

"Find the monthly price of the Pro plan." "Reach the free-trial form." "Add a medium black T-shirt to the cart." Plain English. No scripts, no selectors.

2

Leading AI agents run them on a schedule

We run each task with the agent models behind ChatGPT, Claude and Gemini, in a real browser. Every result shows which agent passed, and how many times.

3

Outcomes are verified, not self-reported

Pass or fail is decided by checks on the page, URL, cart contents or extracted value. The agent saying "done" never counts.

4

Every failure gets a cause and a suggested fix

Hidden behind an overlay. Not discoverable. Misread. Ambiguous action. Broke mid-flow. Site bug. Each comes with the step trace, screenshots and a change your team can make.

5

You hear about regressions, with proof

Slack or email when a journey that used to pass stops passing. Re-test after a deploy and see the fix proven.

Who it's for

Built for the people who own the funnel, not the test suite.

Reports are written for marketing, growth and e-commerce leads. No setup call, no engineer required to read them.

SaaS

Can an agent find your pricing and start a trial?

Buyers research through ChatGPT and Claude before they ever see your homepage. Check that agents read the right price, compare plans correctly and reach the trial form, and re-check on every deploy.

  • find price · compare plans
  • find a doc or policy
  • reach trial signup

E-commerce

Can an agent find the right size and reach your cart?

Agentic commerce is a board topic. Before an agent can buy from you, it has to find the product, pick the variant and reach the cart without getting lost in a mega-menu or blocked at the door.

  • find product · select variant
  • reach cart
  • find shipping and returns

Agencies

Give every client an AI mystery-shopping report this month

One account, many client sites. Pooled tests, client groups and branded reports. Add AI mystery shopping to the maintenance package and keep clients paying for the alerts.

  • all journeys, grouped per client
  • shareable branded reports
  • monthly regression alerts

Pricing

Pay per agent test.

One test is one AI agent attempting one task on your site. Every plan starts with a free report.

Free report
$0
  • One site
  • 2 agent tests
  • Full evidence: trace, screenshots, extracted answers
  • Shareable report

Everyone. Start here.

Get my free report
Starter
$29/month
  • 1 site
  • 30 agent tests a month
  • Weekly scheduled runs
  • Email alerts
  • 90-day history

Indie SaaS and small shops.

Get started
Agency
$149/month
  • 10 sites
  • 300 pooled agent tests a month
  • Client groups and viewer roles
  • Branded, shareable reports
  • Everything in Pro

Web, growth and SEO agencies.

Get started
Scale
from $499/month

100+ sites or daily agent runs on priority journeys, API access, SSO, and competitor benchmarks: the same journeys run on up to three competitor sites, side by side.

Talk to us

Add-ons: one-off audit with a fix plan, $299. Extra competitor benchmark, $49 a month.

Questions

Straight answers

Do you test the real ChatGPT or Claude?

We run the same agent models those assistants are built on, in a real browser. Every result names the agent that produced it.

What is in the free report?

For each task: pass or fail, the step trace with screenshots, the answer the agent extracted, and for failures the likely cause and a suggested change. The sample at the top of this page shows the format.

How is this different from Checkly, Datadog Synthetics or Lighthouse?

Scripted monitors replay clicks you recorded. Readiness audits check that files like robots.txt or llms.txt exist. We give an agent a goal, let it find its own route, and verify the outcome independently.

What if my site blocks bots?

That is a finding in itself. The report shows where the agent was blocked, and you can allowlist our runner to let it through.

Why "2 of 3" instead of a score?

Agents are probabilistic, so one run proves little. Each task runs three times per agent, and we report the count.

Free report

Find out before your customers' agents do.

Tell us the site and the tasks that matter. We'll email you the evidence.

  • Pass or fail for every task
  • Step trace, screenshots, extracted answers
  • Likely cause and a suggested fix for every failure

Get your free report

If the report is useful, which plan would you pick?