The AI Mystery Shopper · agent-tested journeys

See whether AI agents can complete your customer journeys.

An AI agent is an assistant that can browse pages, compare options, use tools and work through forms on a customer’s behalf.

The AI Mystery Shopper is a controlled AI agent, sent through five important journeys on your website. You see what it completed, where it stalled, what it misunderstood and what to fix first. Every material result is checked by a named human reviewer before the report is issued.

Five real journeys Repeat runs Captured evidence Human-reviewed verdict

Scope, fee and timing are agreed before testing starts.

A human owns the verdict
An illustrative run of a quote journey. The agent opened the business cover page, read the three levels of cover and started a quote for twelve employees. One question was not clear, so it guessed. It then read back a price of £0.00, when the published price for twelve employees is £1,284. It stopped before anything was sent. Verdict: fail. A person checked the run before this was called.
The gap most teams cannot see

Your information can be correct and still be unusable by an AI assistant.

A customer may ask an AI assistant to find the right service, compare two options, identify a relevant expert or reach the next step. The answer may exist on your website, but that does not mean an agent can find it, interpret it correctly or act on it safely.

An agent can reach a form and still miss a required field. It can find a client story and choose the wrong one. It can finish a journey and return the wrong price, recommendation or next step.

It missed a required field.

The label made sense to a person. The agent guessed, and the guess was wrong.

It chose the wrong one.

Five case studies, one relevant. It quoted the one from a different sector.

It returned the wrong price.

Every step completed. The number at the end was not the one you publish.

A normal website scan does not show that. The journey has to be attempted.

Five customer jobs

Start with the journeys where failure would matter.

We agree five representative tasks before testing begins. Each task has a starting point, a correct outcome, clear pass conditions and a safe stopping point.

A review may test whether an agent can

Task definition Agreed before testing

Find the right option

Starting point

A search result or a landing page, with the customer’s need stated in plain language.

Correct outcome

It names the option a knowledgeable person would have chosen.

Pass condition

The named option matches the need, not just the search words.

Safe stopping point

Nothing is submitted and no account is created.

Task definition Agreed before testing

Compare the choices

Starting point

Two candidate options the customer is weighing up.

Correct outcome

It states the real difference in price, terms or eligibility.

Pass condition

Every figure it reports matches your published source.

Safe stopping point

No quote is requested and no data is sent.

Task definition Agreed before testing

Find the right person

Starting point

A stated need that requires a named specialism or sector experience.

Correct outcome

It returns a person or team who genuinely has that experience.

Pass condition

The evidence for the match is findable on your own website.

Safe stopping point

No message, call request or meeting is booked.

Task definition Agreed before testing

Find convincing proof

Starting point

An objection or a due-diligence question a buyer would actually raise.

Correct outcome

It retrieves the case study, evidence or policy that answers the question.

Pass condition

The proof is relevant, current and quoted accurately.

Safe stopping point

No gated asset is requested with real contact details.

Task definition Agreed before testing

Reach the next step

Starting point

The intent to buy, book, apply or enquire.

Correct outcome

It reaches the final confirmation step with every input valid.

Pass condition

All required fields understood, and nothing is submitted.

Safe stopping point

It halts on the confirmation screen, before submission.

The aim is not to crawl every page. It is to test a small number of customer jobs with real commercial, operational or control consequences.

Tested, not assumed

The agent attempts the journey. The evidence shows what happened.

Pick

We define the five customer jobs, the correct outcomes and what would count as a serious failure.

Test

A pinned reference agent attempts every journey twice. A third run is used only when the results disagree.

Review

I check the evidence, reject anything unreliable and decide whether each journey passed, partly worked or failed.

Prioritise

The 22 supporting measures explain why the journeys behaved as they did and what to fix first.

Retest

The same frozen journeys can be run again after changes so you can see whether performance moved.

The same journeys, rerun after changes
One verdict per journey

Pass, partial or fail. No technical translation required.

Pass

The agent reached the agreed outcome, returned the right information and respected the stopping rule.

Partial

The agent made useful progress but missed part of the outcome, needed an avoidable workaround or could not complete every safe step.

Fail

The agent could not reach the outcome, returned the wrong answer or hit a critical failure that made completion unsafe or misleading.

The journey scorecard

One line per agreed task.

01 Find the right option 2 / 2 6 captures PASS
02 Compare the choices 3 / 2 11 captures PARTIAL
03 Find the right person 2 / 2 5 captures PASS
04 Find convincing proof 2 / 2 8 captures PARTIAL
05 Reach the next step 2 / 2 14 captures FAIL

Illustrative layout. Every verdict is set for your journeys after human review.

A completed flow is not automatically a pass. If the recommendation, price or next step is wrong, the report says so.

The output

A plain-English baseline your team can act on and retest.

For the decision-makers

An executive verdict.

Which journeys work, which do not and where the greatest risk sits.

01 · Executive verdictPage 1

Two journeys work. Two are unreliable. One is unsafe to trust.

2

Pass

2

Partial

1

Fail

Greatest risk

The quote journey completes and returns a figure that does not match the published price.

For the decision-makers

A journey scorecard.

Pass, partial or fail for every agreed task.

02 · Journey scorecard5 tasks
01Find the right optionPASS
02Compare the choicesPARTIAL
03Find the right personPASS
04Find convincing proofPARTIAL
05Reach the next stepFAIL
For the decision-makers

A plain-English findings report.

What was tested, what happened, why it matters and what good looks like.

03 · Findings reportJourney 05

What happened, in words your team can use.

For the decision-makers

A prioritised action plan.

What to fix first, what to test next and what can wait.

04 · Prioritised action planFix first can wait
P1Label the required fields the agent could not interpret in the quote form.
P2Make the published price reachable without a calculation step.
P3Tag case studies by sector so the right proof is chosen.

Each item names the journey it unblocks.

For the implementers

A technical evidence pack.

The detailed measurements and captured evidence for the people implementing the changes.

05 · Technical evidence pack22 measures
step 03 · required field “employees”, no label association
step 04 · returned £0.00 · published £1,284
step 04 · source: /quotes/summary · capture 09
run 1 / run 2 agreement: disagree run 3
14 captures Run logs Agent + version Measurements
Signed under my name

A documented human review.

Every material amendment, limitation and final judgement recorded under my name.

06 · Documented human reviewAudit trail
Machine proposedPARTIAL
Amended: wrong price returned
Human verdictFAIL

The original result, the decision and the evidence all remain visible.

Reviewed and signed · S. Quinlan
So you can measure again

A frozen retest baseline.

The tasks and success criteria needed to measure improvement after changes.

07 · Frozen retest baselineLocked

The five tasks, the correct outcomes and the pass conditions, held still so the next run is comparable.

01Task, outcome and pass conditionFROZEN
02Agent and model versionFROZEN
03Environment and stopping rulesFROZEN
Rerun after changes

The client report explains the decision. The evidence pack shows how it was reached.

Not an automated score

The machine records the run. A human owns the verdict.

Automation provides repeatability. It does not get the final say.

I check whether the evidence is complete, whether the agent reached the correct outcome and whether a serious mistake is hidden inside an apparently completed journey. If I change the proposed result, the original result, the decision and the evidence remain visible.

The report is not released until the scope, evidence, journey decisions, critical failures, limitations and final verdict have been signed off.

Scope Evidence Journey decisions Critical failures Limitations Final verdict
Steve Quinlan
Sign-off

No report leaves without my name on the verdict.

Steve Quinlan · a decade of testing, kept only what paid back.

Experimenter, not guru
A benchmark that earns the name

Your first useful benchmark is your own baseline.

The same frozen journeys can be repeated under the same conditions and rerun after changes. That answers the question that matters first: did the journey improve?

As more reviews are completed, comparable journeys can reveal recurring failure patterns across website types and customer jobs. Those comparisons only count when the task, correct outcome, agent version, environment, evidence and human-review process are genuinely comparable.

I will not quote an industry average or percentile until there are enough comparable, permissioned real reviews to defend it. Synthetic test runs never become client-facing benchmark evidence.

Baseline vs retest
Find the right option80% 95%
Compare the choices55% 90%
Find the right person70% 85%
Find convincing proof40% 80%
Reach the next step25% 75%
First run Retest Illustrative movement

Real movement is measured against your own frozen baseline.

Available now

Your own baseline, rerun on the same frozen journeys.

Not yet

Industry averages and percentiles. Not until they can be defended.

Where this is useful

Built for organisations with customer journeys that matter.

Usually a good fit when
  • Your website includes an important form, quote, booking, application, comparison or service-selection journey.
  • A wrong result could mean lost revenue, poor service or a control problem.
  • Your team needs evidence before a rebuild, release or automation investment.
  • You can define what the correct customer outcome should be.
  • You want to retest improvements, not receive a one-off list of recommendations.
Not the right fit when
  • You only want to know whether AI tools can find or cite your website.
  • You need accessibility certification, penetration testing or legal assurance.
  • You want to test your own chatbot rather than an independent customer-side agent.
  • Nobody can name the customer job or the correct result.
  • You want a badge or a guarantee that every future AI agent will behave the same way.
Common questions

The questions that come up first.

01What is an AI agent?

An AI agent is an assistant that can browse pages, compare information, use tools and take steps towards a goal. The AI Mystery Shopper tests a customer-side agent attempting to use your website, not your organisation’s own chatbot.

02What exactly do you test?

We agree five important customer journeys and a correct outcome for each one. The agent attempts every journey through repeat runs. The supporting review examines the website structure, forms, performance, content and error handling that influenced the outcome.

03Is this the same as usability testing?

No. Human usability testing studies how people experience a journey. The AI Mystery Shopper tests how an AI assistant operates and interprets it. The two can reveal different problems and work well together.

04Is this an AI visibility or GEO audit?

No. Visibility asks whether AI can find, understand and cite the organisation. The AI Mystery Shopper asks whether an agent can complete a customer task correctly. Visibility may explain part of a failure, but it is not the outcome being sold.

05Is the final verdict automated?

No. Automation carries out repeatable checks and organises the evidence. I review correctness, evidence gaps, critical failures, amendments, limitations and the final verdict before the report is issued.

06Which AI agents do you test?

The agent and model version are agreed when the work is scoped and recorded in the report. A baseline normally starts with one pinned reference agent. Wider cross-agent comparison can be added when it would change the decision.

07Do you need access to our systems?

Many public journeys can be tested from the outside. Authenticated or transactional journeys may need a safe test account, sandbox or agreed stopping point. Access requirements are settled before testing.

08What happens after the review?

Your team or delivery partner can work from the prioritised action plan. The same journeys can then be rerun to show whether the changes improved the result.

09How long does it take and what does it cost?

The scope depends on the number and complexity of the journeys, access requirements and depth of evidence needed. You receive the agreed timing, fee, deliverables and boundaries before testing begins.

Start with one journey

Choose one customer job you cannot afford an AI agent to get wrong.

A quote A booking An application A comparison A product or service choice

Tell me the URL, what the customer is trying to achieve and what the correct outcome should look like. I will tell you whether The AI Mystery Shopper is the right way to test it and where the scope should stop.

Send me one journey to test

Direct reply from me within two business days. No sales sequence.