See whether AI agents can complete your customer journeys.
An AI agent is an assistant that can browse pages, compare options, use tools and work through forms on a customer’s behalf.
The AI Mystery Shopper is a controlled AI agent, sent through five important journeys on your website. You see what it completed, where it stalled, what it misunderstood and what to fix first. Every material result is checked by a named human reviewer before the report is issued.
Five real journeys Repeat runs Captured evidence Human-reviewed verdict
Scope, fee and timing are agreed before testing starts.
Your information can be correct and still be unusable by an AI assistant.
A customer may ask an AI assistant to find the right service, compare two options, identify a relevant expert or reach the next step. The answer may exist on your website, but that does not mean an agent can find it, interpret it correctly or act on it safely.
An agent can reach a form and still miss a required field. It can find a client story and choose the wrong one. It can finish a journey and return the wrong price, recommendation or next step.
It missed a required field.
The label made sense to a person. The agent guessed, and the guess was wrong.
It chose the wrong one.
Five case studies, one relevant. It quoted the one from a different sector.
It returned the wrong price.
Every step completed. The number at the end was not the one you publish.
A normal website scan does not show that. The journey has to be attempted.
Start with the journeys where failure would matter.
We agree five representative tasks before testing begins. Each task has a starting point, a correct outcome, clear pass conditions and a safe stopping point.
Find the right option
A search result or a landing page, with the customer’s need stated in plain language.
It names the option a knowledgeable person would have chosen.
The named option matches the need, not just the search words.
Nothing is submitted and no account is created.
Compare the choices
Two candidate options the customer is weighing up.
It states the real difference in price, terms or eligibility.
Every figure it reports matches your published source.
No quote is requested and no data is sent.
Find the right person
A stated need that requires a named specialism or sector experience.
It returns a person or team who genuinely has that experience.
The evidence for the match is findable on your own website.
No message, call request or meeting is booked.
Find convincing proof
An objection or a due-diligence question a buyer would actually raise.
It retrieves the case study, evidence or policy that answers the question.
The proof is relevant, current and quoted accurately.
No gated asset is requested with real contact details.
Reach the next step
The intent to buy, book, apply or enquire.
It reaches the final confirmation step with every input valid.
All required fields understood, and nothing is submitted.
It halts on the confirmation screen, before submission.
The aim is not to crawl every page. It is to test a small number of customer jobs with real commercial, operational or control consequences.
The agent attempts the journey. The evidence shows what happened.
Pick
We define the five customer jobs, the correct outcomes and what would count as a serious failure.
Test
A pinned reference agent attempts every journey twice. A third run is used only when the results disagree.
Review
I check the evidence, reject anything unreliable and decide whether each journey passed, partly worked or failed.
Prioritise
The 22 supporting measures explain why the journeys behaved as they did and what to fix first.
Retest
The same frozen journeys can be run again after changes so you can see whether performance moved.
Pass, partial or fail. No technical translation required.
Pass
The agent reached the agreed outcome, returned the right information and respected the stopping rule.
Partial
The agent made useful progress but missed part of the outcome, needed an avoidable workaround or could not complete every safe step.
Fail
The agent could not reach the outcome, returned the wrong answer or hit a critical failure that made completion unsafe or misleading.
One line per agreed task.
A completed flow is not automatically a pass. If the recommendation, price or next step is wrong, the report says so.
A plain-English baseline your team can act on and retest.
An executive verdict.
Which journeys work, which do not and where the greatest risk sits.
Two journeys work. Two are unreliable. One is unsafe to trust.
2
Pass
2
Partial
1
Fail
The quote journey completes and returns a figure that does not match the published price.
A journey scorecard.
Pass, partial or fail for every agreed task.
A plain-English findings report.
What was tested, what happened, why it matters and what good looks like.
What happened, in words your team can use.
A prioritised action plan.
What to fix first, what to test next and what can wait.
A technical evidence pack.
The detailed measurements and captured evidence for the people implementing the changes.
A documented human review.
Every material amendment, limitation and final judgement recorded under my name.
The original result, the decision and the evidence all remain visible.
A frozen retest baseline.
The tasks and success criteria needed to measure improvement after changes.
The five tasks, the correct outcomes and the pass conditions, held still so the next run is comparable.
The client report explains the decision. The evidence pack shows how it was reached.
The machine records the run. A human owns the verdict.
Automation provides repeatability. It does not get the final say.
I check whether the evidence is complete, whether the agent reached the correct outcome and whether a serious mistake is hidden inside an apparently completed journey. If I change the proposed result, the original result, the decision and the evidence remain visible.
The report is not released until the scope, evidence, journey decisions, critical failures, limitations and final verdict have been signed off.
No report leaves without my name on the verdict.
Steve Quinlan · a decade of testing, kept only what paid back.
Your first useful benchmark is your own baseline.
The same frozen journeys can be repeated under the same conditions and rerun after changes. That answers the question that matters first: did the journey improve?
As more reviews are completed, comparable journeys can reveal recurring failure patterns across website types and customer jobs. Those comparisons only count when the task, correct outcome, agent version, environment, evidence and human-review process are genuinely comparable.
I will not quote an industry average or percentile until there are enough comparable, permissioned real reviews to defend it. Synthetic test runs never become client-facing benchmark evidence.
Your own baseline, rerun on the same frozen journeys.
Industry averages and percentiles. Not until they can be defended.
Built for organisations with customer journeys that matter.
- Your website includes an important form, quote, booking, application, comparison or service-selection journey.
- A wrong result could mean lost revenue, poor service or a control problem.
- Your team needs evidence before a rebuild, release or automation investment.
- You can define what the correct customer outcome should be.
- You want to retest improvements, not receive a one-off list of recommendations.
- You only want to know whether AI tools can find or cite your website.
- You need accessibility certification, penetration testing or legal assurance.
- You want to test your own chatbot rather than an independent customer-side agent.
- Nobody can name the customer job or the correct result.
- You want a badge or a guarantee that every future AI agent will behave the same way.
The questions that come up first.
01What is an AI agent?
An AI agent is an assistant that can browse pages, compare information, use tools and take steps towards a goal. The AI Mystery Shopper tests a customer-side agent attempting to use your website, not your organisation’s own chatbot.
02What exactly do you test?
We agree five important customer journeys and a correct outcome for each one. The agent attempts every journey through repeat runs. The supporting review examines the website structure, forms, performance, content and error handling that influenced the outcome.
03Is this the same as usability testing?
No. Human usability testing studies how people experience a journey. The AI Mystery Shopper tests how an AI assistant operates and interprets it. The two can reveal different problems and work well together.
04Is this an AI visibility or GEO audit?
No. Visibility asks whether AI can find, understand and cite the organisation. The AI Mystery Shopper asks whether an agent can complete a customer task correctly. Visibility may explain part of a failure, but it is not the outcome being sold.
05Is the final verdict automated?
No. Automation carries out repeatable checks and organises the evidence. I review correctness, evidence gaps, critical failures, amendments, limitations and the final verdict before the report is issued.
06Which AI agents do you test?
The agent and model version are agreed when the work is scoped and recorded in the report. A baseline normally starts with one pinned reference agent. Wider cross-agent comparison can be added when it would change the decision.
07Do you need access to our systems?
Many public journeys can be tested from the outside. Authenticated or transactional journeys may need a safe test account, sandbox or agreed stopping point. Access requirements are settled before testing.
08What happens after the review?
Your team or delivery partner can work from the prioritised action plan. The same journeys can then be rerun to show whether the changes improved the result.
09How long does it take and what does it cost?
The scope depends on the number and complexity of the journeys, access requirements and depth of evidence needed. You receive the agreed timing, fee, deliverables and boundaries before testing begins.
Choose one customer job you cannot afford an AI agent to get wrong.
Tell me the URL, what the customer is trying to achieve and what the correct outcome should look like. I will tell you whether The AI Mystery Shopper is the right way to test it and where the scope should stop.
Direct reply from me within two business days. No sales sequence.