Real AI experiments on real work.
I test AI on the work that fills your week: decisions, documents, research, reviews and repeated tasks.
Each experiment compares an AI approach with how the work gets done now. I measure the difference, show the evidence and publish the honest verdict.
-
Real taskwork worth doing better.
-
Fair testone change against a credible baseline.
-
Measured resulttime, quality or output.
-
Honest verdictkeep it or bin it.
A 733-word prompt scored 52.5. An 87-word brief scored 100.
I gave Claude Opus 5 the same evidence and the same decision six times.
Three runs used a 733-word prompt with 20 prescribed steps. Three used an 87-word goal-led brief. The longer prompt found the right problem, then followed its instructions into the wrong answer.
The prompt told the model which analysis to run. That analysis excluded the problem in the evidence. All three runs recommended the wrong fix.
The brief defined the outcome, evidence, constraints and checks. It left the route through the work open. All three runs reached the supported decision.
define the destination, evidence, constraints and checks.
prescribe the route before the model has seen what the evidence requires.
One task. One model. One day. Six runs. This does not prove short beats long. It shows what can happen when the route in a prompt excludes the problem hiding in the evidence.
A demo shows what AI can do. An experiment shows whether it helps.
A polished output proves very little. The useful question is whether AI made a real task better than the way you do it now.
Every published experiment needs four things.
The task
A recognisable piece of work with a reason to improve it.
The test
One AI approach compared with a credible way of doing the work today.
The result
A difference in time, quality or output, with enough evidence to check the claim.
The verdict
Keep what worked. Bin what did not. State the limits either way.
Start with the problem, not the tool.
The model will change. The work will still need to be done.
These experiments, tools and build notes are organised around what they teach you, not which AI product happened to be used.
The Outcome Brief
Define the outcome, reader, source, constraints and quality checks. Leave the route through the work open, then test what comes back.
Build your briefI built an AI team around the way I work
What earned a place, what I got wrong and why the order of the work mattered more than the number of agents.
See what earned a placeWhat does a CV become when you can talk to it?
ChatGPSteve tested whether a source-grounded conversation could replace the usual hunt through roles, dates and project descriptions.
See the ChatGPSteve experimentWhy your first AI project should save hours
Choose work small enough to test, important enough to matter and measurable enough to earn the next experiment.
Read the eight-week approachWhy CRO people are the best AI adopters
The advantage does not come from knowing more tools. It comes from knowing how to test a claim before turning it into a workflow.
See the experimentation habitsAI is changing buyer research before the first click
Buyers increasingly form a view before reaching a website. This turns the change into an experiment you can run rather than a trend you have to believe.
Run the buyer-research testDo not copy my result. Run the test on your work.
My experiment can show you where to look. It cannot tell you what will work inside your role, your team or your constraints.
That is why the method starts with your task.
Five steps. One proven gain at a time.
Pick a real task. Test an AI approach. Measure the difference. Keep what works. Stack the win into the way you work.
Brief the result. Leave the route open.
Turn the outcome, reader, source, constraints and checks into a brief you can test. The builder runs in your browser.
Get the next experiment, not another list of tools.
Every fortnight I send one real AI experiment: the task, the test, the number and the honest verdict.
You also get one experiment you can run on your own work.
Free. Fortnightly. Leave whenever you like.
The archive is bigger than the hub.
This page is curated around the work that best represents the brand now.
The blog keeps the complete record, including earlier writing on experimentation, product, UX, AI consulting and GEO.
Browse the full writing archiveBring the same discipline into your business.
I help firms turn the AI they have already bought into systems that pay back. We find the biggest opportunity, prove the value and build from there.
Explore AI consulting