I built an AI team around the way I work.
11 specialist roles. Three workflows. A 400-line written standard.
The models were the least interesting part. The useful work was turning my own standards into a system AI could follow.
The same tools give everyone the same starting point
Most people look for more value from AI by finding a better model or saving another prompt. That helps until everyone else can do the same thing.
Your edge starts when AI knows your work, your source material, your standards and your definition of good. That cannot be downloaded from a prompt pack. You have to build it through your own experiments.
This website became one of mine.
From representing me to working with me
In 2023, I built ChatGPSteve to test whether AI could represent my career. It answered questions using 12 years of work history, in my voice, with sources.
Later, I changed the question: could AI work with me on an ongoing job?
Maintaining a website involves research, writing, search checks, editing, prioritisation and decisions that need to survive between sessions. One general chatbot could produce an answer. I needed a system that could maintain the thread.
So I divided the work into specialist roles. Each role owns one question, one set of standards and one place in the sequence.
What I built
The system has four numbers behind it:
The core roles include:
This is what makes it a working system. Each role has a bounded job. Each output has somewhere to go.
The order is part of the system
The reviews run in a fixed sequence.
SEO comes first to establish intent and structure. GEO follows so the finished structure works for AI search as well as traditional search. The content review then puts the page into my voice without breaking that structure.
Consultancy pages get one final clarity check for the firm buyer. People-level pages skip it. The reporter consolidates the findings after the revisions are complete.
Each stage works on the output of the previous one. If I optimise for search after the voice pass, I risk breaking the writing. If I polish for clarity before the voice pass, the next rewrite can undo it.
The sequence is a quality control. It stops one specialist fixing its own problem by creating another problem further down the line.
Nothing publishes automatically. The team proposes changes. I decide what ships.
The part I got wrong
I built the system before I measured the starting point.
I can tell you it has 11 roles, three workflows and a 400-line guide. Those are build numbers. They are not result numbers.
I did not record a clean manual baseline before building it. I cannot tell you honestly that the website team saves a specific number of hours or catches a specific percentage more issues. Retrofitting a number would make the case study look stronger and the experiment weaker.
The next test is straightforward. I will take two comparable pages, give both routes the same brief, and record:
That will tell me whether the system is faster, better, or simply more elaborate.
What earned a place
Five parts of the system have survived repeated use.
Give each role one question
A role becomes useful when its responsibility is narrow enough to judge. “Make this page better” is vague. “Check whether every claim has evidence” can pass or fail.
Write down what good means
The 400-line writing guide matters more than any individual prompt. It gives every role the same standards and stops the system guessing what “sounds like Steve” means.
Design the order
The output of one stage becomes the input to the next. A good role in the wrong place can still damage the result.
Remember decisions, not just outputs
The session manager records what changed and why. That stops each working session becoming another restart.
Keep accountability clear
AI can produce more options and run more checks. It does not inherit responsibility for the result. I still decide what is accurate, what is useful and what ships.
Run the experiment on one part of your work
Start smaller than I did. Choose one repeated task where the output is easy to judge.
Add another role only after the first one produces a measured gain and the next part of the work has a genuinely different way to fail.
That is how a personal AI system grows without turning into a collection of impressive-looking machinery.
The role structure, sequence, written standards and session memory stay. They have become part of how I run the site.
The performance claim remains pending until the matched test produces a number.
That is the point of testing AI on your own work. A convincing system is still a hypothesis until you measure what changed.
The next experiment goes in The Experiment Log
Every fortnight I share one AI experiment on real work: the setup, what changed, the number and the honest verdict.
You also get one experiment to run on your own work before the next issue.
Fortnightly. Keep only what works for you.
Bringing the same discipline into a firm?
I help firms turn scattered AI use into systems that work on repeat.