Field note · AI systems

I built an AI team around the way I work.

11 specialist roles. Three workflows. A 400-line written standard.

The models were the least interesting part. The useful work was turning my own standards into a system AI could follow.

Author: Steve Quinlan Topic: AI systems · Experimentation Read: 7 min
The AI team Hover a role  ·  tap on mobile
11
Eleven roles, one operator
Each role owns one question, one set of standards and one place in the sequence. Hover or tap a node to read its job.
One person, 11 specialist roles and a defined route through the work.
01The edge

The same tools give everyone the same starting point

Most people look for more value from AI by finding a better model or saving another prompt. That helps until everyone else can do the same thing.

Your edge starts when AI knows your work, your source material, your standards and your definition of good. That cannot be downloaded from a prompt pack. You have to build it through your own experiments.

This website became one of mine.

Same tools, same start
Your edge
Your own experiments
02Working with me

From representing me to working with me

In 2023, I built ChatGPSteve to test whether AI could represent my career. It answered questions using 12 years of work history, in my voice, with sources.

Later, I changed the question: could AI work with me on an ongoing job?

Maintaining a website involves research, writing, search checks, editing, prioritisation and decisions that need to survive between sessions. One general chatbot could produce an answer. I needed a system that could maintain the thread.

So I divided the work into specialist roles. Each role owns one question, one set of standards and one place in the sequence.

The ChatGPSteve chat interface answering a question about Steve’s career, in his voice, with sources.
ChatGPSteve · 2023 · in my voice, with sources Try ChatGPSteve
03The system

What I built

The system has four numbers behind it:

11
11 specialist roles covering planning, writing, review, links and reporting.
3
Three workflows for creating content, reviewing pages and maintaining context.
400
400 lines in the current writing guide that defines voice, evidence standards and what good looks like.
1
One final editor: me. The system recommends changes. I decide what ships.

The core roles include:

01A site planner chooses the page and the depth of review.
02An SEO checker establishes search intent, headings, metadata and internal links.
03A GEO optimiser checks whether AI tools can extract, understand and cite the page.
04A content reviewer applies my voice, evidence standards and writing rules.
05Consultancy pages get a separate plain-English check for the firm buyer.
06A reporter turns the findings into one prioritised plan.
07A session manager records what changed, what I decided and what happens next.
+Other roles handle drafting, link checks and supporting quality controls.

This is what makes it a working system. Each role has a bounded job. Each output has somewhere to go.

04The sequence

The order is part of the system

The reviews run in a fixed sequence.

A fixed route conditional
Each stage works on the output of the previous one. Hover or tap a stage to see its job, and why its place in the line matters.
Planner → SEO → GEO → voice → conditional clarity → reporter → decision.

SEO comes first to establish intent and structure. GEO follows so the finished structure works for AI search as well as traditional search. The content review then puts the page into my voice without breaking that structure.

Consultancy pages get one final clarity check for the firm buyer. People-level pages skip it. The reporter consolidates the findings after the revisions are complete.

Each stage works on the output of the previous one. If I optimise for search after the voice pass, I risk breaking the writing. If I polish for clarity before the voice pass, the next rewrite can undo it.

The sequence is a quality control. It stops one specialist fixing its own problem by creating another problem further down the line.

Nothing publishes automatically. The team proposes changes. I decide what ships.

05The honest part

The part I got wrong

I built the system before I measured the starting point.

I can tell you it has 11 roles, three workflows and a 400-line guide. Those are build numbers. They are not result numbers.

I did not record a clean manual baseline before building it. I cannot tell you honestly that the website team saves a specific number of hours or catches a specific percentage more issues. Retrofitting a number would make the case study look stronger and the experiment weaker.

The next test is straightforward. I will take two comparable pages, give both routes the same brief, and record:

Record Baseline · not recorded
Elapsed review time.
Issues found and accepted.
Changes I reject.
Rework needed after the review.

That will tell me whether the system is faster, better, or simply more elaborate.

06What stuck

What earned a place

Five parts of the system have survived repeated use.

01

Give each role one question

A role becomes useful when its responsibility is narrow enough to judge. “Make this page better” is vague. “Check whether every claim has evidence” can pass or fail.

02

Write down what good means

The 400-line writing guide matters more than any individual prompt. It gives every role the same standards and stops the system guessing what “sounds like Steve” means.

03

Design the order

The output of one stage becomes the input to the next. A good role in the wrong place can still damage the result.

04

Remember decisions, not just outputs

The session manager records what changed and why. That stops each working session becoming another restart.

05

Keep accountability clear

AI can produce more options and run more checks. It does not inherit responsibility for the result. I still decide what is accurate, what is useful and what ships.

07The method

Run the experiment on one part of your work

Start smaller than I did. Choose one repeated task where the output is easy to judge.

Use The Edge Loop:
01
Pick a task you complete at least once a week.
02
Test one AI approach against how you do it now.
03
Measure at least three runs. Record time, accepted changes and rework.
04
Keep the approach only if the improvement matters to you.
05
Stack the winner into your workflow, then test the next stage.

Add another role only after the first one produces a measured gain and the next part of the work has a genuinely different way to fail.

That is how a personal AI system grows without turning into a collection of impressive-looking machinery.

The verdict
Kept. Claim pending.

The role structure, sequence, written standards and session memory stay. They have become part of how I run the site.

The performance claim remains pending until the matched test produces a number.

That is the point of testing AI on your own work. A convincing system is still a hypothesis until you measure what changed.

08The Experiment Log

The next experiment goes in The Experiment Log

Every fortnight I share one AI experiment on real work: the setup, what changed, the number and the honest verdict.

You also get one experiment to run on your own work before the next issue.

AI consulting

Bringing the same discipline into a firm?

I help firms turn scattered AI use into systems that work on repeat.

Explore AI consulting