Stop scripting the model. Brief it.
Most people improve their prompts by making them longer. It rarely works, and you cannot tell which of your 40 instructions did the job.
State what you need, say how you will judge it, and let the model work out the how.
Lead with the outcome, not the instructions.
A long prompt feels like control. It is not. It tells the model what to do without ever telling it what the work is for, and no line in it can pass or fail.
A brief does the opposite. It names the result, the reader, the checks and the boundaries, then leaves the approach open. It is usually shorter than the prompt it replaces.
What the prompt has to carry
Then → now- Every step, in order
- Every edge case, pre-empted
- Your chosen approach
- The output format
- The outcome
- Who it is for
- What good looks like
- What must not slip
The instructions did not disappear. They moved from describing the route to describing the destination and the test.
Six parts, and none of them describes how to do the work.
Part 3 is the one people leave out, and it is the one that changes the result. Every part is worked through in full in the second half of this page.
That is the whole idea. The tool below turns it into something you can use in the next five minutes.
Build your brief.
Seven questions. Your answers assemble into a brief you can paste into ChatGPT, Claude, Gemini or anything else. About five minutes the first time, and about one minute once the questions are familiar.
Think of it as scaffolding. The seven questions are the method: the tool shows you the shape they make, and you shape the brief from there to fit your own work. Once the questions are familiar you will write them straight into the chat box.
Nothing you type here is sent anywhere. It stays in your browser.
What kind of task is this?
This tunes the suggested checks in step 4. Skip it if none of them fit.
What has to exist when this is done?
One or two sentences. Say what it is and what it has to achieve.
For example: a one-page summary of this report that lets my director decide whether to fund phase two.
Who receives it, and what do they do next?
Naming the reader and their next action decides tone, length and structure without you specifying any of them.
For example: our finance director, who will read it in a meeting and either approve the spend or send it back.
What does good look like?
Tick the checks that matter, then add your own. A good check is one someone else could apply to your output and reach the same answer as you.
What must not slip?
The deal-breakers. State them separately, because a model trying to be helpful will otherwise trade one away to satisfy something else.
What should it work from?
What form, and how much latitude?
For example: one page of plain prose. An email under 200 words. A table with one row per option.
Ask first, for anything that matters. One clarifying question costs less than a confident wrong answer.
Make it mark its own work, then mark it yourself.
The last line of your brief asks the model to check its draft against your numbered criteria and report what it met, what it missed and what it assumed. Read that list before you read the output. It tells you where your brief was thin.
You get a short list back: met, not met, assumed. Do not treat it as proof, because a model can mark its own homework generously.
What it is reliably good at is telling you what it had to invent, which is the failure you most need to catch.
Then run your own three questions.
- Did it pass every check you wrote?
- Is it better than what you would have produced in the same time?
- Would you send it, counting the edits you actually had to make?
Fix the brief, not the output
When a briefed output disappoints you, you can point at the line that caused it: check 3 was vague, the boundary was missing, you never named the reader. Change one line and run it again.
When a 600-word script disappoints you there is nothing to point at. You add another rule and hope. That is why a brief improves over a month and a mega-prompt does not.
When a brief works, the part worth keeping is not the text the model produced. It is the answers that made the difference: the check that caught something, the boundary that stopped a bad habit, the way you described the reader.
Reuse those on the next task of the same kind and adjust them as you learn. Prove it across three pieces of work before you make it part of how you work.
Everything above is enough to run your first brief. What follows is why this works, how to write a check that is actually checkable, and how to test the whole idea against the prompt you use today.
What each part is doing.
Nobody hands a good freelancer a 2,000-word script. You tell them what you need, who it is for and how you will judge it. Then you get out of the way. Treat the model the same.
The outcome
What has to be true when this is done.
Write me a summary of this report.
I need a one-page summary that lets a busy director decide whether to fund phase two.
The second version tells the model what the summary is for, so it knows what to keep and what to cut.
The reader
Who receives it, and what they do next.
Make it professional.
It goes to our finance director, who will read it in a meeting and either approve the spend or send it back with questions.
Naming the reader and their next action decides register, length and structure without you specifying any of them.
What good looks like
The checks the output must pass.
Make it high quality and well structured.
It must name the decision on the first line, give the three numbers that most affect it, fit on one page, and flag anything the report leaves unresolved.
Every item can be checked by someone who did not write it. That is what makes it a check rather than a wish.
Writing down what good looks like.
This is the part people leave out, and it is the part that changes the result. Ask most people how they judge an AI output and the answer is a feeling: it reads well, it seems about right, I would have written it differently.
A feeling cannot be handed to a model and cannot be improved. A check can.
A good check is one someone else could apply to your output and reach the same answer as you.
Fits on one page. Under 200 words. Names three sources. Exactly one recommendation.
The easiest to write and the easiest to verify. Start here.
Names the decision on the first line. Includes the new date. Flags anything unresolved. Attributes every figure to a source.
Nearly as reliable as countable, and it covers most of what actually matters.
Reads like a person wrote it. Lands the point without overselling.
Real, but useless as written. Make it checkable by naming the test: a colleague reading it cold could restate the point in one sentence.
The check worth adding to everything
Every fact must come from the material I gave you. Anything you could not source, list rather than fill in.
Models close gaps politely. Asked to write meeting minutes without being told who took them, a model will often supply a plausible name and put it in the document. It looks correct, which is exactly the problem.
The fix is not a better model. It is a brief that says what to do when something is missing.
You do not need 12. Three checks that genuinely decide whether you would send the work beat a dozen you wrote to feel thorough. Add another only when a real failure shows you the gap.
What must not slip
The quality that cannot drop, and the things that must not happen.
Do not make mistakes.
Do not invent a figure. Do not recommend anything the report does not support. Do not exceed one page.
Deal-breakers are worth stating separately. A model trying to be helpful will otherwise trade one of them away to satisfy something else.
What to work from
The material, and which source wins in a conflict.
You know the context.
Work from the attached report and the two emails below. Where they disagree, the report is correct. If a figure is in neither, say so rather than filling the gap.
Most invented facts are gaps the model closed politely. Naming the authoritative source and giving explicit permission to say “not in the material” closes them properly.
The latitude
Where the model may choose, and permission to ask first.
Follow these steps exactly.
Choose the structure you think serves the decision best. If anything essential is missing, ask me before you start.
This is the part people leave out second most often, and it is the part that pays. You are buying the model’s judgement rather than renting its typing.
Six parts, and not one of them describes how to do the work. That is the shift. You specify the destination and the test, then trust the model with the route.
The same task, scripted and briefed.
One task, two ways of asking. This is illustrative, not a result. The point is the shape of the difference, not a score.
You are an expert business writer with 20 years of experience. Write a professional summary of the attached report. Use a formal but approachable tone. Open with an executive summary of no more than three sentences. Then write a section on background, a section on findings and a section on recommendations. Use bullet points where appropriate. Bold your key terms. Avoid jargon. Be concise but thorough. Make sure it is well structured and easy to scan. Include a call to action at the end.
Shortened. The original runs to about 600 words.
- It never says the summary exists to get phase two funded, so the model optimises for coverage instead of a decision.
- It specifies four sections, which is your structure and not necessarily the right one.
- Nothing in it can fail. There is no line you could point at when the output disappoints you.
- The model knows what the document is for, so it can decide what to cut.
- Every line under “what good looks like” can pass or fail, so a disappointing output points at a specific line.
- It is shorter than the script it replaced, and it does more.
The brief is not longer than the script. It is pointed somewhere.
Mega-prompts solved a problem the models no longer have.
Long instruction stacks were a sensible workaround. Earlier models could not hold a goal across a long task, could not plan an approach and could not tell you when they were stuck. Spelling out every step compensated for all three.
Frontier models plan. They hold an objective, choose a route, and tell you what is missing before they start. The workaround outlived the weakness.
Three things happen when you keep scripting anyway.
You get compliance, not the outcome
A long script tells the model what to do. It never tells the model what the work is for. So the model satisfies your instructions and misses your intent, and technically it did exactly what you asked.
You cannot tell what worked
When a 40-instruction prompt produces something good, you do not know which instruction earned it. Nothing in the stack can be tested, so nothing in the stack can be improved. You keep all 40 out of superstition.
You cap the output at your own imagination
Every step you specify is a step the model cannot improve on. If your approach is the third-best approach, a detailed script guarantees you the third-best answer.
A prompt you cannot grade is a prompt you cannot improve.
This started as an argument about mechanism rather than a result I had measured. I have now run the head-to-head once: a 733-word prompt with 20 numbered steps against an 87-word brief, three runs each, on a task where the biggest support category by volume and the biggest by value pointed in opposite directions. The scripted prompt scored 52.5 out of 100 and put £11,000 behind the wrong problem. The brief scored 100. Read the full experiment.
One test is not a law. The task was built to contain that trap, so it does not show that briefs beat long prompts in general, and I would bin any claim that it does. What it shows is narrower: when the route you prescribe is wrong for the data, a capable model follows it anyway, writes down that it is doing so, and hands you exactly what you asked for.
The next section gives you the protocol to run it on your own work, which is the only result that should change what you do.
Do not take my word for it. Run the test.
You almost certainly have a long prompt you rely on. Put it against a brief and find out which one earns its place. This is a real experiment with a hypothesis, a method and a number, and it takes about an hour.
About one hour · six runs, one scoring sheet
Pick a task that is actually hard
A short email is not a test. Both approaches will pass and you will learn nothing. Choose something with source material to work from, at least three constraints and a real quality bar. If you would be embarrassed to send a bad version, it is a fair test.
Write the checks before you run anything
Take three to five checks from step 4 of the builder. Write them down now, while you cannot see any output. Criteria written afterwards bend to fit whatever arrived.
Run both arms three times
Same model, same source material, same day. Your existing prompt three times. Your brief three times. Start a fresh chat every time so nothing carries over. Edit nothing.
Score all six against the same checks
Mark each output against each check: pass or fail. Score them without knowing which arm produced them if you can manage it, and at minimum score all six in one sitting so your standard does not drift.
Read the verdict honestly
- Which arm passed more checks?
- Which arm failed the same way twice?
- When something failed, could you point at the line that caused it?
The third question decides it. A repeatable failure you can trace to a specific line is fixable in two sentences. A repeatable failure inside a 600-word script is not.
The mega-prompt may win on your task. If it does, that is a real finding and worth keeping. I ran this same test on my own work and published the result, including the limits that stop it proving as much as the headline suggests.
Run it through The Edge LoopThe next experiment lands in the newsletter.
Get it in The Experiment LogWhen a detailed script is the right tool.
Briefing is a better default. It is not a rule. Four cases where spelling out the steps is correct, and pretending otherwise would be the same hype this page argues against.
The format is the requirement
A report that must match a template, a data extract that feeds a spreadsheet, anything where a different-but-better structure counts as a failure. Specify it exactly. There is no judgement to buy.
You have already found the winning recipe
If you have tested an approach on 20 pieces of the same work and it wins every time, that is not a mega-prompt. That is a documented method, and you earned it by experimenting. Write it down and reuse it.
The process is what gets audited
Regulated work, clinical or legal steps, anything where someone will later ask how the output was produced. The sequence is the deliverable. Specify it.
You are using a small or fast model
The reasoning that makes briefing work is not evenly distributed. A small, cheap model handling high volume often does need the steps spelled out. Match the instruction to the model, not to a principle.
Notice what the exceptions have in common. In every one of them you already know the right approach: because it is fixed, because you tested it, or because someone else mandated it.
Brief when you do not know the best route. Script when you do.
This is what Test looks like when the thing under test is a prompt.
The Edge Loop is five steps: pick a real task, test an AI approach against how you do it now, measure the result, keep only what works for you, stack it so the gains compound.
An outcome brief is what step 2 looks like in practice. The checks you wrote in step 4 of the builder are what make step 3 possible: without them, “measure the result” has nothing to measure against. The head-to-head above is one full turn of the loop, run on your own prompt.
One task. One brief. One number. Then again.
See The Edge LoopQuestions worth asking.
Is this just prompt engineering with a different name?
No, and the difference is the check. Prompt engineering optimises the instruction. An outcome brief states the result you need and the test it must pass, then leaves the approach open. The instruction gets shorter rather than more clever, and you end up with something you can grade.
Does this work with ChatGPT, Claude and Gemini?
Yes. The brief is plain text and nothing in it is specific to one provider. It works best with the reasoning models each of them offers, because the whole approach depends on the model being able to plan its own route.
My prompt already works. Why change it?
If it works on repeat and you can tell why, keep it. Case 2 in the exceptions covers exactly that. The test is whether you could point at the line responsible when it disappoints you. If you could not, you are trusting something you cannot fix.
Is a shorter brief not just less context?
Context and instruction are different things. An outcome brief usually carries more context, not less: the reader, the purpose, the authoritative source, and what to do when something is missing. What it drops is the step-by-step method, which is the part the model can work out for itself.
How many checks should I write?
Three to five. Enough that they genuinely decide whether you would send the work, few enough that you will actually apply them. Add another only when a real failure shows you the gap.
Does the tool send my answers anywhere?
No. The Brief Builder runs entirely in your browser. Nothing is stored, nothing is transmitted and there is no sign-up. Copy your brief before you leave the page, because the answers are not saved.
One real experiment. One result you can use.
I test AI on real work, show the number and share the honest verdict. The head-to-head above is the next experiment on my own list, and the result goes out to subscribers either way.
See the methodOne real experiment every fortnight. Keep only what works for you.