Why I write my own rules before I let a cheaper model do the work

The price of machine intelligence has collapsed faster than any input cost I have worked with in twenty years of engineering and business. Stanford's 2025 Ai Index measured the fall precisely: querying a model that performs at GPT-3.5 level cost $20 per million tokens in November 2022 and $0.07 per million by October 2024, a drop of more than 280 times in under two years. Every provider now sells a menu, with a frontier model at the top and cheaper tiers underneath that handle routine work perfectly well. The obvious move for anyone running a business on Ai is to push as much work as possible down to those cheaper tiers, and I do exactly that across my own products. But nothing gets handed down until the rules for the work are written, and that habit has saved me more than the cheap models have.

The evidence that capability alone does not pay

MIT's Project NANDA spent 2025 studying how generative Ai actually lands inside companies, drawing on 150 interviews with leaders, a survey of 350 employees and an analysis of 300 public deployments. Its report on the state of Ai in business found that 95% of enterprise generative Ai pilots delivered no measurable return, despite an estimated $30 to 40 billion of investment behind them. The diagnosis pointed away from model quality and towards what the researchers called a learning gap: the tools never learn the organisation's workflows, context or standards, so the output never quite fits the work it was meant to improve.

The same pattern appears at the level of a single person. METR, a research group that measures Ai capability, ran a randomised controlled trial in 2025 with sixteen experienced open-source developers working on 246 real tasks in codebases they knew well. With Ai tools switched on, the developers completed their work 19% slower, and afterwards they estimated the tools had made them 20% faster. That gap between felt productivity and measured productivity is the expensive part, because nobody corrects a problem they cannot feel.

A cheap model produces plausible work at extraordinary speed, which means it also produces plausible mistakes at extraordinary speed, and the review burden lands on you either way. However, neither study reads to me as an argument against using Ai. Both describe what happens when capable tools are put to work without conditions attached.

Capability without a standard produces variation

I spent more than fifteen years in regulated manufacturing before I ran a consultancy, across aerospace and high-pressure gas systems, and in those industries nobody hands a task to a capable new operator and hopes for the best. The task arrives with standard work attached: the sequence, the tolerances, the checks to run before starting, and what to do when something does not look right. The operator might be excellent, and the standard still exists, because capability without a defined standard produces variation, and variation is where scrap, rework and risk accumulate. I watched one facility lose more than £1.1 million to scrap in each of two consecutive years, largely through poor control of equipment and process, so I am not speaking hypothetically about what variation costs.

A cheaper model is a capable operator with no memory of your last correction. Each session starts from zero, so if your standards live in your head you end up re-teaching them every time, usually after the work has already gone wrong, and the model fills every gap you leave with a guess delivered in a confident tone. The MIT learning gap is the same failure at enterprise scale, since the organisations getting nothing back were largely the ones that gave capable tools no defined standard to work to.

The reason people skip the rulebook is that a model, unlike a new operator, appears to understand you immediately. It answers in fluent prose, mirrors your terminology and never asks you to repeat yourself, so the natural conclusion is that written instructions are optional. The METR developers fell for a version of that illusion, feeling 20% faster while running 19% slower, and they were experts using capable tools on code they knew well. If experienced specialists misjudge the effect that badly, the rest of us should assume we will too, and put the standard in writing where feeling cannot argue with it.

What my rulebooks actually contain

Every project a model works on in my business has a standing rulebook, written before the work starts and updated whenever the work drifts. The one covering my code repositories fits on a single page, while the one governing how models write in my voice runs to several thousand words, because voice turned out to be the harder thing to specify. The rulebook lives next to the work, in the folder the model reads before it starts, and the same five things earn a place in every one of them.

House style comes first, and it is more specific than most people expect to need: the spelling conventions, the punctuation habits, the words I never use, the way numbers are reported. Second is a definition of done that can actually be checked, so a code change is finished when the production build passes cleanly, and a researched article is finished when every claim traces to a named source. Third are the boundaries, the short list of actions the model may never take without my approval, which in my case covers publishing anything, sending anything, deleting anything and spending anything. Fourth is honest reporting: the model states what landed, what failed and what was skipped, in that order, because a cheerful summary that hides a skipped step costs more than the failure itself. Fifth is a standing instruction to ask rather than guess whenever a question is genuinely open.

None of this is sophisticated, and that is the point. These are the controls a good operations manager would put around any new starter, applied to a worker that costs a fraction of a penny per task. The payoff is measured in review time: when the rules are written, my review of a model's work is a check against a known standard, ten minutes with a diff or a source list, instead of an open-ended edit that quietly takes the hour I thought I had saved. Delegation only pays when the checking costs less than the doing, and written rules are what tilt that equation in your favour.

Five rules to write before you delegate

If you are about to hand real work to a cheaper model, write these five rules first, in a file the model reads every time it starts.

  1. Capture your standards the first time you correct the model, not the third. Every correction you make is a rule you already hold and have never written down, so move it from your head into the rulebook the moment you notice it.
  2. Define done as something checkable. A passing build, a verified source or a reconciled total can be tested, whereas 'good enough' cannot, and a model will always claim good enough.
  3. List the actions that always need your approval. Publishing, sending, deleting and spending make up my list, and yours will probably look similar.
  4. Require honest reporting. Instruct the model to state what failed and what was skipped as well as what succeeded, and treat a report that hides either as a defect in the rulebook.
  5. Fix the rule, not just the output. Correcting an output fixes one task, while correcting the rulebook fixes every task that follows it.

The 280-fold collapse in the price of intelligence means capability is no longer the constraint on what you can delegate; the quality of your written standards is. That is good news for anyone who has run a team, a production line or a project office, because specifying work clearly is an old skill applied to a new worker. Write your rules before the next task goes out, keep them current, and the cheaper model will repay you for as long as you do.

Tommy Findlay

Chartered Engineer, MBA and Lean Six Sigma Black Belt. Founder of iS3, helping UK businesses adopt Ai with the discipline of an engineer.

How ready is your business for Ai?

Get your free readiness score in ten minutes.

Get Your Ai Readiness Score