Running an agency · September 4, 2026 · 6 min
How to tell if AI work is actually getting better
One number tells you whether AI is doing your client work or drafting near it: how much of each draft you rewrite before it goes out. If it is the same in month three as in month one, you bought a tool, not a colleague — and most vendors have never measured it.
By Islam Hachimi, Founder
There is one number that tells you whether AI is doing your client work or just drafting near it: how much of each draft you rewrite before it goes out. If that number is the same in month three as it was in month one, you have bought a tool, not a colleague.
Everyone selling AI to service businesses will show you a good first draft. A good first draft is table stakes now — the models are excellent and they are the same models everyone else is calling. The question that separates the products is what happens to the second month, and almost nobody will answer it, because almost nobody is measuring it.
Why the first draft stopped being evidence
Two years ago a competent draft was the whole demo. Today any vendor can produce one, because the intelligence is bought from the same three or four places and the price of it keeps falling. What a vendor cannot buy is knowledge of how you write, what your client will not tolerate, which numbers you always lead with, and the paragraph you delete every single time.
That knowledge only exists in one place: the edits you make. A product either captures them and changes because of them, or it hands you a slightly different first draft forever.
The measurement, and you can run it yourself
Keep the draft. Keep what you actually sent. Compare them. Not by feel at the end of a quarter — literally keep both files and count what changed, month over month.
- Count in words, not characters. Fixing a typo and rewriting an argument are different events, and a character count rates them by length.
- Measure per person. Two partners in one firm have different standards and averaging them hides both.
- Give it two months before you conclude anything. A curve drawn through four points is a hope, not a trend.
Ask any vendor for their edit rate over time, per customer. The answer tells you what you need to know either way. If they have never measured it, they have been selling you the first draft.
Why so few can answer it
It is not shyness. Measuring this means keeping the draft and the approved version as separate records, which means an approval step that is part of the product rather than an email thread. Most tools generate into a document you then edit in place, and the original is gone the moment you touch it. The evidence is destroyed by the workflow.
The wider numbers suggest this matters more than the category admits. Retention across AI-native software runs at roughly half that of ordinary business software, and the failures are not usually failures of capability — they are products that were impressive in month one and identical in month six. One well-funded bookkeeping service that shut down in late 2024 had raised over $100M and died of churn, not of bad output.
What we do about it, plainly
We keep both versions of everything. The draft the machine produced and the version you approved are two separate records, and the difference between them is measured and shown back to you in your own console. It is the number we ask to be judged on.
We are early enough to be honest about the state of that evidence: the mechanism is built and the curve is being collected now, per customer, and it is allowed to say we cannot tell yet rather than flatter itself on thin data. When there is enough of it to publish, we will publish it — including if it is flat, because a claim that only reports good news is not a measurement.
Whether you buy anything from us or not, keep your drafts. It is the cheapest possible insurance against paying a subscription for two years to a product that never learned your name.