███╗   ███╗ ██╗   ██╗  ██████╗ ███████╗ ██╗
████╗ ████║ ╚██╗ ██╔╝ ██╔════╝ ██╔════╝ ██║
██╔████╔██║  ╚████╔╝  ██║      █████╗   ██║
██║╚██╔╝██║   ╚██╔╝   ██║      ██╔══╝   ██║
██║ ╚═╝ ██║    ██║    ╚██████╗ ███████╗ ███████╗
╚═╝     ╚═╝    ╚═╝     ╚═════╝ ╚══════╝ ╚══════╝
DemoAcademyPricing
Sign inBook a meeting

Running an agency · September 4, 2026 · 6 min

How to tell if AI work is actually getting better

One number tells you whether AI is doing your client work or drafting near it: how much of each draft you rewrite before it goes out. If it is the same in month three as in month one, you bought a tool, not a colleague — and most vendors have never measured it.

By Islam Hachimi, Founder

There is one number that tells you whether AI is doing your client work or just drafting near it: how much of each draft you rewrite before it goes out. If that number is the same in month three as it was in month one, you have bought a tool, not a colleague.

Everyone selling AI to service businesses will show you a good first draft. A good first draft is table stakes now — the models are excellent and they are the same models everyone else is calling. The question that separates the products is what happens to the second month, and almost nobody will answer it, because almost nobody is measuring it.

Why the first draft stopped being evidence

Two years ago a competent draft was the whole demo. Today any vendor can produce one, because the intelligence is bought from the same three or four places and the price of it keeps falling. What a vendor cannot buy is knowledge of how you write, what your client will not tolerate, which numbers you always lead with, and the paragraph you delete every single time.

That knowledge only exists in one place: the edits you make. A product either captures them and changes because of them, or it hands you a slightly different first draft forever.

The measurement, and you can run it yourself

Keep the draft. Keep what you actually sent. Compare them. Not by feel at the end of a quarter — literally keep both files and count what changed, month over month.

  • Count in words, not characters. Fixing a typo and rewriting an argument are different events, and a character count rates them by length.
  • Measure per person. Two partners in one firm have different standards and averaging them hides both.
  • Give it two months before you conclude anything. A curve drawn through four points is a hope, not a trend.
Ask any vendor for their edit rate over time, per customer. The answer tells you what you need to know either way. If they have never measured it, they have been selling you the first draft.

Why so few can answer it

It is not shyness. Measuring this means keeping the draft and the approved version as separate records, which means an approval step that is part of the product rather than an email thread. Most tools generate into a document you then edit in place, and the original is gone the moment you touch it. The evidence is destroyed by the workflow.

The wider numbers suggest this matters more than the category admits. Retention across AI-native software runs at roughly half that of ordinary business software, and the failures are not usually failures of capability — they are products that were impressive in month one and identical in month six. One well-funded bookkeeping service that shut down in late 2024 had raised over $100M and died of churn, not of bad output.

What we do about it, plainly

We keep both versions of everything. The draft the machine produced and the version you approved are two separate records, and the difference between them is measured and shown back to you in your own console. It is the number we ask to be judged on.

We are early enough to be honest about the state of that evidence: the mechanism is built and the curve is being collected now, per customer, and it is allowed to say we cannot tell yet rather than flatter itself on thin data. When there is enough of it to publish, we will publish it — including if it is flat, because a claim that only reports good news is not a measurement.

Whether you buy anything from us or not, keep your drafts. It is the cheapest possible insurance against paying a subscription for two years to a product that never learned your name.

Everything described here is in the kernel that runs Mycel — the scheduler, the wedges, the guards, and the tests that hold them.

Read the kernel →More writing →

Read next

  • Why your cold email stopped arrivingCold email rarely dies from one bad day. It dies from a fortnight of small mistakes — a ramp that looks like a leak, a sequence that retries dead addresses, one domain quietly carrying the whole day — and by the time you notice, the fix is a new domain and three months of waiting.
  • What a junior actually costsA junior on $55,000 costs about $100,000 in their first year once you count payroll tax, health insurance, recruiting, equipment and the fifteen days of your own time it takes to ramp them. Here is the full arithmetic, and the test for which work is worth hiring for and which is a process problem wearing a headcount costume.
  • Chasing an invoice without losing the clientMost overdue invoices are not a refusal to pay — they are an invoice that reached one person who is not the person who pays. A four-rung ladder where the escalation is in the specificity rather than the tone, the three rules underneath it, and the four things to fix before you fix the chasing.

Take the client you turned down last month.

Describe what you deliver and the first draft exists before you have finished your coffee.

Start 7 days free

The first AI delivery firm. You sign.

All systems operational

Ask an AI about us

  • Claude
  • ChatGPT
  • Perplexity

It reads the site and answers on its own. We do not get to edit what it says.

Product

  • What you get
  • Pricing
  • Changelog
  • What it runs
  • Free reports
  • Product map
  • Team
  • Blog
  • Glossary
  • AI Visibility Index
  • Sign in
  • Docs

Compare

  • vs ChatGPT, Claude, or whichever tab is already open
  • vs Grok Bot and the AI-employee platforms
  • vs Hiring an account manager
  • vs Profound
  • vs Otterly
  • vs Building it yourself
  • vs Zapier & n8n
  • vs Temporal
  • vs LangGraph
  • vs CrewAI & AutoGen
  • All comparisons

Legal

  • Privacy
  • Sub-processors
  • Terms
  • DPA
  • Security
███╗   ███╗ ██╗   ██╗  ██████╗ ███████╗ ██╗
████╗ ████║ ╚██╗ ██╔╝ ██╔════╝ ██╔════╝ ██║
██╔████╔██║  ╚████╔╝  ██║      █████╗   ██║
██║╚██╔╝██║   ╚██╔╝   ██║      ██╔══╝   ██║
██║ ╚═╝ ██║    ██║    ╚██████╗ ███████╗ ███████╗
╚═╝     ╚═╝    ╚═╝     ╚═════╝ ╚══════╝ ╚══════╝