███╗   ███╗ ██╗   ██╗  ██████╗ ███████╗ ██╗
████╗ ████║ ╚██╗ ██╔╝ ██╔════╝ ██╔════╝ ██║
██╔████╔██║  ╚████╔╝  ██║      █████╗   ██║
██║╚██╔╝██║   ╚██╔╝   ██║      ██╔══╝   ██║
██║ ╚═╝ ██║    ██║    ╚██████╗ ███████╗ ███████╗
╚═╝     ╚═╝    ╚═╝     ╚═════╝ ╚══════╝ ╚══════╝
DemoAcademyPricing
Sign inBook a meeting
Changelog
2 August 2026
  • kernel
  • fix

Runs now stop when the work is done

A real invoice chase wrote its answer five minutes in. Everything the task asked for, correct. Then it spent twenty more minutes reading the same policy file again until a time limit killed it.

The run was recorded as failed. The customer would have been told their invoice chase did not work, while the finished answer sat in a sandbox we deleted a moment later.

The problem: the run ended when the agent said it was done. That is not the same question as whether the work is done. So the harness watches the file instead. When the result file is there and matches what the task promised to produce, the run ends. Nothing left to pay for.

Free runs also got faster and cheaper. Free was capped to the cheapest model. That model kept getting one tool call wrong and repeated the same broken call eleven times without ever getting it right. It is 4x cheaper per token and far more expensive per job, because it loops until a ceiling stops it — and the ceiling is where the bill lands. Free now reaches the standard model. Spend is still capped in dollars, which was the control doing this job all along.

Bring one file. Ship one deliverable. See what stays.

Start free. Upload one piece of past work. Correct it once.

Start 7 days free

Drafts the work. Correct it once.

All systems operational

Ask an AI about us

  • Claude
  • ChatGPT
  • Perplexity

It reads the site and answers on its own. We do not get to edit what it says.

Product

  • What you get
  • Pricing
  • How do you make it?
  • What are you owed?
  • Changelog
  • What it runs
  • Free reports
  • Product map
  • Team
  • Blog
  • Glossary
  • AI Visibility Index
  • Sign in
  • Docs

Compare

  • vs ChatGPT, Claude, or whichever tab is already open
  • vs Grok Bot and the AI-employee platforms
  • vs Hiring an account manager
  • vs Profound
  • vs Otterly
  • vs Building it yourself
  • vs Zapier & n8n
  • vs Temporal
  • vs LangGraph
  • vs CrewAI & AutoGen
  • All comparisons

Legal

  • Privacy
  • Sub-processors
  • Terms
  • DPA
  • Security
███╗   ███╗ ██╗   ██╗  ██████╗ ███████╗ ██╗
████╗ ████║ ╚██╗ ██╔╝ ██╔════╝ ██╔════╝ ██║
██╔████╔██║  ╚████╔╝  ██║      █████╗   ██║
██║╚██╔╝██║   ╚██╔╝   ██║      ██╔══╝   ██║
██║ ╚═╝ ██║    ██║    ╚██████╗ ███████╗ ███████╗
╚═╝     ╚═╝    ╚═╝     ╚═════╝ ╚══════╝ ╚══════╝