Results of the main run

27 September 2026, 21:09 to 21:21 CEST. 24 tasks, three ways each, 72 runs in a random order, one fresh session and one budget per run, model Claude Sonnet 5, total cost 4.04 USD. The specialists ran live on this domain.

Headline numbers

Checking tasks correct (of 16)Help usedA · Work alone11/160 of 24 delegatedB · Find help12/161 of 24 used the directoryC · Know who to call13/1611 of 24 delegated
routechecking taskssimple tasksdelegatedmean costmedian time
A alone11 of 167 of 80 of 240.056 USD10.3 s
B with directory12 of 167 of 81 of 240.058 USD10.5 s
C with fixed list13 of 167 of 811 of 240.055 USD11.2 s

Success criteria, fixed before the run

Five of six passed. The one that failed is the one the experiment was for: with the directory, the agent was not clearly better than alone.

What happened, per task type

task typeABCnote
Valid links (2 tasks, 10 links each)222
Broken links (2 tasks, 10 links each)222
Redirects with expected target (2 tasks, 8 each)212B accepted an http page for an expected https target
Page anchors (2 tasks, 10 each)222
Source supports the claim (2 tasks, 5 claims each)122A called a supported number unsupported
Source supports it only partly (2 tasks)011the hardest category for every route
Source contradicts the claim (2 tasks)111one claim was answered contradicted by all three routes against the key; recorded as a key dispute
Evidence absent (1 task)111
Source not readable (1 task)000all routes read a login page as readable; the key says unverifiable

The simple tasks (sorting, extracting, shortening a headline) were solved alike by all routes, 7 of 8 each; the one miss was a headline that all three cut to seven or eight words instead of six, a defect of the task.

The observation that matters

Route B was told it could search the directory and hire a specialist, and did so once in 24 tasks. Route C had the same specialists named in its instructions and delegated 11 times. The one directory lookup worked: it returned a reachable Link Auditor, the call succeeded, the task was solved. The directory was not broken, it was not consulted. Whether that is a matter of how the tool was described, how the task was phrased, or a sound judgement by the agent that it could do the work itself, is the next question, testable on the same 24 tasks with one change at a time.

Before the main run

On eight small development tasks (two or three links each) the agent alone was already perfect, and specialists added nothing. On eight harder development tasks (twelve links, ten anchors, five claims each) the agent alone solved 3 of 6 checking tasks and with the specialists listed 6 of 6, at equal cost. That gate was passed before the main run was built.

Limits

Experiment by Adrian Föhl. Method, rules and the decision fork were fixed before the run; a blind second opinion from a model from another vendor reviewed the design. Write-up to follow at adrianfoehl.com.