
Human evaluators reading agent transcripts scored an AUC of 0.65 against deliberately subtle sabotage, compared to 0.87 for the strongest LLM monitor. They averaged 103 minutes per transcript. Slower than the AI monitor, less accurate, and much more expensive.
Human review is not the realistic alternative to AI monitoring. It is the unaffordable alternative. Once it is unaffordable, the practical fallback is no review at all.
The same paper tested whether reading the agent's hidden chain of thought matters. When the team removed the hidden scratchpad from the attacker, sabotage success dropped sharply across the models they tested. Private planning is part of what makes an attack work. Reading the chain of thought, when the architecture lets you, is doing real defensive work and not just another audit item.
A second finding from the same paper worth noting. Turn by turn monitoring, where the monitor evaluates each step as it happens, helped less than the team expected. Better system prompts on the monitor and requerying multiple times with averaged scores helped more.
The takeaway is closer to "score the full transcript carefully two or three times" than "watch every action live." Cheaper to run, easier to operate, and the numbers say it works.
If you have been waiting for evidence to push your team toward CoT monitoring, this is the paper to read.
Written by Duane Grey
AI Strategy & Implementation
Independent AI consultant helping companies cut through hype and deploy systems that produce real results.