CROSSPOST: ARIN DUBE: "Clever Hans" at Speed: Implications for "AI" Workflows & Spread
Agentic AI shines on verifiable, testable work—and quietly falls apart when the answer isn’t checkable. “Run and find out” only works when you can quickly and cheaply assess the answer. We don’t have thinking machines; we have very fast runners-who-find-out who we can wire into our workflows—if, that is, we can either (a) find out, or (b) are happy with slop for purposes of summarization, boilerplate, or ritual…
This this this this this:
Arin Dube: “AI” Take of the Day <https://x.com/arindube/status/2074620137106285007>: ‘Fundamentally, there is a world of difference in usefulness of agentic AI in doing tasks that are to a large extent verifiable, versus not. If you want to replicate something, or have verifiable targets the code needs to reproduce, it can be a highly effective tool. For example, the final product can be a software that “works” and passes some tests; or could be a figure that plots (known) data; etc. When there isn’t, it can be a mess, and whether it’s a mess is largely unknown (without incurring large time cost).
This is why many researchers I know have boatloads of Claude Coded (or Codex Coded) proto-findings that haven’t seen the light of day because … it would take a long time to actually have any confidence. We can try to come up with milestones, interim checks, etc., but frankly there are severe limits on what that can look like without incurring a large cost. This is why I think focusing AI tools to be more limited but effective at accomplishing verifiable tasks—and helping users actually find those limited spaces where it can be effective—would can be productivity enhancing than chasing a pipe dream of crossing some mystical AGI threshold…
Consider Clever Hans. Clever-Hans-the-horse could not add. But Clever-Hans-the-system could. The system consisted of Clever Hans, the floor against which the horse stamped its hoof, the air through which the sound travelled, and the human who got excited when he had stamped his hoof the requisite number of times. The whole system could add because it had a test of accuracy: excitement on the part of the human component.
LLMs-in-harness are Clever Hans at speed, stamping their hooves much much faster than a human possibly could. So as long as there is a test of correctness (which implies tests of incorrectness), run-and-find-out will get you to a right answer, and get you to a right answer quickly.
Perhaps I should say not “Clever Hans” but ‘“Rikki-Tikki-Tavi”?
And Peyman Shahidi chimes in:
Mert Demirer, John J. Horton, Nicole Immorlica, Brendan Lucier, & Payman Shahidi: Chaining Tasks, Redefining Work: A Theory of AI Automation <https://peymanshahidi.github.io/assets/pdf/chaining_tasks_ai_automation.pdf>: ‘Production is a sequence of steps that can be executed (1) manually, (2) augmented with AI, or (3) fully automated within contiguous AI-executed steps called “chains.” Firms optimally bundle steps into tasks and then jobs, trading off specialization gains against coordination costs. We characterize the optimal assignment of humans and AI to steps and the firm’s resulting job structure, showing that comparative advantage logic can fail with AI chaining. The model implies non-linear productivity gains from AI quality improvements and admits a CES representation at the macro level. Empirical evidence supports the model’s key predictions that (1) AI-executed steps co-occur in chains, (2) dispersion of AI-exposed steps lowers AI execution at the job level, and (3) adjacency to AI-executed steps increases the likelihood that a step is AI-executed…
Perhaps the fundamental divide in AI usefulness is not between “narrow” and “general” intelligence, but between tasks where built-in very cheap and very quick tests of correctness—or close-enough-ness—can be easily added, and those without them. Today LLMs can race through verifiably-correct steps by running-and-finding-out, but without the “finding-out” part, they flail in the dark elsewhere. Dube has the prose diamond, Shahidi and his coauthors have a simulation framework.
And now my task is to try to see if this bears in any obvious and insightful way on exactly what the secret sauce of Anthropic’s harness is.
