CROSSPOST: ANDREW GELLMAN: “Replication Crisis”, Dan Ariely Files

“Does not replicate” is a very polite way of putting this. Overly polite. By a lot. By a factor of x100, perhaps:

Every time something of mine replicates, I breathe easier. Because it might be the case that J.P. Morgan did not add value, that equipment investment is not tightly correlated to industrial-age economic growth, that merchant rule is not very good and princely rule (cough, Habsburgs, cough) not very bad for mercantile-age prosperity, that stock-market valuations as a whole might easily be driven by earnings-growth extrapolation, et cetera, et cetera, et cetera. It might be the case that wishful thinking led me into a misleading specification search.

Share

For, often, things do not replicate. For two big kinds of reasons. And so here we are today:

Give a gift subscription

The first reason things do not replicate is this: When performing pretty much any empirical study, there will be somewhere between 10 and 20 decisions that are matters of craft, art, and feeling with respect to what is the right thing to do with the data.

If one does one’s best and makes the choices based on one’s vision of how the world and how knowledge work, about half of those work to strengthen and half of those work to weaken the pattern that is in the data. But if one decides for all 10 or 20 to choose the one that strengthens the pattern, then you wind up with the largest of the 1024 or perhaps 1,048,576 possible numbers one could wind up reporting. Ditto if you choose all 10 or 20 to weaken the pattern. Since we are all good Bayesians, if we believe in the pattern that will lead us to make more choices that strengthen it, and if we disbelieve in the pattern that will lead us to make more choices that weaken it.

This is the biggest reason that there is a replication crisis. Look, say, at Josh Angrist and my friend the late Alan Krueger’s 1991 “Does Compulsory School Attendance Affect Schooling and Earnings?” <https://www.jstor.org/stable/2937954>. There is no way their particular choice of baseline specification was not affected by the facts that they believed in the pattern they were finding and were good Bayesians. (Cf.: John Bound, David Jaeger, and Regina Baker, in “Problems with Instrumental Variables Estimation When the Correlation Between the Instruments and the Endogenous Explanatory Variable Is Weak” <https://www.jstor.org/stable/2291055>. Zvi Griliches warned me about this back in 1979: “Do not believe too strongly in your theories about what the data ought to say”, he told me, “for then you will torture the data, and they will confess”.

For outsiders, figuring out what to do with claims of failures to replicate is difficult: which was the study where the authors’ Bayesianship wound up putting its thumb on the scales, the original study or the replication study?

But then there is the other kind of replication crisis, the second reason things do not replicate: people who demand that their research assistants produce the right results on pain of not getting a good recommendation letter, or who choose coauthors who fudge the data, or who fudge the data themselves.

In my view, the evidence is now overwhelming that Dan Ariely’s work is in this second bucket. Believe nothing he has ever published. Andrew Gelman has a write-up. Kyle Hyndman and Aberto Bisin and Uri Simonsohn, Joe Simmons, and Lief Nelson have the goods: “Replication of ‘Procrastination, Deadlines, and Performance: Self-Control by Precommitment’” <https://journals.sagepub.com/doi/full/10.1177/09567976261460772> and “[138] Artificial Deadlines (Part 1): Evidence of Fraud in an Influential Study About Procrastination” <https://datacolada.org/138>:

Leave a comment


CROSSPOST: ANDREW GELLMAN: “Replication Crisis”, Dan Ariely Files

<https://statmodeling.stat.columbia.edu/2026/08/31/shocked/>

I am shocked, shocked to learn that this 2002 paper by Dan Ariely is based on dubious data and doesn’t replicate.

Posted on August 31, 2026 3:15 PM by Andrew

Dan Ariely—the much-decorated business school professor, retired Wall Street Journal columnist, Ted-talk star, NPR hero, Jeffrey Epstein contact, Founding Partner of Irrational Capital, insurance agent, inspiration for a hit TV sitcom, and author of the instant-classic children’s book, “The Adventures of Professor D.”—has had pretty much the worst professional luck that any scientist could have.

Over the period of decades, he keeps ending up as coauthor on journal articles don’t replicate and that turn out, through absolutely no fault of his own, to be based on dubious or fake data. Bad collaborators, lazy research assistants, missing computer files . . . who knows how this is all happening, but it’s bad luck for sure.

Uri Simonsohn, Joe Simmons, and Lief Nelson just came across another one:

A new paper in Psychological Science reports a failure to replicate Study 2 of Ariely and Wertenbroch’s influential article entitled, “Procrastination, Deadlines, and Performance: Self-Control by Precommitment.” The original study, published in Psychological Science in 2002, found that people performed better on a set of tasks when each task had its own externally imposed deadline than when people set their own deadlines or faced a single last-day deadline for all tasks. The paper has had a lasting influence. It has been assigned reading in many economics and psychology courses, and has more than 2,100 citations on Google Scholar. . . .

About 20 years ago, on April 20, 2006, one of the authors of the forthcoming replication, Kyle Hyndman, received the original data files in an email sent from ariely@mit.edu . . . .

But then when Hyndman and his coauthor, Aberto Bisin, attempted to reanalyze the original data as part of their replication effort, they report that this happened:

In October 2024, at the request of the editors, we shared with Dan Ariely an analysis of the contents from the file purportedly for their Study 2 and asked for permission to include a summary of it in the paper. Dan Ariely denied our request, arguing, among other things, that the files we received may not be the actual data. He did not subsequently provide us with any additional data from the original paper.

Damn. I hate when that happens.

Simonsohn, Simmons, and Nelson continue:

This motivated us to return to this paper and fully analyze the original data for the two main studies. We conclude that the data in Studies 1 and 2 were tampered with. . . .

Our assessment that the data were tampered with are based entirely on the analyses presented in our posts. Readers can review the evidence and draw their own conclusions. . . .

Now comes the fun stuff.

And, by “fun,” I mean “horrible.”

Here’s one of the graphs from the 2002 paper:

Wow—that looks pretty impressive! Everything’s exactly in order, the standard errors are small, but the gaps between conditions 1, 2, and 3 aren’t exactly equal. They’re slightly irregular: exactly equal could raise suspicion, but these data look like they could be real . . . at least they do, until you look at them more carefully.

Here come Simonsohn, Simmons, and Nelson to rain on the parade:

Red Flag #1: The Effect Is Too Big
As shown in the reprinted figure above, Ariely and Wertenbroch report a perfect pattern of results, for all three dependent variables, with a sample size of only 20 per condition. The effects are also large. Extremely, implausibly large. . . .

Red Flag #2: Duplicate Observations . . .
18 of the 20 participants in the Last Day Deadline condition had a “Corrections Twin”, another participant who found exactly the same number of errors for each of the three proofreading tasks. Interestingly, these twins have ID numbers that are exactly 10 positions apart (e.g., subject S1 and subject S11 are twins; so are S7 and S17; etc.). (There were no error twins in the other two conditions.)
The existence of so many of these twins – and all of them in only one condition – is inconsistent with these data being real. . . .

Red Flag #3: Things That Should Be Very Highly Correlated Aren’t Correlated At All
At the end of their study, Ariely and Wertenbroch purportedly “asked participants to evaluate their overall experience [of the proofreading task] on five attributes . . .
You might expect these judgments to be correlated. For example, if someone says they liked the task, you might also expect them to say that it was interesting.
In the replication, this was (super) true. Controlling for experimental condition, the partial correlation between liking and interest was, quite sensibly, close to perfect:

But in the original data, this relationship was not only imperfect; it was not there at all. Participants who said they liked the task more did not say that they found the task to be more interesting:

The problem is not limited to these subjective measures. Consider the fact that people did three very similar proofreading tasks, each with 100 mistakes. Surely, we’d expect people who do better on one task to also do better on another, nearly identical task. That simple fact should manifest in extremely large correlations between performance on one task and performance on another. And in the replication data it does, as the correlations range from +.74 to +.90. But in the original data it doesn’t, as the correlations range from +.03 to +.27. . . .

Red Flag #4: No Rounding In Self-Reported Minutes
As you’ll recall from a minute ago, Ariely and Wertenbroch (2002) purportedly asked participants to “estimate how much time they had spent on each of the three tasks” (p. 223). When people provide estimates like this, they tend to round. They usually say “20 minutes” or “30 minutes” instead of “17 minutes” or “32 minutes”. And, indeed, when the replicators asked people to report how many minutes they spent on each of the three tasks, 85% of them gave a round number:

This is what we’d expect humans to do.

But in the original data, they did not do that. Only 11.7% of estimated minutes were round, consistent with the 10% you’d expect by chance alone:

Simonsohn, Simmons, and Nelson conclude:

We are unable to generate a benign explanation for all of the anomalies presented here. The original findings are too large and yet they do not replicate; there are duplicated observations; correlations that should be very strong are often non-existent; and values that should be rounded are not rounded. Based on this evidence, we believe the data for Study 2 of Ariely and Wertenbroch (2002) were severely tampered with or fabricated to produce the desired results.

Can you believe the bad luck of Ariely, to have this happen to him over and over again? I imagine he’ll want to launch an investigation to catch the real faker or fakers. What a waste of time, though. This is effort that could otherwise be devoted to designing psychology experiments. Or delivering Ted talks. Or modifying paper shredders. Or something.

I recommend to Ariely that, moving forward, he choose his collaborators and research assistants more carefully, so that he doesn’t again get stuck with fabricated data.

Maybe he could have them sign some sort of honesty pledge?

In Ariely’s own words, Professor D is “a charming character doing his best, struggling, and using social science to his advantage.” That’s one way to put it!

P.S. I employ some humor in these posts, but this sort of story doesn’t make me happy, it makes me sad. It doesn’t surprise me—lots of people enjoy cheating, and some of them will find their way into academic research, and some of those people will be successful at it—indeed, cheating can make it easier to be successful, as you’re no longer bound by the truth—so I’m not surprised, but I hate to see it.

In all seriousness, this sort of fraud is disgraceful, and whoever did it should admit it and pay some restitution to the Association for Psychological Science and the thousands of subsequent researchers who have relied on these fake findings. They should also reimburse the government for funds received in any later research grants whose proposals relied on these results. Ariely’s a nice guy, I get that, but I think that at this point he should just share the names of his collaborators and research assistants who’ve been doing all this faking It’s not right that he gets all the blame and they get off scot-free.

P.P.S. See here for more from Simonsohn et al.

P.P.P.S. I also again want to thank Simonsohn et al. for their efforts here. They got sued for doing this sort of thing! But they’re not intimidated and they keep at it. Good for them.

I hope that, at the very least, Ariely can send a public note to Simonsohn et al. thanking them for uncovering all these problems in his published papers. When people point out problems in my published work, I appreciate it and I thank them.

<https://statmodeling.stat.columbia.edu/2026/08/31/shocked/>


Brad DeLong here: What do I think of this? I think Uri Simonsohn, Joe Simmons, and Lief Nelson are right. The data is not data. The study is not a study. Where in the workflow that produces the studies published under the name of Dan Ariely the fraud lies can still be debated, but that there is an overwhelming pattern of fraud somewhere in the workflow cannot be denied any longer.

Thus, first, as Andrew Gelman says: “Simonsohn & al…. got sued for doing this sort of thing! But they’re not intimidated and they keep at it. Good for them”. We all need to have their backs all the time. Hyndman Bisin’s, too.

Second, Dan Ariely now has a statement:

Dan Ariely: An Update: “Procrastination, Deadlines, and Performance: Self-Control by Precommitment” (2002) by Dan Ariely and Klaus Wertenbroch <https://danariely.com/dan-ariely-statement-on-2002-procrastination-study/>: ‘As a social scientist, I am hardwired to believe that when evidence underlying a hypothesis or conclusion changes, it is critical that the record be updated to reflect those changes – both as a means of preserving integrity and facilitating the pursuit of truth.

Recently, I was made aware that data underlying a 2002 paper about deadlines and procrastination that I co-authored contained serious anomalies. The documentary record I have at my disposal today about those experiments isn’t sufficient to answer the questions that have been raised, and more than two decades, and hundreds of experiments later, my memory is similarly insufficient. Moving forward, my responsibility lies in ensuring accuracy – in updating the record on these experiments and, along with my co-author, cooperating with the journal that first published our paper to support their reviews and retraction processes.

Irrationally yours, Dan Ariely

Share DeLong's Grasping Reality Weblog

I find some downers here: the elliptical nature of this statement, the failure to acknowledge what the “anomalies” are, or, indeed, to point readers at who and what the criticisms are; none of these is a very good look here. Especially not in light of the 2024 claim that the 2006 data file was, perhaps, not the data file.

Going forward, what should we do, besides being sure not to take Ariely’s word for anything?

I say: three big and two smaller things:

  • First, the real damage is the decade of downstream behavioral work that took this as foundational: This really bothers me. I invested a substantial amount of my human capital in trying to build a map between behavioral cognitive regularities and systemic-level outcomes. This polluted that task.

  • Second, make John Bound and company god-kings: Bound, Jaeger & Baker (1995) on “weak instruments” showed a near-zero first stage means even trivial instrument-error correlation swamps the estimate — random noise instruments produced similar “results.” The identification revolution that was supposed to fix social science had its own replication problem baked in at the start. Often internal validity is not external validity. And a distressing share of the time internal validity is invalid even on its own terms.

  • Third, we need to value post-publication peer review: The “top-five” journals do not want the work done. The people actually doing the work that journals won’t are volunteers doing a very important public good with no career reward and real reputational risk. This is a devastatingly damaging institutional gap.

Plus two more smaller things:

  • The fact that we audit the surprises and not the confirmations is backwards: This is the Bayesian trap that makes fraud and honest error hard to catch. And it generalizes.

  • The garden of forking paths, not p-hacking, is the real danger: p-hacking is detectible in a way that the “1024 forking paths” problem is not.

Subscribe now

Leave a comment

If reading this gets you Value Above Replacement, then become a free subscriber to this newsletter. And forward it! And if your VAR from this newsletter is in the three digits or more each year, please become a paid subscriber! I am trying to make you readers—and myself—smarter. Please tell me if I succeed, or how I fail…

##crosspost-andrew-gellman-replication-crisis-dan-ariely-files
##public-reason
##moral-responsibility
##crosspost
#andrew-gellman-replication-crisis-dan-ariely-files
#andrew-gellman
#replication-crisis-dan-ariely-files
#dan-ariely
#replication-crisis
#researcher-degrees-of-freedom
#forking-paths
#data-fabrication
#post-publication-review
#data-colada
#academic-incentives