A caller insubstantial successful Psychological Science (.htm) reports a nonaccomplishment to replicate Study 2 of Ariely and Wertenbroch’s influential article entitled, “Procrastination, Deadlines, and Performance: Self-Control by Precommitment.” The original study, published successful Psychological Science successful 2002 (.htm), recovered that group performed amended connected a group of tasks erstwhile each task had its ain externally imposed deadline than erstwhile group group their ain deadlines aliases faced a azygous last-day deadline for each tasks. The insubstantial has had a lasting influence. It has been assigned reference successful galore economics and psychology courses, and has much than 2,100 citations connected Google Scholar.
![]()
Because this insubstantial has been truthful influential, it is worthwhile to return a adjacent look astatine the original study to effort to understand why it did not replicate. We did that. This station – and the adjacent 1 – is astir what we found.
***
About 20 years ago, connected April 20, 2006, 1 of the authors of the forthcoming replication, Kyle Hyndman, received the original information files successful an email sent from [email protected] [1]. And astir 3 years ago, connected August 9, 2023, a week aft Francesca Gino sued america for $25 million, we received an out-of-the-blue email from Hyndman successful which he sent those files to us. We performed speedy analyses of the data, and past had a speech pinch Hyndman and his co-author, Alberto Bisin. In that conversation, they told america they were going to behaviour a replication, and, uncovering ourselves engaged pinch the lawsuit, we near it astatine that.
We precocious learned that their replication was forthcoming successful Psychological Science. And upon reference Footnote 14 of their paper, we besides learned this:
. . . In October 2024, astatine the petition of the editors, we shared pinch Dan Ariely an study of the contents from the record purportedly for their Study 2 and asked for support to see a summary of it successful the paper. Dan Ariely denied our request, arguing, among different things, that the files we received whitethorn not beryllium the existent data. He did not subsequently supply america pinch immoderate further information from the original paper. Consequently, we are incapable to supplement our replication workout pinch immoderate further study of the files we received successful 2006 aliases immoderate different data.
This motivated america to return to this insubstantial and afloat analyse the original information for the 2 main studies. We reason that the information successful Studies 1 and 2 were tampered with. In 2 posts, we coming the grounds that led america to this conclusion. Today’s station focuses connected the study that grounded to replicate (Study 2), and our adjacent station is astir Study 1.
Our appraisal that the information were tampered pinch are based wholly connected the analyses presented successful our posts. Readers tin reappraisal the grounds and tie their ain conclusions.
To the champion of our knowledge, Klaus Wertenbroch has ne'er had entree to immoderate type of the information for immoderate of the studies. And, we judge it is acknowledgment to him that we do. When Kyle Hyndman reached retired to the authors backmost successful 2006, Klaus replied pinch this email [2]:

Our ResearchBox contains the information and codification to reproduce each of the results successful this post.
Finally, it should beryllium noted that erstwhile we shared these posts pinch Ariely and Wertenbroch a fewer weeks ago, they reached retired to Psychological Science to petition that the article beryllium retracted. As of this writing, that process is ongoing.
The Study That Did Not Replicate: Study 2 of Ariely and Wertenbroch (2002)
As noted above, Ariely and Wertenbroch explored really deadlines power performance. In a discourse successful which group had aggregate tasks to perform, the authors hypothesized that group would execute amended successful the look of evenly spaced deadlines for those tasks, alternatively than erstwhile they were each owed astatine the end.
The research progressive an incentivized proofreading task. Each subordinate received 3 10-page documents, each containing 100 “grammatical and pronunciation errors” (p. 222). Participants were tasked pinch uncovering and correcting those errors.
Sixty participants were randomly assigned to 1 of 3 conditions, precisely 20 participants successful each condition:
Condition 1. Evenly Spaced Deadlines. One archive was owed each week, truthful aft 7, 14, and 21 days.
Condition 2. Set Your Own Deadlines. Participants chose their ain deadlines (within 21 days).
Condition 3. Last Day Deadline. All 3 documents were owed connected the last (21st) day.
The results perfectly and powerfully supported the authors’ hypothesis. Participants fixed evenly spaced deadlines did overmuch better, successful position of performance, delays, and net [3].

Do We Have The Original Data?
As a reminder, successful 2023 Hyndman sent america files he received from [email protected] successful 2006. There were 3 Excel files – information for a aviator study, for Study 1, and for Study 2 – each pinch record properties indicating that the information were “Last saved by” “Dan Ariely”.
With these files we are capable to reproduce each 9 intends and each 9 modular errors shown successful the fig above, arsenic shown visually successful this footnote: [4]. We besides successfully reproduce the six different intends reported successful the matter [5].
Red Flags
In our analyses we identified 4 awesome reddish flags. We talk each successful turn.
Red Flag #1: The Effect Is Too Big
As shown successful the reprinted fig above, Ariely and Wertenbroch study a cleanable shape of results, for each 3 limited variables, pinch a sample size of only 20 per condition. The effects are besides large. Extremely, implausibly large.
Consider the proofreading capacity results. Participants pinch Evenly Spaced Deadlines made an mean of 136.1 corrections, whereas those pinch the Last Day Deadline made an mean of only 71.1 corrections, astir half arsenic many. This effect has a Cohen’s d = 2.5, indicating that the information intends are 2.5 modular deviations apart. The relationship betwixt experimental information and number of corrections is r = .79.
To admit that this effect is conscionable excessively big, see it successful the discourse of different effect sizes. An effect size of d = 2.5 is larger than evident effects we announcement successful mundane life, effects that tin easy beryllium seen pinch the naked eye. For example, it is overmuch larger than the effect of gender connected tallness (men are taller: d ≈ 1.8) and connected number of shoes owned (women ain much shoes: d ≈ 1.2; spot Colada[18]). It is besides larger than immoderate manipulation checks. For example, Petty and Cacioppo (1984) study that participants exposed to messages containing 9 arguments said that they encountered much arguments than group exposed to messages containing 3 arguments. This has to beryllium true. And it was true, but only to the tune of d = 1.49 [6]. It is not plausible that deadlines power proofreading capacity much powerfully than the number of arguments influences the perceived number of arguments.
Effect sizes greater than aliases adjacent to 2.5 are not intolerable – they are sometimes observed pinch manipulation checks – but they are extraordinarily uncommon for non-obvious psychological findings, peculiarly for a measurement for illustration proofreading correction detection, which is apt to beryllium noisy, and highly adaptable crossed people.
Another measurement to admit the enormousness of this effect is to look astatine the distribution of the limited adaptable crossed conditions. The fig beneath shows that they hardly overlap. For instance, whereas cipher successful the Last Day Deadline information made much than 100 corrections, 90% of the participants successful the Evenly Spaced Deadlines information did:
Red Flag #2: Duplicate Observations
If looking astatine Figure 2 you thought, “wait, why are location truthful galore reddish bars pinch 2s?”, bully catch. That is weird. The 2s correspond group who recovered precisely the aforesaid full number of corrections made crossed 3 tasks. But it’s really weirder than that. These participants recovered not conscionable the aforesaid number of corrections successful total, but besides made the aforesaid number of corrections for each of the 3 separate proofreading tasks.
Here is simply a screenshot of the original information file, formatted and sorted to beryllium easier to digest:

We spot that 18 of the 20 participants successful the Last Day Deadline information had a “Corrections Twin”, different subordinate who recovered precisely the aforesaid number of errors for each of the 3 proofreading tasks. Interestingly, these twins person ID numbers that are precisely 10 positions isolated (e.g., taxable S1 and taxable S11 are twins; truthful are S7 and S17; etc.). (There were nary correction twins successful the different 2 conditions.)
The beingness of truthful galore of these twins – and each of them successful only 1 information – is inconsistent pinch these information being real.
Red Flag #3: Things That Should Be Very Highly Correlated Aren’t Correlated At All
At the extremity of their study, Ariely and Wertenbroch purportedly “asked participants to measure their wide acquisition [of the proofreading task] connected 5 attributes: really overmuch they liked the task, really absorbing it was, really bully the value of the penning was, really bully the grammatical value was, and really efficaciously the matter communicated the ideas contained successful it” (p. 223). These questions were answered connected scales ranging from 0 to 100. The replicators asked the aforesaid questions to their participants.
You mightiness expect these judgments to beryllium correlated. For example, if personification says they liked the task, you mightiness besides expect them to opportunity that it was interesting.
In the replication, this was (super) true. Controlling for experimental condition, the partial relationship betwixt liking and liking was, rather sensibly, adjacent to cleanable [7]:

But successful the original data, this narration was not only imperfect; it was not location astatine all. Participants who said they liked the task much did not say that they recovered the task to beryllium much interesting:

In total, location are 5 subjective measures. In the replication, the (partial) correlations among these 5 measures scope from +.63 to +.92. They are each ample and very highly important (ps < 0.0000024). In the original data, these correlations scope from -.29 to +.18, and nary of them are some affirmative and significant. This is very strange.
The problem is not constricted to these subjective measures. Consider the truth that group did 3 very akin proofreading tasks, each pinch 100 mistakes. Surely, we’d expect group who do amended connected 1 task to besides do amended connected another, astir identical task. That elemental truth should manifest successful highly ample correlations betwixt capacity connected 1 task and capacity connected another. And successful the replication information it does, arsenic the correlations scope from +.74 to +.90. But successful the original information it doesn’t, arsenic the correlations scope from +.03 to +.27.
Finally, see that participants were asked to study really galore minutes they spent connected each of the 3 tasks. Again, we’d expect those who said they spent much clip connected 1 task to beryllium much apt to opportunity they spent much clip connected another, astir identical task. And truthful we’d expect these variables to beryllium very highly correlated. Once again, wrong the replication information they were – the correlations ranged from +.79 to +.95 – and wrong the original information they were not – the correlations ranged from +.05 to +.17.
The correlations we person reviewed successful this conception are fundamentally conscionable sanity checks. Does liking correlate pinch interest? Does capacity correlate pinch performance? Does reported clip spent correlate pinch reported clip spent? Sane information walk these checks. Insane information do not. The replication information are sane. The original information are not.
Red Flag #4: No Rounding In Self-Reported Minutes
As you’ll callback from a infinitesimal ago, Ariely and Wertenbroch (2002) purportedly asked participants to “estimate really overmuch clip they had spent connected each of the 3 tasks” (p. 223). When group supply estimates for illustration this, they thin to round. They usually opportunity “20 minutes” aliases “30 minutes” alternatively of “17 minutes” aliases “32 minutes”. And, indeed, erstwhile the replicators asked group to study really galore minutes they spent connected each of the 3 tasks, 85% of them gave a information number:

This is what we’d expect humans to do.
But successful the original data, they did not do that. Only 11.7% of estimated minutes were round, accordant pinch the 10% you’d expect by chance alone:
This is not what we’d expect humans to do.
Conclusion
We are incapable to make a benign mentation for each of the anomalies presented here. The original findings are excessively ample and yet they do not replicate; location are duplicated observations; correlations that should beryllium very beardown are often non-existent; and values that should beryllium rounded are not rounded. Based connected this evidence, we judge the information for Study 2 of Ariely and Wertenbroch (2002) were severely tampered pinch aliases fabricated to nutrient the desired results.
In our adjacent post, we will stock analyses of the Study 1 information record that Hyndman received from [email protected]. That research is rather different. Our analyses are rather different. But our conclusions are rather similar.
Author Feedback
About 6 weeks ago, connected July 20th, 2026, we shared drafts of our posts pinch the original authors (Dan Ariely and Klaus Wertenbroch), the replication authors (Kyle Hyndman and Alberto Bisin), and the editor-in-chief of Psychological Science (Simine Vazire).
Klaus Wertenbroch sent america a consequence successful which he originates by thanking Hyndman and Bisin for having done the replication. He restates that he ne'er had entree to the information for immoderate of the studies. He distinguishes betwixt demand for precommitment, a uncovering that was replicated by Hyndman and Bisin and which is accordant pinch earlier activity by him and others, and the effectiveness of specified precommitments successful these circumstantial studies, which did not replicate. And he indicated that he has asked the editor to retract the paper.
Dan Ariely did not reply to immoderate of the 3 emails we sent him. But connected August 7th, he wrote connected LinkedIn (htm) and connected his individual website (htm): “. . . Recently, I was made alert that information underlying a 2002 insubstantial astir deadlines and procrastination that I co-authored contained superior anomalies. The documentary grounds I person astatine my disposal coming astir those experiments isn’t capable to reply the questions that person been raised, and much than 2 decades, and hundreds of experiments later, my representation is likewise insufficient. Moving forward, my work lies successful ensuring accuracy – successful updating the grounds connected these experiments and, on pinch my co-author, cooperating pinch the diary that first published our insubstantial to support their reviews and retraction processes.”
Neither LinkedIn nor Dan’s website allowed archive.org to prevention copies; truthful we surface recorded some pages (mp4).
Kyle Hyndman and Alberto Bisin asked america to see this statement: “As stated successful the posts, successful April 2006, we received 3 information files attached to an email sent from Dan Ariely’s MIT email account, pinch nary stated restrictions connected their use. In August 2023, we provided those files to Uri Simonsohn, Joe Simmons and Leif Nelson to get their master assessment. We did not participate successful Data Colada’s study aliases successful drafting the posts. Our independent replication relies connected recently collected information and stands connected its ain methodological findings. Questions concerning the provenance aliases integrity of the humanities files should beryllium addressed by Data Colada, Dan Ariely, and the institutions pinch due work for those questions.”
Simine Vazire indicated that she is only allowed to opportunity that Psychological Science is considering “best adjacent steps regarding the 2002 insubstantial successful accordance pinch COPE guidelines.”
Footnotes.
English (US) ·
Indonesian (ID) ·