Oct. 5, 2026

Statistical Sleuthing: Can you spot a problem?

Statistical Sleuthing: Can you spot a problem?

Can you spot problems in a scientific study just by reading the published paper? We put our statistical sleuthing skills to the test with a famous study on procrastination and deadlines that was retracted more than 20 years after it was published. But were there clues hidden in plain sight in the published results—and would we have noticed anything suspicious at the time? We put on our data-detective hats and find a few statistical clues of our own before turning to Data Colada’s investigation of the original data, where the story gets even stranger. Along the way, we hunt for missing participants, question suspiciously enormous effects, meet some unlikely data twins, and discover why sometimes the best statistical tool is simply asking: Do these data behave like real people?

Statistical topics

  • Statistical sleuthing
  • Effect sizes
  • Degrees of freedom
  • Replication
  • Missing data

Methodologic Morals

  • “A data detective's toolbox is surprisingly simple. Check the degrees of freedom, look for twins, question giant effects, and ask whether the data actually make human sense.”
  • “You don't have to be a statistician to catch a statistical criminal.”


References

Kristin and Regina’s online courses:

Demystifying Data: A Modern Approach to Statistical Understanding

Clinical Trials: Design, Strategy, and Analysis

Medical Statistics Certificate Program

Writing in the Sciences

Epidemiology and Clinical Research Graduate Certificate Program


Program that Kristin teaches in::

Epidemiology and Clinical Research Graduate Certificate Program


Kristin’s manuscript writing course:

https://instats.org/seminar/from-data-to-discussion-writing-a-scient

Find us on:

Kristin - LinkedIn & Twitter/X

Regina - LinkedIn & ReginaNuzzo.com


  • (00:00) - Introduction
  • (02:15) - The Procrastination Paper and Study Two
  • (08:51) - Suspiciously Tidy Results
  • (16:18) - Study One: Real Classroom Stakes
  • (20:06) - Missing Data and First Red Flags
  • (28:17) - How Data Colada Got the Data
  • (33:43) - When the Replication Failed
  • (35:36) - Red Flags: Data Twins and Sanity Checks
  • (47:21) - Study One Gets the Spreadsheet Treatment
  • (51:45) - The Grading Scheme Unravels
  • (57:51) - Manuscript Versions Reveal the Truth
  • (01:08:11) - Rating the Claim and Methodological Morals

00:00 - Introduction

02:15 - The Procrastination Paper and Study Two

08:51 - Suspiciously Tidy Results

16:18 - Study One: Real Classroom Stakes

20:06 - Missing Data and First Red Flags

28:17 - How Data Colada Got the Data

33:43 - When the Replication Failed

35:36 - Red Flags: Data Twins and Sanity Checks

47:21 - Study One Gets the Spreadsheet Treatment

51:45 - The Grading Scheme Unravels

57:51 - Manuscript Versions Reveal the Truth

01:08:11 - Rating the Claim and Methodological Morals

[Regina] (0:00 - 0:07)
The R-squared of 0.57 says that 57% of the variability in students' grades is explained by their grades.


[Kristin] (0:08 - 0:15)
I'm right. And then there's a fudge factor. How much coffee did you bring the teacher, right?


That's not in the grade sheet.


[Regina] (0:16 - 0:16)
Apparently.


[Kristin] (0:22 - 0:32)
Welcome to Normal Curves. This is a podcast for anyone who wants to learn about scientific studies and the statistics behind them. I'm Kristin Sainani.


I'm a professor at Stanford University.


[Regina] (0:33 - 0:38)
And I'm Regina Nuzzo. I'm a professor at Gallaudet University and part-time lecturer at Stanford.


[Kristin] (0:38 - 0:43)
We are not medical doctors. We are PhDs. So nothing in this podcast should be construed as medical advice.


[Regina] (0:43 - 0:53)
Also, this podcast is separate from our day jobs at Stanford and Gallaudet University. Kristin, today we're talking about a study on procrastination.


[Kristin] (0:53 - 0:56)
Oh, that is a topic that fascinates just about everyone.


[Regina] (0:56 - 1:13)
But the story is also about the paper about procrastination, not just procrastination. It's about a paper that was once influential, but then was retracted 20 years after publication.


[Kristin] (1:13 - 1:25)
Wow. And Regina, the fact that it took so long means it's had time to be widely cited and used in the classroom, meaning people have trusted the results, even though, as we're going to find out, it's not trustworthy.


[Regina] (1:27 - 1:54)
It's fascinating, actually, Kristin, because if you read the paper carefully, in some ways the story was just a little too neat, as we will see. Okay, but for the claim, I've already given away that this paper had been retracted, so it doesn't make sense to make the claim about what the paper concluded. That would be too easy.


So our claim today is this. Many people who make up or alter data are surprisingly not very good at it.


[Kristin] (1:55 - 2:03)
We have to get a little creative with these claims when we're not, like, actually talking about the topic, procrastination. We're talking about the scientific process. Exactly.


[Regina] (2:03 - 2:15)
All right, good. All right, so today we're going to be doing some statistical sleuthing, both by us and by a group that published a series of blog posts on their detective work findings on this paper.


[Kristin] (2:15 - 2:35)
And this group was Data Colada. I love them. They are great statistical detectives, and their blog is a fascinating read.


Forget, like, Agatha Christie. This is like a murder mystery novel, but with data instead of people. And they have been involved in some very high-profile investigations.


[Regina] (2:36 - 3:12)
Yes, Data Colada, I love them. Three amazing researchers are behind this. They started their blog in 2013, and they have just gotten better and better, honestly, over the years.


And we'll talk in a moment more about what they found, but first I want you and I to walk through this paper together and do our own sleuthing. All right, so the paper is called Procrastination, Deadlines and Performance, Self-Control by Precommitment. And the question is whether people can set their own deadlines well or whether you kind of need other people to give you a good deadline structure.


[Kristin] (3:12 - 3:34)
Oh, that is actually a great question. And I know that some college professors actually just have one deadline for all the homework, all the projects at the end of the term, rather than spacing things out. And of course, that doesn't mean that students can't turn things in early, but human nature then a lot of students just wait to the last minute.


So I don't give my students quite that much flexibility.


[Regina] (3:35 - 3:58)
That is exactly the idea. So the paper was published in Psychological Science back in 2002, and it was by two researchers, a behavioral economist named Dan Ariely, who was at the time at MIT, now at Duke, and Klaus Wertenbroch, a professor of marketing at INSEAD, which is a business school in France.


[Kristin] (3:59 - 4:42)
Oh, you know, interesting, Regina. Dan Ariely has actually been on the Data Colada radar before this paper. There is another famous case that involves a researcher named Francesca Gino, and she's been in the news lately because she just lost her tenured position at Harvard, and she was making like a million dollars a year or something.


And she lost her tenured position because of problems that Data Colada uncovered. It's a whole saga, but it's relevant to today's episode because one of Francesca Gino's retracted papers from 2012 happened to be co-authored with Dan Ariely. So the case we're talking about today is actually tied to the case of Francesca Gino, which we are going to talk about in the next upcoming episode.


[Regina] (4:43 - 5:12)
That is a good story. And you're right, it is not coincidence that all of this is coming together in front of Data Colada. So that one has lots of juicy twists and turns.


We're just going to tease it today. Today we're going to focus on this procrastination paper. And the Data Colada blog post that revealed these problems were recent, actually.


So they published at the end of August, beginning of September of this year. And then the article itself was retracted shortly after that in September.


[Kristin] (5:13 - 5:20)
Well, Regina, considering it's September 29th when we're taping this, I feel like we are covering breaking news here at Normal Curves.


[Regina] (5:20 - 6:07)
We are. It's very exciting. We're so timely.


Okay. So the paper reported primarily on two small studies. I am going to start with study two because that's the order that the statistical sleuthing happens.


So I want to preserve, you know, that narrative suspense. All right. Study two.


They recruited 60 students at MIT to do a proofreading task. Each participant had to proofread three papers and the researchers had set it up ahead of time so that each paper had exactly 100 mistakes, spelling errors or grammar errors. And participants got paid 10 cents for each mistake that they found.


But they also lost a dollar every day that they were late after their deadlines.


[Kristin] (6:08 - 6:25)
First of all, Regina, I'm imagining that they were doing this offline, like 2002. Is this pen and paper? Red pen?


Oh, I love it. But I'm doing the math here. And 100 mistakes per paper, 10 cents each.


They can make up to $30 if they find all of the mistakes and if they don't have any late penalties.


[Regina] (6:25 - 6:57)
Right. And there were three different deadline groups. Each participant was randomly assigned to just one of them.


So group one, they had to turn in one proofreading report every week for three weeks. Group two had just one deadline at the end of the three weeks, turn it in wherever. Group three got to choose their own three deadlines whenever they wanted within those three weeks.


They could make them all in the last day if they wanted, but they had to pre-commit in the beginning of the study and they couldn't change their deadlines later.


[Kristin] (6:58 - 7:07)
All right. So we've got like the regular deadlines that the teacher imposed. We've got the last day deadline, and then we've got the choose your own adventure deadline group.


Yes.


[Regina] (7:08 - 7:24)
Exactly. Nice summary. So the researchers hypothesized that having someone else impose a regular deadline schedule, that would lead to better outcomes than if you let people pick deadlines on their own or if you just, you know, give them one lax deadline at the end.


[Kristin] (7:24 - 7:29)
And what was the primary outcome? Was it just how much money they made in the end in this task?


[Regina] (7:29 - 7:43)
Good question. So the money is determined by two things, right? The number of mistakes that they caught in each paper and also how late they were.


That was the penalty. So they kept those two performance metrics separate.


[Kristin] (7:43 - 7:43)
Ah, okay.


[Regina] (7:44 - 8:04)
Right. Because maybe they're independent. So those were the big outcomes, but they also looked at two other more minor things, subjective experience for each participant.


Like how much did they enjoy the proofreading? And how much time each participant reported spending proofreading each paper?


[Kristin] (8:05 - 8:18)
Oh, because maybe if you wait to the last minute, if you procrastinate, then maybe it's not as pleasant because you're under a lot of pressure and maybe you have to rush. So those outcomes might indeed be related to deadlines. Yeah.


[Regina] (8:18 - 8:40)
Yeah. You can see that.


So here's what I did when I went through the paper, Kristin. I decided to treat it just like fresh, just blindly, pretending like I didn't know it had been retracted. I didn't know anything about the data manipulation accusations because I just wanted to see if I was just reading the published paper, what, if anything, would make me pause and say, hmm.


[Kristin] (8:41 - 8:51)
Right. So you're going to start with, if I was just the reviewer for this paper, what I have caught that there were some suspicious things here? Yes.


Okay. There you go. What did they find in this paper?


Okay.


[Regina] (8:51 - 9:06)
Let's talk about performance first. That was the number of mistakes they caught, how late they were. And here's the frustrating thing, Kristin.


They never report the actual values for the individual group means or the standard deviations in the text of the paper or a table.


[Kristin] (9:06 - 9:21)
Right. So let me guess. Did they just give a bar chart that shows the means and the standard deviations in a picture?


And we've talked a lot about why we hate that on this podcast, and it seems to be very common in old papers from the turn of the century. We've seen a lot of these.


[Regina] (9:22 - 9:28)
Very old. So I had to use our boyfriend, Graph2Table, to extract the data.


[Kristin] (9:29 - 9:40)
I'm going to put the plug in for Graph2Table and remind our listeners that you can use Graph2Table yourselves and with a discount, normalcurves20, and that helps support the podcast when you use that and you get a discount.


[Regina] (9:41 - 10:15)
And it is such a fun app. Okay. So when I looked at the numbers, Kristin, things were just pretty tidy.


Let's say that. Just a little clean. So let's talk about how many mistakes they caught.


People did the best and they caught the most mistakes when they were in that strict, regularly spaced deadlines group. They had an average of 136 mistakes caught. So that's, if you can do the math, that's just about half of all of the available mistakes.


[Kristin] (10:15 - 10:25)
Right. So 136. Remember, there were 300 possible.


So these students didn't do the best job ever. This was pre-AI, and there were a lot of little things to spot in there.


[Regina] (10:26 - 10:41)
Okay. So they did the worst when it was that last day deadline, when there was basically no deadline. They found an average of only 71 mistakes.


And that self-imposed group, the choose your own deadlines group, was in the middle with an average of 107.


[Kristin] (10:41 - 10:58)
Wow. That lined up really, really nicely, perfectly in the order that probably the authors predicted and nice and evenly spaced. I just want to note, though, it's not necessarily suspicious that things worked out as you hoped.


It's just a little convenient.


[Regina] (10:59 - 11:32)
It is very neat, small error bars, p-value less than 0.01. So everything was nice there. And then we see the same pattern with the lateness, right? How many days late they were.


Beautiful pattern. People were more on time in the regularly spaced deadline, an average of only 3.5 days late. The worst was that last day deadline, an average of about 12.5 days late. And that choose your own deadline group was in the middle, an average of eight days late.


[Kristin] (11:32 - 11:43)
That's actually a lot. And people weren't necessarily taking this study as seriously as they would like a class, I think. But again, evenly spaced and exactly in the order that the authors, I think, were hoping.


[Regina] (11:43 - 12:22)
Mm-hmm. Okay. Now let's go on to how people rated their subjective experience doing these proofreading tasks.


[Kristin]
And what was the scale for this, Regina?


[Regina]
Yeah. This one was kind of interesting.


So they asked them five questions. How much did you like the task? How interesting was the task?


How good was the quality of writing? How good was the grammatical quality? And how effectively did the text communicate the ideas in it?


And they answered each of the five questions on a one to 100 point scale. And then the researchers added them up and then took the average.


[Kristin] (12:22 - 12:36)
That's a bit of a weird variable because it's combining like how much they liked doing it with how good they thought the papers were. So those aren't really the same construct that we're measuring. That's weird that they put those together.


But okay.


[Regina] (12:36 - 13:05)
Okay. Okay. I'm not sure what that's about, but okay.


I'm not sure either. I kind of chalked that one up to silly, sloppy research, right? Why would your perception of the quality of the writing be related to your deadline schedule?


I am not really sure. But let's just go with it. Okay.


So here we have the means. They actually gave us the exact group means in the text. And again, it's in this beautiful order.


The last day deadline group had the best experience with an average score of 38.


[Kristin] (13:06 - 13:07)
So kind of opposite direction.


[Regina] (13:09 - 13:36)
They enjoyed it better when they got to put everything off until the last minute, even though they were also later when they turned it in and they caught fewer errors. Okay. So the group with the worst subjective experience was the regular deadlines, right?


One every week. They had an average of 22. And then the choose your own deadline, self-imposed deadline group was in the middle at 28.


[Kristin] (13:36 - 13:56)
And this is out of 100. So people didn't really enjoy this task very much is what you're telling me. And maybe it's like eating your vegetables.


Like if you do it the way you're supposed to on time, meet these regularly spaced deadline, it's good for you. You're going to do a better job, but maybe it doesn't feel as good. Is that possible?


Would you rather have cake?


[Regina] (13:57 - 14:17)
I guess this is possible. I'm not really sure what they were after by asking this question in the first place. Like, does this really tell us anything?


But no matter, their result was statistically significant. And then the last one is how much time they reported spending proofreading each paper.


[Kristin] (14:17 - 14:29)
And let me guess. The effect is perfectly ordered and nicely spaced out between the three groups. And this one, I bet the regular deadline group spent the most time and the last day deadline group spent the least time.


[Regina] (14:29 - 14:51)
You are psychic. The regular deadlines spent the most time on each paper, an average of 84 minutes. The last day deadline spent the least, an average of 51 minutes.


So like a difference of half an hour in there on each paper. And the choose your own deadline group spent an average of 70 minutes on each paper.


[Kristin] (14:51 - 14:59)
So again, the same kind of beautiful pattern. Although this one makes a lot of sense to me, because if you waited till the absolute last minute, you just have less time to get it done.


[Regina] (15:00 - 15:01)
This is true.


[Kristin] (15:01 - 15:02)
Could be real.


[Regina] (15:02 - 15:13)
So Kristin, here they finally gave us an actual F statistic for this hypothesis test. F, of course, is referring to the F distribution, not profanity.


[Kristin] (15:14 - 15:18)
Oh, and not a grade either, which would be relevant here.


[Regina] (15:18 - 15:28)
Yeah, right. All right. F distribution.


So the F statistic for this test on the amount of time that the students spent, it was 46.


[Kristin] (15:28 - 15:34)
46. Wow. That is a huge F statistic.


And they only gave an F statistic for this one outcome, not for the others?


[Regina] (15:35 - 15:36)
Yep.


[Kristin] (15:36 - 15:38)
Oh, that's weird. OK, but that is huge.


[Regina] (15:38 - 16:01)
Yep. The p-value, I calculated it out, is actually one with 11 zeros in front of it. Wow.


One trillionth. So the researchers report it as just 0.001, which is hiding how surprising this effect really is. Like, shockingly surprising, especially with only 20 subjects in each group.


[Kristin] (16:01 - 16:07)
That's pretty small sample size. Right. With a small sample size to get a p-value that small, you actually end up with a pretty big effect size.


[Regina] (16:07 - 16:17)
Yes, exactly. Hold on to that thought. We will revisit that.


So not impossible, but a little too tidy.


[Kristin] (16:18 - 17:12)
So I can see where you're going here, Regina. We need more information to know if there is anything weird in the data. And you know, it could just be that the deadlines do really have this huge effect and it does come out this neat.


It could be. But you know, your spidey sense is up a little bit that maybe it's a little too good to be true. Right.


The effect is really big and it's nicely ordered for every single outcome. And the one outcome that really gets me suspicious is that weird subjective experience outcome because I don't even know what that outcome is measuring. It's kind of measuring noise because it's combining how much I like something with what I thought about the writing.


And for that one to come out in this perfect pattern makes me a little even more skeptical because that's just a weird variable that I don't even know what it means. So Regina, when we see something like this, the next question we're going to ask is, can you hand me the data? Can I see the data so that I can look and see what's going on?


Is that what happened next here?


[Regina] (17:12 - 18:10)
Mm-hmm. The Data Colada folks, Kristin, actually got the individual level data. They got a spreadsheet of data and they found some really bizarre things in here.


But before we get to what the Data Colada folks found out with their spreadsheet, let's see what we can find out in the rest of the paper just looking at the published results. So now we're going to talk about study one. And here, Kristin, they had two sections of a course at MIT, an executive master's course.


And both sections had, as part of their grade, they had to write three short papers. And section one, 48 students, the instructor gave them deadlines. They were strict, regularly spaced deadlines throughout the term.


Section two was a special section, 51 students. They could choose their own deadlines. Again, they had to set them right away.


They couldn't change them, but that was the fun one.


[Kristin] (18:11 - 18:44)
So we're now down from three groups to two groups. So we lost the group that just got the end of the quarter deadline, right? It's either choose your own adventure or set in stone, spaced out deadlines.


And Regina, technically though, the students were not randomly assigned to these conditions, right? They just signed up for a section that fit their schedule, and then the condition got assigned to that section. So we do need to worry a little bit about if there was any systematic difference between who signed up for each section.


Because maybe one was at 8.30 in the morning, and one was at 1 p.m., and we get all the slackers in the 1 p.m. section, right?


[Regina] (18:45 - 18:55)
Yeah. Yeah. The paper did not say, pointedly did not say that they were randomized, but they did say that the two sections were about equal in their academic performance.


Okay. Okay.


[Kristin] (18:56 - 18:58)
So they did some kind of check to make sure they were somewhat balanced. Yeah.


[Regina] (18:59 - 19:08)
Yeah. Good. I think it was a good check.


So penalty for turning in your paper late in both of the classes was 1% deducted for each late day.


[Kristin] (19:08 - 19:32)
Wow. That is a very generous late penalty. Usually teachers take off like 10% per day.


Okay. Very nice. Mm-hmm.


And Regina, this is a lot like study two, the proofreading task, except now it's actually real life. There are stakes, right? Before it was a couple of bucks, but now their actual grade depends on this.


So that's good. There's some real stakes here where maybe people will take it very seriously.


[Regina] (19:33 - 20:06)
Mm-hmm. All right. So let me start off by saying when I read the results for study two, I was annoyed right off the bat because they did some things that made it look like they had mysteriously missing data.


And just for the choose-your-own-adventure section, the choose-your-own-deadline section, they had 51 students in the section, but they reported degrees of freedom in some t-tests that suggested they used data from only 45 students, not 51.


[Kristin] (20:06 - 20:41)
Okay. So degrees of freedom are one of these great little tricks in the statistical detectives toolkit. And actually, we talked about degrees of freedom and this trick back in the episode we did on ultramarathons and vitamin D.


But the trick is even a little easier to implement here because we are just talking about a simple t-test and you can tell pretty easily from the degrees of freedom in a t-test what the sample size was. And so this is a great way of making sure that the sample size the authors say they are using is actually what they analyzed because often those don't match.


[Regina] (20:41 - 21:10)
Exactly. So here they were just a simple one-sample t-test on the deadlines that each student chose, which means degrees of freedom is n minus 1, and they reported 44 degrees of freedom. So that means they were using data from only 45 students.


So they were basically saying they only had deadline information from 45 students in a section of 51 students who had to choose their own deadlines.


[Kristin] (21:10 - 21:13)
Which begs the question, where did the other six students go?


[Regina] (21:14 - 21:22)
Where did they go? Right. So that just sets the stage for me being a little annoyed, a little suspicious because they didn't mention.


[Kristin] (21:22 - 21:33)
So, you know, missing data happens sometimes. Maybe somebody forgot to record that. But if there's six people with missing data, they ought to tell you somewhere in the paper, we have missing data because, and I'm assuming they did not here.


[Regina] (21:34 - 21:45)
They did not. OK, so the most important outcome that they looked at in this study was the student's overall grade in the entire course.


[Kristin] (21:46 - 22:00)
So wait, Regina, why didn't they look at like the average paper grade since the deadlines were on the paper? Do we know how much those papers counted to the final grade? Was the final grade entirely based on those papers?


Like why look at the final grade?


[Regina] (22:00 - 22:25)
Why look at the final grade? Excellent question. I have no idea.


I had to reread the paper over and over to try to figure it out. It just said grades and it was pretty clear they were not talking about paper grades. They did not say what the grading scheme was.


So who knows? Maybe the course grades are, you know, 90 percent based on the paper grade, but they don't say at all.


[Kristin] (22:25 - 22:32)
Yeah, that's something they ought to tell you because the natural outcome to me there would be the grades on the papers, the things that the deadlines were for. But OK.


[Regina] (22:32 - 22:46)
Mm hmm. Oddly enough, they also looked at the students' project grade, which they didn't describe at all other than to say it was unrelated to the papers and it was due on the last day of class for everyone.


[Kristin] (22:46 - 22:55)
OK. So wait a minute. Why would you look at the project grade if everyone had the same deadline for that and it wasn't related to the condition?


Like why is that relevant?


[Regina] (22:55 - 23:05)
Also excellent question. The author said, oh, deadlines can have a secondary effect on other aspects of performance. And so that's why we're looking at the project grade.


[Kristin] (23:05 - 23:16)
Oh, OK. Maybe if you wait until the last day to do all three papers and you also wait to the last day to do the project, then that negatively affects your project. Right.


Right.


[Regina] (23:16 - 23:38)
OK. Secondary effects. OK.


Secondary effects. OK. So for both of these outcomes, they found a nice clear pattern.


The overall average course grade in that no-choice section was about 89. And in the choose-your-own-deadline section, it was 86, p-value 0.003. Oh, so the effect size here is not quite as huge as we were seeing before.


[Kristin] (23:38 - 23:51)
This is the difference between an 86 would be like a B and 89 is a B plus. So I mean, my Stanford students do sometimes argue over differences that small, but really in the grand scheme of things, B or B plus, not a huge difference.


[Regina] (23:51 - 24:10)
Not a huge difference, but the project scores, Kristin, same pattern, but stronger. The no-choice section average was 86 compared to 77 in the choose-your-own-deadline section. P-value was 0.00007. So that is a pretty big difference.


[Kristin] (24:10 - 24:44)
That's the difference between a B and like a C plus. And that's weird because I certainly would have expected the deadlines to have a bigger impact on the papers and the overall grade than to an unrelated project. This does make me think, Regina, maybe this is a real difference, but this makes me wonder about any biases in choosing the sections, right?


So people self-selected to these sections. So again, could it be that we get all the slackers in Section 2 and all the good students in Section 1 and that's why the project scores differ and it has nothing to do with the deadlines?


[Regina] (24:44 - 24:59)
That was my question, my immediate thought too. Like that is such a big difference that there's a section effect going on here. So the authors didn't say, and it's also not clear, Kristin, whether the sections were taught by different instructors.


[Kristin] (24:59 - 25:03)
Which is kind of important because one instructor might grade differently than the other.


[Regina] (25:04 - 25:11)
Exactly. Oh, and remember the missing data that we were talking about, like missing the six students?


[Kristin] (25:11 - 25:11)
The six, yeah.


[Regina] (25:11 - 25:39)
Are you right? Okay, it happened again here, but in a different way. So again, I could see from the degrees of freedom in the analysis of the overall course grade that they included all the students.


They had a sample size of 99 degrees of freedom, 97 and minus two. But in the analysis of the project grade, they were missing information from two students. They only reported a 97, not 99.


[Kristin] (25:39 - 25:48)
So we've lost two students, don't have a final project grade. Maybe they didn't submit it and they just took the zero and still pass the course. Could be.


[Regina] (25:48 - 26:02)
Could be, but they didn't explain it again. The authors didn't explain it. So on top of the missing information from before, so we're missing six students' deadlines and now we're missing two students' project grades, you know.


[Kristin] (26:02 - 26:12)
Right. And again, missing data happen, but if you're not upfront with it, if it's hidden, that's when we consider that kind of another statistical cockroach.


[Regina] (26:12 - 26:56)
Another statistical cockroach, exactly. Okay. And lastly, they did something else that was not a main analysis, but it ended up being really important later.


So remember this. Okay. They looked at the choose your own deadline section and they extracted just the subgroup of the students who chose deadlines that were evenly spaced or almost evenly spaced, right?


So you could choose any kind of deadline and some of them chose ones that were evenly spaced and others chose different ones. And they just chose the evenly spaced ones and then they compared that subgroup to everyone in the regular no choice section where they also had regularly spaced deadlines imposed on them.


[Kristin] (26:56 - 27:16)
So I guess, Regina, here they're trying to separate out, is the effect they're seeing, is that due to the fact that somebody else imposing the deadlines, maybe you're going to adhere to them more because you're scared than if it's just a self-imposed deadline versus is it about evenly spacing out the deadlines rather than waiting till the last minute? They're trying to tease out those two factors?


[Regina] (27:17 - 27:56)
Exactly. Exactly. That is perfect. And so they gave the results for this analysis and they said, hey, those two groups were not different, which means they're basically saying students who chose regularly spaced deadlines did about as well as those who had them imposed on them. And the ones that did the worst of all were the ones who chose deadlines that were not regularly spaced.


Now, Kristin, it was weird because they didn't actually give any results at all, like no p-value, no means, nothing, other than to say not significantly different and that the effect size was reduced by 59 percent.


[Kristin] (27:56 - 28:06)
Well, I'd really like to see some real numbers there because not significant doesn't necessarily mean that the groups aren't different. But we didn't get any other numbers.


[Regina] (28:06 - 28:17)
We did not. But this is why we have Data Collada folks because, Kristin, they managed to get their hands on an actual spreadsheet with the real data from the paper.


[Kristin] (28:17 - 28:57)
Regina, this is interesting to me because you're telling me that now 24 years after this paper was published, somebody dug up the data and is scrutinizing it now. Like, how did that happen? Usually people are just saying, well, the dog ate my data.


That's common, right? People don't share their data 24 years later. Or it was on my laptop that I retired five years ago.


I've gotten that one. I'll have you know, though, somebody a couple of years ago asked me for my Ph.D. thesis data from like 20 something years before. And I had it.


I was able to find it. So, you know, good file management, as we talked about in the last episode.


[Regina] (28:57 - 29:14)
Yes, that is one of your superpowers among many. But this is an excellent question. Why are we revisiting this now?


And there's an interesting story behind that. But that's a teaser. I will tell you about it after our short break.


[Kristin] (29:34 - 29:57)
Welcome back to Normal Curves. Today we're talking about a 2002 paper on procrastination. And Regina, you have been just teasing away to keep people on the hook to hear the story behind this.


And now we're going to finally hear about how the Data Colada folks started investigating this paper and some of the details of the actual data set which we have.


[Regina] (29:58 - 30:14)
Publicly available now. All right. So this story with Data Colada actually starts three years ago when a researcher named Kyle Hindman sent the Data Colada folks what appear to be the original data files that go with this paper.


[Kristin] (30:14 - 30:29)
OK, so wait, how did he get the data? As most 2002 papers didn't have publicly available data just because the internet thing was just getting going. So how did he happen to have the data?


He wasn't a co-author on the paper, so he just got it from the authors at some point?


[Regina] (30:30 - 31:12)
Well, that goes back to 17 years earlier, back to 2006. This is an epic saga here. So Kyle at the time was a PhD student in economics at New York University and he wanted the data for some reason.


It's not entirely clear, but maybe he wanted to do a little computational, you know, reproducibility, just recreate some of the results. And he did manage to get the data by emailing the papers to authors. And after some back and forth, eventually he was sent what appears to be the paper's original data files in an email from Dan Ariely's email address at MIT.


[Kristin] (31:13 - 31:29)
Well, you know, again, file management. I'm impressed that he got the data in 2006 and he had it available, held on to it for 17 years and was able to find it. But how did he happen to revive this data set 17 years later and give it to the Data Colada folks?


[Regina] (31:30 - 31:43)
So we don't know exactly why, but we do know this. He sent the data to the Data Colada folks one week after Francesca Gino sued the Data Colada folks for $25 million. You remember that?


[Kristin] (31:45 - 32:12)
Ah, so again, we're going to be talking about Francesca Gino in the next episode. That's a really great story as well. But I'm just going to guess, hypothesize, maybe when the whole thing with Francesca Gino came out, which included a paper coauthored by Dan Ariely, maybe Kyle then said, hey, I have that data set from Dan Ariely from way back.


And he decided to go back and look at it and see if there was anything suspicious. Exactly.


[Regina] (32:13 - 32:22)
I mean, I feel like the picture is starting to make sense. The thing is, the Data Colada folks were a little distracted at the time with the whole $25 million lawsuit thing.


[Kristin] (32:23 - 32:27)
Right. They were in the middle of being sued. So maybe they didn't have time to deal with this.


[Regina] (32:28 - 33:05)
Right. So at the time, three years ago, they actually said in a blog post, they managed to do a quick analysis of the data. It missed everything else.


They talked with Kyle and his colleague, Alberto Bison, and Kyle and Alberto said that they were going to conduct a replication of this procrastination study, which means they would try to reproduce it with their own experiment, right, using the same methodology and collect their own data and reproduce it that way. And since the Data Colada folks were already a little busy, they just left it at that until just recently.


[Kristin] (33:06 - 33:25)
Oh, so maybe what happened is Kyle was planning to do a replication at some point, and he never got to it until 17 years later. But when he finally got to it, he was like, okay, now I'm going to go look at the data set so that I can try to replicate it exactly. And at that point, maybe noticed some funny things and sent it to the Data Colada folks.


We don't know. I'm just guessing.


[Regina] (33:25 - 33:42)
That is highly likely, and we will see why in a moment, because Data Colada folks picked it back up because they saw that Kyle and Alberto were publishing their replication study. And it hadn't...


[Kristin]
Oh, and did it replicate?


[Regina]
It did not replicate.


[Kristin] (33:43 - 33:43)
Oh. Okay.


[Regina] (33:44 - 33:49)
I know. But this is not about their replication paper, but we will hear from it. It informs something. We will hear about it in a moment.


[Kristin] (33:49 - 34:00)
Well, I mean, the fact that they couldn't replicate, it doesn't mean that the original findings weren't real or true. Because you can have findings that don't replicate.


But it does make us wonder about the original study if it's not able to be replicated.


[Regina] (34:01 - 34:56)
It does. And especially when you hear about this very interesting footnote in the replication paper. So Kyle and Alberto said they had done their own data analysis of the data in those files that Kyle got back in 2006.


And they said in the footnote that they shared that analysis, their independent analysis with Dan Ariely and asked for permission to include a summary of it in this replication paper. But Dan said no.


[Kristin]
Oh, that is interesting.


Why would you say no to that?


[Regina]
The footnote said Dan had argued, among other things, that the files might not actually be the original data. The file that he gave Kyle might not actually be the original data, but he also did not give Kyle and Alberto any other data.


So they were stuck.


[Kristin] (34:56 - 35:13)
Why would you come back and say those aren't the data? Maybe he knew that if people started looking at the data, they would find weird things. And therefore, he was trying to pretend that that was just, what, the decoy data?


Right. But it does smack of somebody trying to cover their tracks.


[Regina] (35:14 - 36:18)
OK. And then that's when the data collateral folks said, oh, yeah, how about we go back and look at this now? And that's when they picked it back up.


So they started with study two. That was the proofreading study, because that is the only study that Kyle and Alberto tried to replicate. And the data collateral folks dove into study two, and they found four red flags.


Now, remember, they have the original data, so they're able to go even deeper than you and I did just walking through it just now. Red flag number one. The blog post said the effects are extremely implausibly large.


So remember we talked about that finding with how many mistakes the participants in each group caught? This is a proofreading one. And when you look at the regular deadlines group versus the last day deadlines group, that effect has a Cohen's d of 2.5. They were able to calculate that because they had the original data.


[Kristin] (36:18 - 37:10)
Cohen's d of 2.5. That is actually huge. And we had some numbers like, what was it, like 70 mistakes versus 107 versus 130 something. So you can kind of see that that's a large effect from that difference, but it's a little hard to contextualize that. The Cohen's d puts that in context because now we're in standard deviation units.


A 2.5 standard deviation difference between groups is humongous. In fact, I was involved in scrutinizing some meta-analyses in sports science. We published a paper on this and we found that most of the effect sizes that these meta-analyses reported that were 3.0 or larger were actually errors because effect sizes like that don't generally happen in nature. Unless it's something really big like, you know, if I take aspirin, it helps headaches. That might have an effect size of 3. But deadlines, you would expect the effect to be more subtle.


[Regina] (37:10 - 37:39)
You would. Exactly. 2.5. And I love how the data colada folks put it into context. They said, hey, this is much larger than the obvious effects we notice in everyday life. They pointed out that this is larger even than the effect of gender on height. So men are taller on average than women with a Cohen's d of only 1.8. And now here we're talking about 2.5 just for some deadline structures.


[Kristin] (37:41 - 37:47)
It seems like if it was that big, we'd all be always regularly spacing deadlines for our students. It would be so obvious.


[Regina] (37:48 - 37:59)
Okay. Red flag number two. Ready?


When you look at the results, people in that last day deadlines group had twins.


[Kristin] (38:00 - 38:22)
Oh, data twins. Okay. Let me explain that. What we mean by that is two rows of data have identical outcomes.


And in this case, that's not easy to be coincidental because we're talking about the number of mistakes on three different papers. And it can range from one to 100. So if you got somebody with the exact same numbers on that, that would be a little suspicious.


[Regina] (38:23 - 39:07)
That is hugely suspicious. So Kristin, it's not just that these data twins found identical numbers of mistakes overall, over all three papers, but for each individual paper. And okay, you'll love this.


For example, subject one caught nine errors on the first paper, 15 on the second, 26 on the third. Subject 11 caught the exact same numbers for all three papers. Subject two also had a twin.


Who was? Subject 12. Subject three also had a twin.


Guess what their subject number was?


[Kristin]
I'm thinking 13. Am I psychic?


[Regina]
You are indeed psychic. There were nine twins like this, all with ID numbers that were separated by 10.


[Kristin] (39:07 - 39:52)
Okay, Regina, this reminds me of our discussion in the last episode with Kate Laskowski, who found also duplicate observations in her spider data from Jonathan Pruitt. Same thing. There was a lot of obvious cut and paste.


You know, you'd think people would be more clever when faking data, but I guess people are lazy, or maybe they just think no one's ever going to see the data. But I think this is kind of a smoking gun, right? How is it that you would end up with the exact same numbers across two different participants and their ID numbers are not 1 and 1, which could be an accidental cut and paste, but are 1 and 11?


It seems very, very deliberate. I mean, aren't we done now? Like, can't we all go home and say these data were faked?


[Regina] (39:53 - 40:01)
I guess we could go home, but there are so many more fun goodies to uncover that I think you'll enjoy the rest of the journey.


[Kristin] (40:01 - 40:08)
No, I love statistical sleuthing, so bring it on. But wow, that is just, yeah, that's really a smoking gun.


[Regina] (40:08 - 40:28)
This is pretty bad. So the Data Colada people wrote, quote, the existence of so many of these twins and all of them in only one condition, they were all under the last day deadline condition, is inconsistent with these data being real.


[Kristin] (40:30 - 40:41)
I'm going to translate that into even more plain English. It is consistent with these data being fake or manipulated. Manipulated, yeah.


I hope we are not going to get sued for saying that, Regina.


[Regina] (40:42 - 40:45)
I hope not, because we have no money, though.


[Kristin] (40:46 - 40:56)
Right. There's nothing to sue us for. All right.


But Regina, that's only two red flags, if I'm counting correctly. So what are red flags 3 and 4?


[Regina] (40:56 - 41:07)
Right. Red flag number 3, this is what they called failed sanity checks. OK.


So remember those subjective experience questions? Right.


[Kristin] (41:07 - 41:15)
There were five questions like how much you liked the task, how interested you were in the task, but also how you rated the quality of the writing.


[Regina] (41:16 - 41:29)
Right. So within it, there are some questions that are kind of similar, right? How much you like the task and how interesting you find it.


You would think that people who say they like a task would also tend to think it's interesting, right?


[Kristin] (41:29 - 41:37)
Oh, yeah. I would expect those two to go together, even if those two did not relate to, like, what you think about the writing. But those two should go together.


[Regina] (41:37 - 41:47)
And the data, they were not correlated.


[Kristin]
Not at all?


[Regina]
The correlation coefficient was negative 0.1.


[Kristin] (41:47 - 42:11)
Wow. So they were actually slightly inversely correlated. So, like, if you liked the task, then you tended to find it a little less interesting, which is not a very plausible combination. Just to remind people, correlation coefficient of 0 would mean no correlation.


Negative is inversely correlated. Positive is positively correlated. And the closer you get to 1, the stronger the correlation.


0.1 would be considered minimal or very little correlation.


[Regina] (42:12 - 42:36)
But when you look at that replication study, when they ask the exact same questions, in the replication study, they were highly correlated. So that is a good sanity check, right? Yes.


Also, you would think that people who are good at proofreading paper number 1 are probably going to be pretty good at proofreading paper number 2 and paper number 3, right? Right.


[Kristin] (42:36 - 42:53)
You would expect people to be correlated within themselves. You would expect within-person correlation, meaning, yeah, the people who are good at proofreading on one paper are also good at proofreading on another paper. And the people who are bad at it are also bad across all papers.


That's definitely what we would expect to see here, because those shouldn't be independent.


[Regina] (42:53 - 43:19)
They should not be independent. So in the replication study, those correlations were strong. They ranged from 0.7 to 0.9. But in the procrastination paper data, those correlations were not. They ranged from 0.03 to 0.27. Which is basically almost nothing. Almost nothing. The Data Colada people wrote, the people in the groups didn't behave like people.


[Kristin] (43:19 - 43:25)
That's a great quote. Yeah. People are hardly correlated with themselves.


That does not seem real.


[Regina] (43:25 - 43:39)
Right. Data Colada called this sanity checks, right? They said, sane data pass these checks of people being correlated with themselves.


Insane data do not. The replication data are sane. The original data are not.


[Kristin] (43:39 - 43:46)
That's some great writing right there. I love that guess. Insane data.


We have insane data. All right. What was red flag number four, Regina?


[Regina] (43:46 - 44:00)
Number four. So, Kristin, how about if I ask you to recall a task and estimate how long you spent on it? How long did it take you to drive into Stanford today?


[Kristin] (44:00 - 44:01)
I think it was about 30 minutes.


[Regina] (44:01 - 44:06)
Now, notice that you said 30 minutes, not 31, not 27.


[Kristin] (44:07 - 44:25)
Right. I rounded because I didn't start my watch, my Garmin, when I started my commute, like I would for a run. So I don't have a precise time and I'm just recalling and estimating.


So most people do this, right? We round when we report estimated times like this. We don't get precise values. So does this relate to the paper in some way?


[Regina] (44:25 - 44:37)
This does because, remember, the researchers asked the participants to recall how long they spent proofreading each paper and they asked them after the fact.


[Kristin] (44:38 - 44:50)
Oh. So I'm definitely expecting round numbers. People aren't going to say I spent, you know, 72 minutes.


They might say 60 minutes or 70 minutes or 75 minutes, but not something too precise.


[Regina] (44:51 - 45:05)
Right. Right. Remember, we didn't have the original data, so we couldn't look at that, at what each person was saying.


But once you have the original data, you see that only 12% of the numbers that people reported were round numbers.


[Kristin] (45:05 - 45:05)
Wow.


[Regina] (45:06 - 45:11)
But in the replication study, which is kind of our sanity check, 85% of the numbers were round.


[Kristin] (45:11 - 45:44)
That is actually a great trick. And this is a trick I've seen Data Colada use in their sleuthing before, because it's about how humans work. Right.


When you ask people to give numbers, they often round to the nearest 10 or the nearest 100. But when people are making up data, they sometimes forget how people work. Right.


And they don't think about that. And so they make it up in a way that does not look human. And that's especially true if they use some kind of random number generator to generate the data, which is something Data Colada has found before.


So this is just a cool statistical sleuthing trick to know.


[Regina] (45:45 - 46:04)
Kristin, I'm a little worried right now that we're giving away all the sleuthing tricks, right? And this will help the data cheaters, right? I'm thinking of like the murderers who are sitting at home watching CSI and making notes on how to hide the bodies better or something.


Are we helping data cheaters now?


[Kristin] (46:04 - 46:33)
Somehow I don't think the serial murderers are sitting home watching CSI and taking notes and then using that to escape the cops. First of all, CSI isn't always totally realistic. Okay.


And maybe criminals are not that smart. And from what we've seen so far, the people making up these data don't seem all that clever. So I'm not sure they're listening to normal curves and figuring out how to cheat better.


I'm not too worried about that. Yeah.


[Regina] (46:33 - 47:02)
All right. You've relieved my guilty conscience a little more. Okay.


So in the blog post, the Data Colada folks summarized all of these red flags as this, quote, we are unable to generate a benign explanation for all of the anomalies presented here. And then they said that they believe that the data were severely tampered with or fabricated to produce the desired results.


[Kristin] (47:02 - 47:15)
I agree. I do not have a benign explanation. And you could try to give the others the benefit of the doubt. Kate Laskowski talked about doing this with Jonathan Pruitt, right?


All possible ways this could have arisen innocently. And there's just none of those ways are plausible.


[Regina] (47:16 - 47:20)
All right. But that was study two. They also looked at study one.


[Kristin] (47:20 - 47:21)
Oh, wow.


[Regina] (47:21 - 47:21)
Okay.


[Kristin] (47:21 - 47:27)
That's the one in the real classroom, the MIT executive program, where the students' grades were actually at stake.


[Regina] (47:27 - 47:44)
Exactly. So the Data Colada folks didn't set out actually to do a deep dive into study one because that replication study we're talking about from Kyle and Alberto, it was just for study two with the proofreading tasks. But I think the Data Colada folks, Kristin, just, I don't know, they couldn't help themselves, maybe.


[Kristin] (47:45 - 47:53)
After uncovering all of that in study two, they weren't just going to walk away and say, let's not open the spreadsheet from study one. They were going to look. Curiosity.


[Regina] (47:54 - 48:36)
Curiosity. And in some ways, it turns out that study one was even weirder.


But it does have a fun twist at the end, too. Okay. So recall that in this study one, a key finding was that the students' overall grades are better when the instructor imposed regularly spaced deadlines than when the students chose their own deadlines.


And then they also did a subgroup analysis and reported that this difference in course grades went away when you compared only those students who chose evenly spaced deadlines versus students who had evenly spaced deadlines imposed on them.


[Kristin] (48:36 - 48:42)
Meaning that it was about the spacing of the deadlines rather than about the lack of choice.


[Regina] (48:42 - 49:32)
Right, right. And this really tied into the big hypothesis and theory in the beginning, so everything was nice and tidy. So the Data Colada folks just started out with what I think is a very reasonable thing.


They tried to use the spreadsheet data to recreate the results in the paper, right? The p-values and the means and everything. And this is when they found a key weirdness that eventually unraveled the whole thing.


So they started out just simple. And it was this. Remember, Kristin, you asked me how they calculated the overall course grade, right?


What the grading scheme was. And I said it was not clear that they didn't specify in the paper. Well, that was key.


That was a good data detective nose that you had because that was the thing.


[Kristin] (49:32 - 49:38)
Oh, well, now that we have the data, we can actually look and probably we can see what the grading scheme was, right?


[Regina] (49:38 - 50:00)
Mm-hmm. So the spreadsheet had six columns with grade information. There was a column that had the course grade, and then there were five columns with the grade components, with the individual things.


So they had one on the papers. They had participation grade. They had the project grade.


They had an exam grade. And they had some weird thing called B points.


[Kristin] (50:00 - 50:09)
So then did they try to reverse engineer, Regina, how the final grade was calculated, like what weight the project got versus the exam?


[Regina] (50:09 - 50:20)
They did. So quiz for you, Kristin. What kind of statistical analysis do you think that they ran using those six columns to answer the question about the grading scheme?


[Kristin] (50:20 - 50:30)
Well, I'm guessing that they ran a multivariable regression analysis where the final grade was the response variable and those five grade components were the predictors.


[Regina] (50:30 - 50:39)
Bingo. That is exactly what they did. So, Kristin, how about you explain why they did this and how it could reverse engineer the grading scheme?


[Kristin] (50:39 - 50:51)
Right. So grading schemes are usually some kind of weighted average, like maybe the exam is worth 40 percent, the homework is worth 30 percent. If you throw that into a regression, you can essentially tease that out.


Those weights come out as the beta coefficients of the regression.


[Regina] (50:52 - 51:09)
So they did exactly that. It's a great detective strategy. But guess what they found?


So they had 99 students. They put all the data into a regression. They got an R squared of 0.57. You want to talk about what R squared is and what it means?


[Kristin] (51:09 - 51:45)
Right. OK. So 0.57 is definitely not what I would have expected. I would have expected here, since the same grading scheme presumably should be applied to every student, if we're being fair, that we should get back an R squared of basically 100 percent. The R squared is telling us what percent of the variability in the outcome is explained by the predictors. And in this case, your final grade should 100 percent be explained by the individual components of that grade.


Otherwise, it would be weird. So what happened? This instructor was doing some voodoo, just deciding grades randomly.


What's going on?


[Regina] (51:45 - 51:53)
I love it. The R squared of 0.57 says that 57 percent of the variability in students' grades is explained by their grades.


[Kristin] (51:53 - 52:00)
Right. Then there's a fudge factor. How much coffee did you bring the teacher?


Right? That's not in the grade sheet.


[Regina] (52:01 - 52:10)
Apparently. OK. So they kept digging.


And then they thought, OK, what if the two sections use different grading schemes?


[Kristin] (52:10 - 52:22)
Oh, OK. That could explain it. Right.


We have these two sections. And maybe the instructors weighted things differently in the two sections. Oh, that would make sense.


And then if you put everything into one model, then you wouldn't get back an R squared of 100 percent. OK. Does that explain it?


[Regina] (52:23 - 52:51)
Well, kind of. They ran a regression on each section separately. Right.


And in the no choice section, they got an R squared of 100 percent. Beautiful. OK.


And they were able to recreate then the grading scheme. It looks like everyone got full points on participation, no matter what was recorded in the spreadsheet. And then all five components of the grade, which was the exam, the papers, et cetera, each were worth 20 percent.


[Kristin] (52:51 - 52:55)
So just a perfect, simple average of the five components of the grade. What about the choose your own deadline section?


[Regina] (52:55 - 53:03)
This is the hilarious part. They got an R squared for this of 0.24. Wow.


[Kristin] (53:03 - 53:10)
So only 24 percent of your grade is explained by the elements of your grade. Real voodoo magic going on here.


[Regina] (53:10 - 53:22)
Oh, even worse, though, Kristin, the results showed negative weight for some components. So the higher a student's exam grade, the lower their course grade. The higher their project grade, the lower their course grade.


[Kristin] (53:23 - 53:29)
Wait. So you lose overall points in your final grade if you did well on the exam? Are you are you?


[Regina] (53:30 - 53:34)
That's that's fun, isn't it? That would be fun to run a course like that.


[Kristin] (53:34 - 53:38)
I don't think the students would like it too much. Maybe the bad students would like it.


[Regina] (53:38 - 53:49)
Well, so the Data Coletta folks said, uh, what the heck? And they wrote, we think the answer to that question is data fraud.


[Kristin] (53:50 - 53:54)
OK, so did they find some actual proof that some of the grades had been altered?


[Regina] (53:55 - 54:20)
They found evidence that looked like grades had been changed, but not all of them, only 13 out of 49. So what they did is recalculate everyone's grade in that section using the grading scheme that they found for the other one, right? The simple weighted average.


And for 36 students, the grades matched perfectly. It was exactly what you would expect. It was 13 students that were very wrong.


[Kristin] (54:20 - 54:33)
That is a great way to check. Just go back and calculate the grades the way you think you should and then check and see if they match. And you're telling us most did match, but there's 13 that are off.


That smacks of highly suspicious that somebody went in and changed those 13 grades.


[Regina] (54:35 - 55:19)
And remember, they had this subgroup analysis, and they found that students who had chosen regularly spaced deadlines had better grades than those who chose deadlines that were not regularly spaced. So here's a thought experiment for you, Kristin. Let's say you were the poor researcher who ran this, and you have your hypothesis.


You want to show that those two subgroups are different, but, oh no, the original data did not show that pattern at all. It showed, let's just say, that it didn't matter, right? That the average grade in those two subgroups was the same.


You wanted to show one subgroup did better, and it was your job to tamper with the data in a way that would show that.


[Kristin] (55:20 - 55:34)
Right. I hope that wasn’t my job. But how would I do it if I had to, if there was a gun to my head? Well, I would either make some of the spread-out deadline grades higher, or I could make some of the non-spread-out deadline grades lower, or I could do both.


[Regina] (55:34 - 55:58)
Or do both. And that is pretty much exactly what they found when they looked at the data. So there were five grades that were changed from the spread-out deadline schedule, and they all had been changed to be higher, exactly the way you would need it to be.


There were eight grades that were changed from the non-spread-out deadline group, and all but one had been changed to be lower.


[Kristin] (55:58 - 56:18)
So that kind of helps rule out a more innocent scenario where maybe those grades accidentally got changed because of bad data handling, because if it was an accident, you'd expect it to be more random. We're equally likely to be higher or lower in both of those subgroups. The fact that it tracks almost exactly with what you need to get a certain result is very suspicious.


[Regina] (56:18 - 56:33)
And it turns out, Kristin, that if you kept those 13 grades as they should have been calculated, right, using the regular grading scheme, everything right, then that finding about the spread-out deadlines completely disappears.


[Kristin] (56:34 - 56:37)
Oh, wow. So somebody was really trying to manufacture that finding.


[Regina] (56:37 - 57:21)
They did. And the Data Colada folks noticed this other very interesting fact. So remember, we're talking now just about the choose-your-own-deadline section.


And the authors reported the mean for this choose-your-own-deadline section, the overall course grade, they reported the mean. Remember that?


[Kristin]
Oh, it was 86, right?


[Regina]
86. Exactly, exactly. Okay.


And now we're finding out, okay, someone went in and it looks like they tampered with the data to be able to get one subgroup higher than the other. But kind of miraculously, the overall average for the entire section did not change.


[Kristin] (57:21 - 57:33)
Oh, wow. Okay. So somebody is doing this in a purposeful way because they are trying to keep the original 86 and not change it, which is, again, suspicious of it not just being accidental.


[Regina] (57:33 - 57:51)
Exactly. And the Data Colada folks then went digging online now. Guess what they found?


An earlier version of the paper that was posted as a working paper that was under review at this journal. They had posted it at that French institute where one of the authors works.


[Kristin] (57:52 - 58:08)
Well, see, the Wayback Machine is your friend for statistical sleuthing. Sometimes you can actually find, because people like to say, hey, look, I'm publishing something, and they post it online somewhere a little early, and you can find it on the Wayback Machine. And then you can compare versions, and that can be very interesting.


[Regina] (58:09 - 58:29)
And it was indeed very interesting. They did not match. So the original version of the manuscript under review did not have that subgroup finding about people choosing regularly spaced deadlines versus having them imposed.


But the published version did. So where do you think that finding got introduced?


[Kristin] (58:29 - 59:01)
So I'm guessing that during the peer review process, right, some peer reviewer said, hey, did you look at this? And now you got to go back and reanalyze. And so maybe in doing that, the authors didn't find the result they wanted from that subgroup analysis.


So they're like, well, we can make it look the way we want. But they had to then, they couldn't change the overall average for that group, because they'd already told the reviewers what that average was. And it would look really weird if the average suddenly changed.


Yeah.


[Regina] (59:02 - 59:46)
Absolutely. This is the most likely scenario. And if you fudge the numbers really carefully, the peer reviewer probably will not be able to tell at all, except, Kristin, there was one clue in the changes, in the difference between these two manuscripts that could have alerted the peer reviewer, but apparently did not.


The Data Colada folks did not mention this, but I noticed it when I started digging in. Remember, I said I was annoyed because it looked like there were missing data that they didn't explain, right, based on the degrees of freedom. It looked like they had course grades on all 99 students, but only project grades on 97 students.


[Kristin] (59:46 - 59:47)
Oh, right. Yeah.


[Regina] (59:47 - 1:00:16)
And it was like, where did those two grades go? You know, did they get eliminated? Well, in the original manuscript that they submitted to the journal, the degrees of freedom were different than in the final published version.


So in the original manuscript, they never actually said what their sample size was, strangely. But we can tell from the degrees of freedom that they reported on 97 students, 97 course grades, 97 project grades, not 99.


[Kristin] (1:00:17 - 1:00:29)
Oh, wait a minute. So somehow in between the original manuscript and the resubmitted manuscript, they found two additional people, two more course grades at least overall, although not two project grades.


[Regina] (1:00:29 - 1:00:48)
I know, right? It did not make any sense. So I went back to the spreadsheet, right, of the original data, and that's when I saw there were two rows that only had the final course grade, no paper grade, no project grade, nothing else, just kind of like a blank line and the final course grade at the end.


[Kristin] (1:00:49 - 1:00:57)
Well, that's suspicious. How could you get data on their course grade but not any of the work they actually did during the term?


[Regina] (1:00:57 - 1:01:19)
Mm-hmm. And I noticed that these were both in that choose your own deadlines group, which was what we needed for that subgroup analysis, and one of them had spread out deadlines and the other did not. And if you took out these two data points, the grade average is not exactly what they had reported in that first manuscript submission.


[Kristin] (1:01:19 - 1:01:40)
Oh, so they weren't clever enough to get it to work out to 86 just by manipulating some grades. So they're like, oh, we'll just add two grades to make the numbers work out. Wow, yeah.


Boy, that is like the height of laziness, though, is like if you're going to make up data, you can't even go back and make up some project grades and some paper grades.


[Regina] (1:01:40 - 1:01:55)
I thought maybe they had lost project grades. No, no, no. They just found two imaginary people.


[Kristin]
Yeah, wow.


[Regina]
That's what it was. So the degrees of freedom could have been a clue to the peer reviewer to probe further, right?


We had the 97, the 95.


[Kristin] (1:01:56 - 1:02:13)
That's why as a peer reviewer, you should pay attention to the degrees of freedom. Well, the weird thing, though, is I can't believe that in their original submission, they didn't put the sample size. I mean, isn't that the most basic information?


Like you say, we had this many participants in this group and they never said that anywhere.


[Regina] (1:02:14 - 1:02:15)
And they never said that anywhere.


[Kristin] (1:02:15 - 1:02:19)
They were hedging their bets. Oh, if we don't, if we don't commit, we can change it later. As needed.


[Regina] (1:02:21 - 1:03:27)
Oh, which is funny because the whole paper is about pre-commitment. We're not going to pre-commit.


OK, but Kristin, I noticed one other thing in the data that was a little weird. So in this spreadsheet, you can see that it looks like the researchers had sorted the data so that everyone in the choose your own deadline section was at the top. Everyone in the regular section was on the bottom.


And the people in the choose your own deadline section, they have information on what deadlines they chose. So they have these three columns with information on the three deadlines. And in the regular section, they just have these columns blank because they didn't choose any deadlines.


OK, so you're scrolling down through the data and you get to the regular section. And all of a sudden, you have this big chunk of six students smack in the middle of the regular section who look like they belong there because the columns with the chosen deadline information is blank. But those six students are labeled as being in the choose your own deadline section.


[Kristin] (1:03:29 - 1:03:37)
So somebody sorted by section and then decided to change six participants from one section to the other. Is that the inference here?


[Regina] (1:03:37 - 1:03:50)
I mean, it is suspicious. There's no smoking gun, but it just adds a little bit more suspicion. And maybe explains why we have these.


Remember, I said we were missing six missing people.


[Kristin] (1:03:50 - 1:04:16)
Yeah, I would explain the six missing data points that you had, Regina, on deadline. So these six are missing deadlines because they weren't actually in the choose your own adventure deadline group. And, you know, this whole sorting trick where people sort things and then they go in and change afterwards.


It makes it kind of obvious that you've changed things because things aren't where they are supposed to be, according to sorting. So that is something that's going to come up in the Francesca Gino case, too, which we're going to talk about in the next episode.


[Regina] (1:04:17 - 1:04:43)
So it's kind of like no smoking gun here, but just like another thing that is making sense of all of those weird sample size things that we noticed before. Adding more suspicion. Oh, one more thing I noticed from looking at the data and comparing those manuscript versions.


Remember, we couldn't figure out why they were reporting on the overall course grades instead of the actual paper grades, right? You asked me about that. Right.


[Kristin] (1:04:44 - 1:04:44)
That didn't make sense.


[Regina] (1:04:45 - 1:04:53)
Uh-huh. Because when you go back in and look at the effect on the paper grades, no statistically significant difference at all.


[Kristin] (1:04:53 - 1:05:01)
Oh, that is interesting. So maybe they tried that out, didn't get a significant result. And so then they said, but what if we look at the overall grade?


[Regina] (1:05:02 - 1:06:05)
Which is why they reported on the project grade and the overall grade, which they were able to get a significant difference for, but not the exam grade, not the papers grade. Right, right, right. Okay, so those are extra little things that I found, which means that there are probably many other statistical cockroaches in here.


This is just what I found. But the Data Colada blog, remember, I said it had a fun plot twist. They had a nice postscript.


And they had described this hypothesized scenario with the peer reviewer, right? They said, oh, maybe this all came because a peer reviewer asked for something. So they described that in their blog post.


They were about to publish. But before publishing, they sent the draft to the two authors of the paper, Dan Ariely and Klaus Wertenbroch. And Klaus asked if there was anything that he could do to help.


And the Data Colada folks said, well, if you happen to have a copy of that peer review back from 2001, that would be cool. But, you know, we know that was 25 years ago. And you know what Klaus did?


[Kristin] (1:06:05 - 1:06:18)
Wow. I'm thinking that he had good file management, and he saved the peer review and sent it to them. And also, this is making me feel like the fact that he's cooperating might mean that he did not know there was tampering of the data. Yes.


[Regina] (1:06:18 - 1:06:23)
Yes. Yes. Yes.


There are many indications in here that he did not see the actual data.


[Kristin] (1:06:23 - 1:06:42)
I think there's a lesson here, Regina, also from the last episode. Keep your files and keep them where you can find them. And just because you're not on that same laptop 20 years later does not mean that you should not have access to the data.


So did the peer review report actually ask for that subgroup analysis? Now I'm dying to know.


[Regina] (1:06:43 - 1:07:13)
They asked for it straight out, very explicit, reviewer two, because it's always reviewer two. I said, you need to show that the people who imposed deadlines similar to those imposed by the experimenters did about as well as the people who had deadlines imposed by the experimenter. And the editor said the same thing.


And then, lo and behold, submitted manuscript had exactly what they asked for in pretty much the exact same words in a way that still supported the original results.


[Kristin] (1:07:13 - 1:07:49)
You know, I don't love that reviewer comment. You can't say, you have to show this. Right?


You should say, did you test this? Not, did you have to show this? That's almost like a leading question for data manipulation.


So that's kind of amazing that we hypothesized that that might have been what happened. And in fact, that is what happened. And it really just points to data tampering in multiple different ways, actually.


There's just so many red flags here. There's not really any other explanation than that somebody went in and messed with these data to get the result they wanted.


[Regina] (1:07:50 - 1:07:53)
Hmm, so many statistical cockroaches underneath the rug, right?


[Kristin] (1:07:53 - 1:08:10)
Oh, yeah. This house is about to collapse. Okay, Regina, this has been fascinating.


But now I think we're ready to wrap this up. And can you state again the claim for today? For this episode, we had to get a little creative since we weren't like looking at a nutritional supplement.


[Regina] (1:08:11 - 1:08:18)
There you go. It was many people who make up or alter data are surprisingly not very good at it.


[Kristin] (1:08:19 - 1:08:42)
And how do we rate strength of evidence in this podcast with our highly scientific one through five smooch rating scale, where one smooch means little to no evidence for the claim, and five smooches means a lot of evidence for the claim. And I know since we've been doing some different episodes the last few times, these claims are a little bit more based on anecdotal evidence than anything else. But what do you think, Regina, kiss it or dis it?


[Regina] (1:08:42 - 1:09:08)
Oh, I'm gonna give five smooches on this one. I feel like the statistical cockroaches, right, were kind of everywhere. I was able to catch a few hints of a problem just reading the published version.


But then, of course, once you get the first submitted draft, and then the data, you find more and more. And you think, well, this should have been better done if you're gonna lie. What about you?


[Kristin] (1:09:10 - 1:09:44)
Yeah, I'm going four smooches, just because I think all of these cases when they come out, you're like, why don't people do this better? But those are only the people we caught, or the Data Colada people caught.


So I'm wondering if there is a whole set of academic criminals who are just better at this and they never get caught. So since I can't rule out that possibility, I'm going four smooches. But definitely, in many of the cases that we see, you do go, I guess they thought nobody was ever going to look at their data, because they weren't too clever about it.


[Regina] (1:09:44 - 1:09:50)
Right. This is survivorship bias. You're right.


Anti-survivorship, we're only seeing the ones that got caught.


[Kristin] (1:09:51 - 1:09:55)
Okay, Regina, how about methodologic morals? What do you have for us today?


[Regina] (1:09:55 - 1:10:17)
Yeah, I wanted to highlight all the cool things that we have gone through today and talk about that. A data detective's toolbox is surprisingly simple. Check the degrees of freedom, look for twins, question giant effects, and ask whether the data actually make human sense.


[Kristin] (1:10:18 - 1:10:34)
Oh, that's a great summary of today's episode, actually, Regina.


[Regina]
What about you?


[Kristin]
I'm going with the theme of, we didn't really have to use much of our statistics degrees today. So maybe knowing something about degrees of freedom, but other than that.


So mine is, you don't have to be a statistician to catch a statistical criminal.


[Regina] (1:10:35 - 1:10:45)
Oh, I like that one. That's exactly right. Some of it is just coming in with the right frame of mind.


[Kristin] (1:10:46 - 1:10:48)
And common sense. And checking anything. Checking at all would be the first place to start.


[Regina] (1:10:49 - 1:10:54)
You don't have to be a statistician, but you just need to be a skeptic, I think.


[Kristin] (1:10:54 - 1:11:01)
Yeah. All right, Regina, this has been super fascinating. And I think we should start a mystery podcast now.


[Regina] (1:11:02 - 1:11:05)
I've heard that true crime is a very popular podcast genre.


[Kristin] (1:11:06 - 1:11:10)
We can cross this over with true crime, Regina. And then we can get more listeners, right?


[Regina] (1:11:10 - 1:11:20)
True statistical crime. I don't know if there's a big crossover in the audience, but there might be.


[Kristin] (1:11:20 - 1:11:30)
True crime with numbers. True crime with data. Data crimes.


[Regina]
Data crimes.


[Kristin]
That's our next podcast. All right. Thanks, Regina.


[Regina] (1:11:31 - 1:11:32)
Thanks, Kristin. Thanks, everyone, for listening.