Ctrl-Z Awards: Will retracting your paper ruin your career?
What happens when you discover that your published paper is wrong? Does retracting it ruin your reputation—or could owning the mistake actually make things better? In this episode, we meet the inaugural winners of the Ctrl-Z Awards, which recognize scientists for correcting or retracting their own work. Behavioral ecologist Kate Laskowski takes us inside a scientific detective story involving suspicious spider data, three retracted papers, and the fear that going public might tank her career. Then computational social scientist Alyssa Smith tells us how one small error in a figure led her to retract the paper at the center of her PhD and job talk. Along the way, we talk about data forensics, the stigma around retractions, and why self-correcting scientists are the heart of science.
Statistical topics
- Data management and documentation
- Open science
- Reproducible workflows
- Research integrity
- Retractions
- Scientific misconduct
- Statistical forensics
Methodologic Morals
- “Science is self-correcting, which means scientists need to be self-correcting too.”
- “Good documentation is boring—right up until it saves you.”
- “If you want scientists to correct their mistakes, reward them for correcting their mistakes.”
- “I'd rather be considered an honest idiot than for people to wonder if I'm untrustworthy.”
References
- Retraction Watch: https://retractionwatch.com/
- Ctrl-Z Awards: https://retractionwatch.com/ctrl-z-award/
- Announcement of inaugural Ctrl-Z winners: https://centerforscientificintegrity.org/2026/08/04/2026-winners-ctrl-z-award-courage-retracting-own-work/
- Kate Laskowski’s blog post: https://laskowskilab.faculty.ucdavis.edu/2020/01/29/retractions/
- Kate Laskowski’s Data Detective Course syllabus
- Alyssa Smith’s site: https://asmithh.github.io/
- Laskowski KL, Montiglio PO, Pruitt JN. Individual and Group Performance Suffers from Social Niche Disruption [retracted in: Am Nat. 2020 Feb;195(2):393. doi: 10.1086/708066.]. Am Nat. 2016;187(6):776-785. doi:10.1086/686220
- Laskowski KL, Pruitt JN. Evidence of social niche construction: persistent and repeated social interactions generate stronger personalities in a social spider [retracted in: Proc Biol Sci. 2020 Jan 29;287(1919):20200077. doi: 10.1098/rspb.2020.0077.]. Proc Biol Sci. 2014;281(1783):20133166. Published 2014 Mar 26. doi:10.1098/rspb.2013.3166
- Modlmeier AP, Laskowski KL, DeMarco AE, et al. Persistent social interactions beget more pronounced personalities in a desert-dwelling social spider [retracted in: Biol Lett. 2020 Feb;16(2):20200062. doi: 10.1098/rsbl.2020.0062.]. Biol Lett. 2014;10(8):20140419. doi:10.1098/rsbl.2014.0419
- Smith AH, Green J, F Welles B, Lazer D. Emergent structures of attention on social media are driven by amplification and triad transitivity [retracted in: PNAS Nexus. 2026 May 07;5(5):pgag137. doi: 10.1093/pnasnexus/pgag137.]. PNAS Nexus. 2025;4(4):pgaf106. Published 2025 Apr 1. doi:10.1093/pnasnexus/pgaf106
Kristin and Regina’s online courses:
Demystifying Data: A Modern Approach to Statistical Understanding
Clinical Trials: Design, Strategy, and Analysis
Medical Statistics Certificate Program
Epidemiology and Clinical Research Graduate Certificate Program
Programs that we teach in:
Epidemiology and Clinical Research Graduate Certificate Program
Find us on:
Kristin - LinkedIn & Twitter/X
Regina - LinkedIn & ReginaNuzzo.com
EPISODE 40 CTRL Z FINAL
[Kate] (0:00 - 0:12)
Because at some point I was just like, you underestimated me.
[Kristin]
I love that.
[Kate]
You think that I care more about the publication, but really what I care about is the right answer.
[Kristin] (0:18 - 0:28)
Welcome to Normal Curves. This is a podcast for anyone who wants to learn about scientific studies and the statistics behind them. I'm Kristin Sainani.
I'm a professor at Stanford University.
[Regina] (0:28 - 0:33)
And I'm Regina Nuzzo. I'm a professor at Gallaudet University and part-time lecturer at Stanford.
[Kristin] (0:34 - 0:39)
We are not medical doctors. We are PhDs, so nothing in this podcast should be construed as medical advice.
[Regina] (0:39 - 0:44)
Also, this podcast is separate from our day jobs at Stanford and Gallaudet University.
[Kristin] (0:45 - 0:47)
Regina, today we're going to do something a little different.
[Regina] (0:47 - 1:05)
That's right. We are breaking from our normal format because usually we examine a paper or multiple papers and talk about the statistics behind the claim. But today, we're going to talk about a recent award that's actually relevant to the kinds of things we talk about on the podcast, so it's not the Emmys.
[Kristin] (1:06 - 1:27)
But from a statistician's point of view, this is actually better than the Emmys, Regina. It's a new award that recognizes scientists for something we usually don't give prizes for, mistakes. It's called the Ctrl-Z Awards, and they say they are recognizing the courage to correct or retract published work.
Today's episode, we're going to interview the two inaugural winners.
[Regina] (1:27 - 1:38)
We were really excited when we heard about these awards. It is such a clever idea. It's wonderful, and it feels like exactly what science needs right now.
[Kristin] (1:38 - 1:48)
Yes, and Ctrl-Z, of course, that's the keyboard shortcut for undo, and the awards are given by the Center for Scientific Integrity, which is the nonprofit behind Retraction Watch.
[Regina] (1:48 - 2:00)
Kristin, you and I are both huge fans of Retraction Watch, and we know the co-founders, Ivan Oransky and Adam Marcus, and also the managing editor, Kate Travis.
[Kristin] (2:00 - 2:05)
Yeah, and Retraction Watch has been around since 2014, and it's been growing ever since.
[Regina] (2:05 - 2:19)
I love their tagline, tracking retractions as a window into the scientific process, and I encourage everyone to go look at their website. And this Ctrl-Z Award fits in perfectly with their goals.
[Kristin] (2:20 - 2:37)
It does. And Regina, even though this episode is a little different, we are going to stick with our traditions on Normal Curves, and so we are going to give a claim to evaluate, even though it's not about a health thing. So, Regina, here is the claim for today. Acknowledging and correcting your mistakes may have better outcomes than you might expect.
[Regina] (2:38 - 3:00)
I love that one. A little background on the Ctrl-Z Award. It was conceived and funded by two guys who actually know their stats, Harvey Motulsky and Earl Buetler, and they are longtime supporters of Retraction Watch and actually co-founders of the company behind GraphPad Prism Software.
[Kristin] (3:00 - 3:06)
And Regina, the Ctrl-Z Award is not just like a certificate. It actually comes with some money, $2,500 for each winner.
[Regina] (3:06 - 3:10)
Which almost makes me want to go out and make some mistakes, right, and correct them.
[Kristin] (3:10 - 4:02)
I know, that's actually like real money. But I don't think the point is to motivate new errors, of course. All right, so there are two awards, one for a senior researcher and one for a junior researcher.
The senior award went to Kate Laskowski, who is an associate professor at UC Davis. And the junior award went to Alyssa Smith, who is an assistant professor at the College of the Holy Cross. They also gave special commendation to Pamela C. Ronald, who is also a professor at UC Davis. We interviewed both of the winners for today's episode. These are really fascinating stories.
I'm so excited to have these guests on. And we're going to start today with the winner of the Senior Researcher Award, Kate Laskowski. Welcome, Kate.
We are thrilled to have you on the podcast. Thanks so much for being here.
[Kate]
Yeah, I'm happy to be here.
[Kristin]
And if you don't mind, could you briefly introduce yourself? Tell us about your background and your current position.
[Kate] (4:03 - 4:13)
I'm an associate professor in the Department of Evolution and Ecology at UC Davis. I consider myself a behavioral ecologist, and I'm interested in the developmental origins of individuality.
[Kristin] (4:13 - 4:28)
So, Kate, I watched one of your talks on the whole experience that you're going to talk about, which got you the Ctrl-Z award. And I have to say, you seem pretty savvy about statistics. So since this is a statistics podcast, I'm curious, are you self-taught, or was this part of your graduate training?
[Kate] (4:29 - 4:56)
I mean, I think mostly self-taught. I certainly took a couple stats courses as a PhD student. But then my research very quickly veered into the realm of random effects.
And at the time, my advisor didn't know how to handle random effects. I didn't know how to handle random effects. And so then that took me down the rabbit hole of figuring out what are mixed models?
How do they work? And so then everything from there on out is basically self-taught.
[Kristin]
Amazing.
[Regina] (4:56 - 5:13)
So, Kate, when you tell this story today, there are two things that are going to play a major role. One is a man named Jonathan Pruitt, and another one is spiders. So how did you start working with Jonathan and with spiders?
[Kate] (5:14 - 7:09)
Yeah, I mean, I never worked with the spiders per se. Jonathan was the one who was handling all the spiders. But so the whole genesis of this is, you know, I'm a PhD student.
I'm finishing up my final year or two of my PhD. And I'm starting to think of like, okay, I need to find a postdoc here. And my advisor, Alison Bell at the University of Illinois, she's really well connected.
She knows lots of people. And she was like, Oh, why don't you start going to these conferences, and you can talk to this person, that person, this person. And so I started trying to put these feelers out.
And it was at the 2012 behavioral ecology conference in Sweden, that I decided I should approach Jonathan Pruitt. And so at the time, Jonathan Pruitt was a faculty member at the University of Pittsburgh. And he was like the rising star in behavioral ecology.
I mean, everybody knew him, he's super gregarious. He had lots of friends, and he's doing just what appears to be like really incredible work using this really cool system, which are these social spiders. And so if you've never heard of social spiders, I warn you, they are the thing of nightmares.
You're not if you're not inclined to like spiders, they are a little bit terrifying. They themselves are quite small. They're only I don't know about the size of like a nickel or a quarter.
But they live in these big colonies with lots of different individuals. And they do this so they can help each other out and take down bigger prey. And so they can take down things as big as birds, I'm told.
They also is a matriarchal society. So it's mostly females, which I've personally loved. And they all take care of each other's young.
So they do alloparental care. But yeah, it's basically a society of spiders. And this was the system that Jonathan had really named his name on.
And so I approached him in 2012 to basically say, Hey, Jonathan, I have this really cool idea for this experiment. And I'd love it if we could test it in your system. And he very much was on board from the get go.
[Regina] (7:09 - 7:17)
So you did not run the experiment with the spiders. You did not play with the spiders. But this was testing one of your academic hypotheses.
[Kate] (7:18 - 9:35)
Yes, it was a hypothesis that I didn't necessarily develop myself. But my whole dissertation was around testing these sort of theoretical predictions that other folks had come up with. And so one of them was the social niche hypothesis.
And I found it very cool, because it kind of like resonates with our human experience. And the general idea of this social niche hypothesis is that if you live in a group, like with other individuals, it might be beneficial to settle into some sort of social role, right. And so you know, individual one does this task, individual two does this other task.
And that way, you can either more effectively cooperate, or more effectively avoid competing with each other. And so the social niche hypothesis is really cool, in part, because one, it's like fun and exciting. But two, it makes really clear testable predictions, which is that the longer you've been living in a group with these other individuals, the stronger your individuality or your personality should be.
And so that was my idea is I said, Oh, I have this really clear testable prediction. And you know, these social spiders who live in groups their whole lives seem like a really great system to test it in. So like Jonathan and I sat down at this conference in Sweden, we designed the whole experiment together.
So it was like, oh, we need these number of colonies, we're going to manipulate them in this way. And we're going to take these behavioral measurements. And then he went back to Pittsburgh, and he said he collected these data.
And then he sent me the file. And so the first thing I did when I got this data file is I import it into R. And I start doing graphical checks.
I did histograms, I did dot charts, I looked for weird outliers, I looked for relationships, I did all the things I know are the appropriate thing to do when you get a data file in the sense of I wasn't looking for fake data, I was just looking to see like, oh, there might be some like mistakes here, let me check and make sure that nothing looks too weird. And it didn't is the answer. I mean, there was some problems of like the behavioral assay that Jonathan said he did was a 10 minute behavioral assay where he scared the animal, and then measured how long it took them to recover.
And so we ended up having, you know, very much censored data where we had a lot of values at the maximum value, because the animal didn't recover in that amount of time. But you know, that's like normal, normal behavior work. And so I saw that pattern, I knew it was there, we figured out a way to deal with it.
And then I moved forward with analysis.
[Regina] (9:35 - 9:45)
So this is going to be important to the story, you got an Excel file, it was not a comma delimited file, it was not just a text file. So in Excel file with tabs.
[Kate] (9:45 - 10:25)
Yes, I'm glad you remembered this detail. Yeah. So I got an xlx file, which for those of us who use R know that it can be notoriously difficult to import an xlx file into R, right?
So what do you do? You say save as comma delimited CSV. And when I don't know if you've ever done this, if you say save as CSV, this little dialog box pops up that says, hey, are you aware that you're only going to save the active sheet?
And you say, yeah, of course, like, that's all I want. I just want sheet one. And so you say yes, and then boom, now you have your CSV that only has the active sheet one sheet from your Excel file in it.
And that's what I imported into R and was doing all my other analysis with.
[Regina] (10:25 - 10:29)
That's a little teaser foreshadowing. That will come up again in the story.
[Kate] (10:29 - 10:30)
It's foreshadowing.
[Kristin] (10:32 - 10:42)
Alright, so you get the data, you analyze it, you find something that supports your hypothesis, I believe, and you publish it. How did that go?
[Kate] (10:43 - 11:36)
Honestly, great. These are some of the easiest and most enjoyable papers to write. Because when we had this really clear hypothesis, like a priori hypothesis, clear predictions, and then lo and behold, the results so clearly fit the predictions, that it's just a breeze to write. It's like, you just can't believe it.
You're like, Oh my god, these social spiders, they have social roles, the longer they live in their groups, the stronger their personalities become. It was so cool. It was one of the coolest when I first get those repeatability statistics, which are what we use to estimate the magnitude of individuality.
I just I couldn't believe it. I was just like overjoyed. And these papers, ultimately, they got a ton of attention, like it got written up in NY Times, because it is like this really cool kind of exciting fun result of like social spiders have social roles, they have strong personalities.
And like, who doesn't love that?
[Kristin] (11:36 - 11:39)
Right? Wow. And so you ended up writing three papers out of this, correct?
[Kate] (11:40 - 11:40)
Yeah, exactly.
[Kristin] (11:40 - 11:48)
Alright, so then fast forward a little bit, and you get tipped off that there might be a problem in one of the papers. So take us back to that time. What happened?
[Kate] (11:48 - 12:42)
Yeah, so these papers were published in 2014 to 2016. I'm in my postdoc at that point. And then flash forward now fall of 2019.
I have just arrived at Davis, I am literally two months into my job, I'm sitting in my office, I don't even have furniture yet. And I get an email from a mentor of mine. He's like a colleague, and he's helped me throughout my career.
And I trust him. And he says, Hey, Kate, I downloaded the publicly available data set for one of your papers. And I'm noticing these weird patterns.
And so then I'm like, well, what weird patterns could possibly be there. And then that is what basically kickstarted this whole thing is as soon as I he told me this, and I realized what I had to look for in terms of there is these exact duplicate values that didn't seem to make sense based on the experimental design. Then that just sort of opened up this whole new puzzle box that I needed to solve.
[Kristin] (12:43 - 12:45)
I love that you treated it like a puzzle box, I have to say.
[Kate] (12:45 - 12:54)
That's exactly how I felt about it. I mean, that's why I'm a scientist, right? It's because I love puzzles.
I love solving problems. And this, this was like the ultimate puzzle in a weird way. Yeah.
[Regina] (12:55 - 13:03)
So Kate, the data that your colleague was downloading was the original Excel file?
[Kate] (13:03 - 13:41)
No, it was not the original Excel file. So it was the sort of the cleaned up pared down file that I had done my final analysis on, and then uploaded on to dryad. Because this paper that the data was associated with was published in the American naturalist, which shout out to the American naturalist, they were one of the very first journals, if not the first journal to require the public deposition of data.
And so that was their requirement. They were always ahead of curve. They're very big proponents of research integrity.
And so that's what we did. We deposited the data. But yeah, it was like a CSV file that was pared down and just like just the important things that you needed for the analysis.
[Kristin]
Got it.
[Regina] (13:41 - 13:59)
So young, new faculty member, realizing that your paper might have a problem, because the data might have a problem. So were you immediately suspicious of fraudulent data? Or were you kind of trusting, oh, there must be a reasonable explanation here?
[Kate] (13:59 - 15:01)
I mean, I certainly thought there was a reasonable explanation, right? I mean, really, my first hypothesis was like, oh, no, I did something wrong in the data file during analysis. I somehow like copied and pasted a bunch of numbers or like copied a column.
And I've now published a paper that's on totally messed up data. And that's so deeply embarrassing. That was like really my first thought.
Luckily, you know, I had the original data files, I had my code. So I had all my scripts from my R file that I could go back and confirm. No, no, I didn't alter the original data file in any way.
So my first thought was that. And then after that, I was like, well, I don't know why these duplicate values are in here. Let me talk to my collaborator who collected the data.
And then Jonathan had this explanation that I call it the block design explanation. He said that these duplicate values of this behavioral measurement occurred, because the animals were all being observed at the same time. And so if they did the same behavior, basically, at the same time, they just wrote down the exact same number.
So that was his explanation, which I mean, I didn't love it, because it was, to me, it felt kind of sloppy. Oh, yeah.
[Kristin] (15:01 - 15:07)
From a statistical standpoint, then you have correlated observations, which I'm assuming you didn't account for in your analysis, right?
[Kate] (15:09 - 16:14)
100%. Yeah. And so, as you know, I love random effects, random effects are like my favorite thing in the world.
And so this is exactly what I told Jonathan, I said, well, send me the schedule of the blocks. And then I at the very least, you know, we can publish a correction and say, hey, our methods are a little bit different than what we thought or that what we first explained. And I can redo the analyses now with this random effect of block to at least appropriately account for the non-independence.
But Jonathan had moved labs many times by that point. And he said, if the schedule ever existed, it's been lost in the moves. And I said, okay.
And so then, that's when this is my puzzle of I was like, well, you know what I could do, I could recreate the blocks. If the idea is, is that the same duplicate values are happening, because the animals are getting observed at the same time. If I find two animals that have the same value, I'll put them both in block A, and then the next two in block B.
And eventually, my thought was maybe I can more or less recreate the blocks. So that's what I started to do. And then I was color coding these Excel files.
And that's when I realized the duplicates were a lot more problematic and could not be explained by this block design explanation.
[Kristin] (16:15 - 16:25)
You did a ton of statistical sleuthing, actually, Kate, to figure this out. So tell us a little bit about so your color code and everything. And yeah, we love statistical sleuthing on this podcast.
[Kate] (16:25 - 18:03)
It wasn't like statistical sleuthing. It was just like Excel sleuthing.
[Kristin]
That counts.
[Kate]
And so there's a lot of other people who were involved in this, right? Like, as soon as I we published the retraction notices of my papers, this was sort of Jonathan's pattern in the sense he would either approach early career researchers or early career researchers would approach him and he was very generous with the idea of like, come to me with your ideas. I have this incredible study system.
I have an army of undergrads, and I'll collect data for you. And then he would send data to these early career researchers, they'd all publish their papers. And so what that meant is there's now a whole community of early career researchers who are suddenly investigating their data and finding weird things.
The Pruitt survivors, we were all trying to figure out like, what is going on. And eventually, we ended up calling them data quilts. Because when you're like Control-F-ing in your Excel files, and you're like color coding to try and find the pattern, it ends up looking like a quilt.
I love that. But yeah, that's exactly what I was doing. I just Control-F-ing through my Excel file to be like, Okay, I know there's a duplicate, like this is a duplicate value, where's the other duplicate value.
And then I would color code it based on if it was an exact duplicate, if it was almost an exact duplicate, but not quite. And then I would color code it, is it a duplicate in a row? Is it a duplicate in a column?
Because really, what it felt like is it just felt like this Rubik's Cube of like, I was just so close to figuring out the key of there's clearly a pattern to these duplications. And I just can't quite figure it out. I felt like a dog with a bone of just like, I have got to solve this little puzzle that I have in front of me.
[Kristin] (18:04 - 18:11)
At what point did you realize that this might not be just an honest mistake? At what point did that realization start to creep in?
[Kate] (18:11 - 18:41)
I think the first time I started to be very nervous is when I noticed that the duplicates were occurring in the same animals. So it's not like the duplicates are between two different individual animals, which would be consistent with the block design explanation. It was that individual one would have value one at observation one, but then value one at observation five.
That's the same literal individual animal. And that's when I was like, that doesn't make sense. And these are down to several decimal places, right?
[Kristin] (18:42 - 18:42)
Exactly.
[Kate] (18:43 - 18:58)
And so it wasn't just one animal, right? Like one animal, sure. Maybe these things randomly happen, but it was like many, many, many, many, many.
It was like most of the animals. So that's when I started being like, this is really bad. And then the real smoking gun that we alluded to earlier is sheet two.
[Kristin] (18:59 - 19:02)
Yes. That was how you found that.
[Kate] (19:03 - 19:46)
Yeah. So I alluded to earlier, like all of my analysis and most of my data sleuthing up until this point had been done in the CSV files that I had saved from the original Excel files. But at this point, I think I'm like, I've already decided to retract paper one.
I'm now in the process of deciding to retract paper two. And what I say to myself is, okay, I'm going to re-download the Excel file that Jonathan sent me in my email. So that way it's, I know I haven't accidentally messed it up.
And then I'm going to color code it in the way I've been doing it on sheet one. And then on sheet two in the same workbook, I'll click on sheet two and I'll add in all my little notes. So I can send this to the editor and say, here's all the things I'm seeing.
Please retract this paper. It's clearly a lost cause.
[Regina] (19:47 - 19:51)
Because you assume sheet two is empty. So you were just going to use that as like a notepad.
[Kate] (19:51 - 21:34)
Yeah. So I was thinking that exactly. I would use sheet two as a notepad.
And then at least it's all in one file and editor doesn't have to open multiple files. And so I click on sheet two, expecting it to be blank. And it's not.
Sheet two is not blank. It's filled with numbers and they're filled with numbers in like a weird way. So on sheet one, which was the supposed data, everything's in basically the long format.
Every observation is a row. So if we observed the animals five times, every individual had five rows, five observations. And then I noticed on sheet two, it's those same values, but now they're transposed into this weird wide format that doesn't like, I don't know why you would store actual biological data in this format, but that's how it was.
And I noticed that the sheet two data, they, it was not all the data. It was only our treatment colonies. So we had control colonies and then we had the treatment colonies and the treatment colonies are the ones that had to have a certain level of repeatability, a certain statistic in order to be consistent with our hypothesis.
So I have this weird, these weird data formatting on sheet two, and I'm looking at it and I'm like, well, I know there's duplicate values on sheet one. Let me find out where they are in sheet two. And so I start Control-F-ing again, data quilting everything.
And what becomes obvious is when you put the data in this weird sheet two format, it's not just single values that are duplicated. It's just entire blocks of like literally a hundred values that are duplicated identically between multiple different treatment groups. And when you transpose that back to the long format, that like block isn't as obvious when it's in this transposed format, it's super obvious.
[Kristin] (21:35 - 21:40)
So perhaps the data started in that format and then they moved it over where it's less obvious.
[Regina] (21:41 - 21:43)
That is, that is my hypothesis as well.
[Kristin] (21:44 - 21:44)
Interesting.
[Regina] (21:45 - 22:05)
Kate, this is very visually striking. You're talking about this like a quilt and you do a great job in your blog post showing this progression bit by bit. So we will definitely link in the show notes to your blog.
Once you see those patterns, you cannot unsee it. It is so clear in here.
[Kate] (22:05 - 22:24)
The frustrating thing with this sort of thing is that once you know what to look for, it's obvious, but like why on earth would you go looking for this? You would only go looking for these sorts of patterns if you didn't trust your collaborator, but like, you know, at face value you should. Right.
[Kristin]
Why would you be collaborating with them if you didn't trust them in the first place? Yeah.
[Kate]
Yeah. A hundred percent.
[Regina] (22:24 - 22:29)
Getting back to the sequence of events, how many papers did you end up retracting?
[Kate] (22:29 - 22:42)
Ultimately three. Two that I was first author on and that I completely wrote those papers and did the analysis for. And then a third that I helped out with the analysis, but one of Jonathan's other postdocs wrote the paper.
[Regina] (22:43 - 23:07)
So getting back to that first paper that you needed to retract, that must have been difficult to make that decision. And I saw a talk that you gave on this topic a few years ago and you said something that I loved. You said, I'd rather be considered an honest idiot than for people to wonder if I'm untrustworthy.
So tell me about that.
[Kate] (23:07 - 24:41)
Yeah, that was ultimately the reason, that phrase is why I decided to publish the blog post that is associated with my second retraction. And it was because, you know, I had one paper retracted and Jonathan and I agreed to the wording of the retraction notice, but I always found that the wording of the retraction notice was a bit vague. It said like, oh, we found data irregularities.
We cannot explain these anomalies in the data. And, you know, for one retraction, maybe that's fine. People will assume, oh, something weird happened.
Who cares? Whatever. Good on them for retracting.
But then retraction two comes out and now we have basically the same retraction statement that says there's these irregularities in the data. And I mean, it just was, if this was me and I was reading that someone else retracted two papers for data irregularities that cannot be explained, I mean, I'm going to just wonder like, what the hell is going on? Right?
And so my fear, my real, like the thing that kept me up at night and that I was like, I have to do this is I just knew that if I just didn't come out and just say, here's everything that I'm seeing, every single conference I went to for the rest of my career, people were going to ask me about it. If I'm lucky, right? At best, they will ask me.
At worst, they're going to wonder what I knew and why I was being so weird and vague about it. And so I just couldn't live with this idea that people were suddenly going to be like, there goes Kate, what do you think she knows about those data irregularities? I can't live with that.
I'd rather just have you all think I'm kind of dumb and I got duped than to think that I manipulated data in any way.
[Kristin] (24:41 - 24:43)
Honesty, the best policy. Yeah.
[Kate] (24:43 - 25:01)
I mean, I just didn't know what else to do, quite honestly. I felt stuck between a rock and a hard place because I couldn't get the retraction notice to say anything stronger. Obviously, Jonathan didn't want to agree to that.
And so I was like, well, I need to speak my own truth. And then the community can decide if they believe me or not. But at least we'll have the full information, the full picture.
[Kristin] (25:02 - 25:09)
How difficult was this whole experience? I mean, you're starting a new job. Like you said, you didn't even have furniture in your office.
And now you're having to retract papers.
[Kate] (25:09 - 25:37)
I mean, so on the one hand, I was deeply worried and like stressed out about that. I'm like, man, I've really just tanked my career in the sense maybe I get a paycheck, but who knows if I get tenure? Who knows if I ever get invited to a seminar.
But then on the other hand, I literally didn't see any other option. It was like, I couldn't live with the idea of not retracting the papers because the data were clearly so messed up. These were not real animal behaviors.
And so on the one hand, you're deeply worried. And on the other hand, like you have no other option. That's exactly how I felt.
[Kristin] (25:38 - 25:48)
So you were very public actually about this. What was the research community's reaction to your public acknowledgement of error?
And were you surprised by that reaction?
[Kate] (25:49 - 26:41)
I mean, it was the best reaction I could have hoped for. My community of behavioral ecology is just it's a real tight knit community. People know each other.
We're all pretty friendly. And I think they just really responded so well. They were like, Kate, you're telling the truth.
Like you've shown us you're telling the truth. And we're sorry this happened. And now let's all band together and do the best we can to protect the people that need protecting, and to sort out how big of a problem this is.
And so really, it was the best response. I like get a little emotional thinking about it. Like I received some of the nicest emails from people that I don't even know, who just emailed me to say, man, you were doing the right thing.
You're this role model. You give me hope that scientists are fighting for what's good in the world. Like people brought flowers to me.
I'm forever grateful. That is not what I would have pictured at all.
[Regina] (26:42 - 26:44)
That's just beautiful. I'm emotional.
[Kate] (26:44 - 26:57)
I mean, it wasn't what I pictured. Yeah, again, I still even with the publishing of the blog post, I still thought I was going to be this social pariah. And to be proven wrong in the best possible way was truly so special.
[Regina] (26:57 - 27:01)
What was happening with the other Pruitt survivors at this time?
[Kate] (27:02 - 28:15)
Yeah. And so the whole thing blew up once I published my blog post, right? Because when I publish a blog post, clearly, there was like this insinuation that something was happening.
And I remember that day, you know, I mean, that was a viral moment where like, I don't think my phone stopped buzzing for like 72 hours. So I'm getting phone calls and texts from all my friends who've also published papers with Jonathan. And they're like, what is going on?
What is going on? And I just tell them, I was like, check your data. So they all started checking their data.
And then as soon as you start investigating that data with an eye of mistrust, they started finding all sorts of crazy things that don't make biological sense. And so yeah, so I think pretty quickly, they all were freaking out too. And now they're wanting to retract their papers because they're very clearly seeing that these data don't seem to be trustworthy in any way.
So they're all doing it. We're all trying to communicate and figure out like how bad is this problem? Like, at some point, I think a data file, there was a formula left behind that apparently created a response variable that obviously shouldn't be created by a formula.
And so yeah, like, those are ones where it's easy where someone's like, is this, is this weird? And we could be like, yes. Yeah.
[Regina] (28:17 - 28:26)
That is wrong. And did you get any work done? You're a new assistant professor.
[Kate] (28:28 - 29:11)
Yeah. No. And I mean, so to jump ahead a bit, this is only the beginning of my story. Like the publishing and the blog posts and spurred is, you know, a massive investigation. I was involved with it for two years.
The final dossier of evidence that the investigative team at Jonathan's University sent me was 292 pages long. And so when people are like, how was your first year as a new faculty member? I was like, well, I spent a lot of time writing.
But it was obviously not on things that I necessarily wanted to be writing about. So no, I got literally no work done during this time at all, just because you're a dog with a bone, you're trying to solve this puzzle. And then you're dealing with all the legal repercussions after.
[Kristin] (29:11 - 29:21)
Yeah, so legal repercussions. So pretty soon, Jonathan didn't become cooperative, I take it and lawyers got involved. So that's always scary.
Tell us what it was.
[Kate] (29:22 - 30:24)
I mean, the scariest thing I've ever been through, without a doubt. I mean, I'm so lucky. I feel like UC Davis really didn't have to support me in the way that they did.
But you know, I'm a brand new faculty, I'm unproven, but they immediately gave me access to their legal team, which was a big relief. But yeah, I mean, Jonathan is fight was fighting for his career. And he was a tenured faculty at McMaster at the time.
And so it makes sense that he would get legal representation, but getting those emails where they're using the language that they're clearly trying to, I thought intimidate me, and or indicate that they were going to maybe attempt to sue me for defamation was, I don't think I slept for like six months. I've never been more terrified. And you know, my husband and I are checking our savings account.
We're like, do we have enough money for like a legal retainer? We have all these this money for fees. And luckily, UC Davis was like, No, no, this is part of your job.
We have you like you will be represented by us if it comes to that. It never did. He never did sue me.
But I did get lots of very threatening, scary letters.
[Kristin] (30:24 - 30:33)
That must be really scary. Was there any comfort though, to knowing like, you were lock solid sure that these data were false? Did that help at all or no?
[Kate] (30:33 - 31:33)
It did help. And honestly, I remember the first time I got one of these very intimidating, scary letters that are using the words like Dr. Laskowski is acting with malice and bad intent. And those are all the words you use apparently, when they're trying to build a defamation case to you, I sent it to UC Davis lawyers, and I'm freaking out.
I'm like, Oh my god, what's happening? I just remember the lawyer I talked to on the phone, she calls me and I'm like, so what do I do about this? Is he gonna sue me?
And she goes, she just laughs. And she goes, I'd like to see him try. And it was honestly, that was the thing that I was like, okay, this is going to be fine.
Like if the lawyer is laughing about this and use it as a welcome challenge, then I should like, I didn't do it as a welcome challenge. But I viewed it as okay, someone has my back, I'm going to get through this. I'm telling the truth.
Everything I've done has been transparent. And that's what the lawyer said. She says, honey, the truth is a valid defense.
And yeah, that's the only way that I think I survived mentally.
[Kristin] (31:34 - 31:41)
And it must have felt really good that you had done all of that careful, careful documentation to have that truth with evidence.
[Kate] (31:42 - 31:56)
Yeah, exactly. Like, I've tried to give every benefit of every doubt at every step of the way of like, this could be this sort of mistake, it could be that sort of mistake, it could be this sort of mistake. And I could systematically reject every single hypothesis.
[Regina] (31:56 - 32:03)
So do I remember correctly, Jonathan tried to unretract some of your papers?
[Kate] (32:03 - 34:54)
Yeah. So it was the two papers that I was the first author on the one in American Naturalist, and then another one in the Proceedings of the Royal Society B. And so after Jonathan agreed to the retraction statements, you know, I have that documented in writing in emails, my impression is and this is again, my personal impression is that after he realized what a public deal this became, he wanted to unretract them and say they weren't that bad. And so American Naturalist, as far as I can understand, did not entertain that as a possibility, because they said, you know, we've already done our institutional investigation, these are our findings, we don't see any new evidence to suggest that we should reconsider our decision.
Proceedings of the Royal Society B was not as hard line. I had to go back and forth with them a lot, where Jonathan would present some new hypothesis, well, maybe it's like this, and I would then have to do all these reanalyses to be like, No, I can reject this hypothesis. Or he would argue that, like, there's no reason we should throw out the duplicate values, because they could be real.
And I would say, Well, until you show me a hard copy video or lab notebook, or something that proves they're real, we should treat them as not real. And so as a whole, that took a lot of time. And it was extremely frustrating, because I felt like I was having to defend my decision when in my mind, Jonathan should be having to defend the validity of these data, like the onus is not on me, right, is how I felt.
That took up several months. But I don't know if you remember. So all this went public in early 2020.
And something else happened in early 2020. And so we were all in lockdown. And I was like, Well, I have nothing else to do.
So why don't we just keep fighting these retractions?
[Regina]
Was there any point, Kate, where you felt anger?
[Kate]
Yeah, I mean, I consider Jonathan a friend, a good like work colleague friend, I applied to work with him as a postdoc, because I thought he was doing cool work.
We were drinking buddies at conferences, he wrote me letters of recommendation for my jobs. I texted him and emailed him when I got engaged, I really genuinely considered him a friend. And so to me, this was a betrayal, both a professional betrayal, but also a personal betrayal, I felt lied to and hurt.
I'm just like, he did this to me, like, but it wasn't just me, right? Like he, it appears that he did this to multiple people. And so yeah, I felt anger, I felt hurt.
Anger actually is a good way to put it at some point, I was just like, you underestimated me.
[Kristin]
I love that.
[Kate]
You think that I care more about the publication, but really, what I care about is the right answer. And I wanted to know why these spiders were doing what they were doing.
And I thought I discovered something cool. And now you're telling me that I was wrong. That pissed me off.
[Kristin] (34:54 - 34:57)
Tell us how this whole saga finally came to a close for you.
[Kate] (34:58 - 36:29)
Yeah, it was, I mean, years when I first approached Jonathan, he was at the University of Pittsburgh. And then he moved to University of California, Santa Barbara. And then he got one of these very prestigious Canada Research 150 chairs.
So he was an endowed professor at McMaster University. And so he was getting lots of research money. He was very well known, very high publicity.
And then to McMaster's credit, they retained a outside independent investigative lawyer team to do this investigation. And I emailed and zoomed with them multiple times. And I remember being like deeply impressed with how they were just doing their jobs.
I remember at one point, I had to send them all my emails between Jonathan and I of like how the data was collected and all these things. And at some point, they were like, okay, so you emailed Jonathan on this day in December. And Jonathan says that he's about to go to South Africa.
And then when he gets back, he's going to collect your data. But then he sends you the email with the data only four weeks after he returned from South Africa, but your experiment is supposed to be six weeks long. Like they had a timeline, they were mapping out what was feasible and biologically possible.
So I really got the impression that they were like doing their due diligence and really deeply trying to sort out whether or not the collection of this data was even feasible.
[Regina]
I want a job like that. It sounds fun, right?
[Kate]
I mean, talk about a dog with a bone. You're just in puzzles all day.
[Regina] (36:30 - 36:35)
Yeah. So Jonathan lost his job at McMaster. Do we know what he's doing now?
[Kate] (36:35 - 37:08)
He resigned as part of a settlement. I don't know what the details of the settlement are. It's confidential.
My understanding is he then was teaching high school biology in Florida. But most recently, I believe he's in Louisiana somewhere. I don't know what he's doing.
I don't follow him. I honestly don't want to know. But I talked to a journalist about a year ago, and she said that's where she found him.
[Regina]
I think he's written a fantasy novel. He's a novelist now.
[Kate]
Yeah, I'm told it's a self-published novel.
I think it's available for 99 cents on Kindle at the moment.
[Kristin] (37:08 - 37:09)
Well, fiction. He should be good at fiction.
[Kate] (37:09 - 37:13)
That was a long-standing joke. Yes.
[Regina] (37:13 - 37:34)
Kate, you changed the way you check data now. You talked about how you put everything in the R, and you're doing the histograms, and the box plots, and looking at everything. But you weren't able, those methods are not designed to catch problems like this.
So do you approach data differently now?
[Kate] (37:34 - 38:36)
I mean, yes and no. I actually don't know that I've collaborated so much with folks where they're purely generating data, and I'm then doing analysis. So I don't know that I've had quite the opportunity.
But something we've started doing in my lab that I actually think is a great idea, and I would recommend to other labs, is we do a data forensics day. And so what we did is everyone who has a data set that your paper is somewhere in the pipeline, we trade data sets. And we basically say, investigate this data set and try to prove that it's fake.
And what this has ended up doing now, we do this in a lab meeting. Of course, we know the data is not fake. We know we collected these data.
But the idea is we're just finding mistakes, right? And you're like, oh, look, this got copy pasted, or like, oh, there's a missing value here. And so it's just been incredibly helpful, just for mistake checking.
It's kind of fun too, because it's data sleuthing. You get to work through these puzzles and see if you can see anything weird. And it helps in terms of data management, making sure your data is clean and organized in the way you want it to be organized.
[Kristin] (38:36 - 38:42)
I love that idea. What other advice based on this experience would you have for other scientists, especially young scientists?
[Kate] (38:42 - 40:04)
I mean, the biggest thing I say is like, if nothing else, you need to have really good data management, right? And file management. I think the only reason I survived this whole thing, right, is because one, I was completely transparent about everything I was seeing.
But two, I had like reasonably good file management that I could go back and immediately find the file that I had analyzed. I had all my scripts. I had evidence and receipts of everything that I did to the data once it entered my possession.
I now teach undergrad courses, and I am shocked at the lack of file management. I just feel like these students, they just download things, and then every file ever is in their download folder. And I'm like, how do you find anything?
And so something we do in my lab is we make sure we have really good file management of like having a folder structure that makes sense. And we also do lots of version control. So we never control save data files, that's not allowed.
You always save as and then append the date. So that way, you can always go follow your breadcrumbs back. If some mistake gets entered in your data file, you don't want to overwrite it, right?
And so I think those are probably like the easiest things you can do is have proper file management, do version control. And if you ever find yourself in a weird situation, just be honest. Don't try to cover anything up.
You'd rather be considered an honest idiot than someone who's trying to cover something up.
[Regina] (40:04 - 40:21)
I love that you mentioned all of these things. Kristin, and I talk about them on the podcast. And we joke that file control, version control, file management is all quite sexy and underappreciated as an aphrodisiac.
[Kate] (40:21 - 41:11)
I agree. I think this is like I tell so like in our first year grad student thing, I like tell them this, I'm like, go to a project you did five years ago. And within five minutes, I want you to tell me what you did, show me the final data file and show me your final analysis file. And I mean, most folks can't do that unless you have good file management because you have all these different files that is like our code for this our code for that are like manuscript final manuscript final final manuscript for real final and you're like, which is it, you know?
[Regina]
But we often don't teach that in school.
[Kate]
No, no, absolutely not. I know.
It's crazy. I was like, this is the thing with all my first year grad students. I'm like, guys, I don't know how to tell you this, but like, you have to document everything because you will forget.
[Regina] (41:12 - 41:41)
I talked to a researcher once a Nobel laureate, he called it scientific hygiene. He said, it's just like you brush your teeth, you wash your hand. This is just what you do. Is it fun?
No. Do you enjoy brushing your teeth? No, but you do it anyway.
Yes. Exactly. Yeah.
So talk to me about open science, because the reason that this initially came to light is because a collaborator of yours found your data, open data posted on American Naturalist. So tell me about open data.
[Kate] (41:41 - 42:30)
I mean, I'm just such a believer now. And I think I noticed there's a generational shift as well. I think when I started grad school, I noticed that there is a bit of resistance to like depositing your data, I think because people feel, you know, they feel ownership of their data and everyone is living in perennial fear of being scooped. But at least in my field of behavioral ecology, like no one's going to scoop me.
Come on, like I study this one weird fish, no one's going to be able to do the same thing I can do. So I'm not I think it's silly. But then in a bigger sense, if you're doing federally funded research, it's getting harder to do.
But assuming you have federal funding, your data don't belong to you. They belong to the taxpayer. And so yeah, I think depositing data is critical.
And as important I think now is depositing code. I think it's like a public good. And I think we'd all be better for it.
If we can deposit more things openly like that.
[Kristin] (42:31 - 42:42)
Yeah, we wholeheartedly agree. So Kate, I'm dying to talk to you about you teach a data detective course at UC Davis. I love this.
And it's very relevant to the podcast. So what do you teach and who's it for?
[Kate] (42:42 - 43:31)
So data detectives, it's a first year seminar. So we talk about science. And we talk about how science is reported and what retractions mean.
But then we also talk a lot about how science is communicated. Like what have you heard in popular media about something about science? And can you trust that?
And how would you use your scientific reading skills to investigate that? So one activity I have them do is they go on TikTok or Instagram or whatever, and they find some influencer who's like touting some health thing of like, oh, you need to wear a weighted vest because that prevents osteoporosis. Or oh, those black utensils have all these carcinogens in them.
And so then we go and we try to find literature that either supports or refutes those claims. So my goal there is to have them, you know, be skeptical consumers of media, but ideally not fully distrust science as a whole.
[Kristin] (43:31 - 43:34)
That sounds like just what we do on this podcast, actually.
[Kate] (43:35 - 43:40)
Perfect. Well, yeah, maybe now I'll be using your podcast as some examples or something.
[Kristin] (43:41 - 43:43)
Oh, absolutely. We have some good examples. Yeah.
[Regina] (43:44 - 44:04)
Congratulations on being one of the first winners of the Control Z award. It's very exciting. Do you think that awards like this can help change the culture around errors and retractions and mistakes?
[Kate] (44:05 - 44:34)
I mean, God, I hope so. Right? I think, you know, retractions have this serious stigma around them. And I understand that because no one wants to be wrong. But my hope is that me and Alyssa and plenty of other folks at this point, who have pushed for their own retractions of their papers, like indicates that, you know, I like I don't care so much about the paper, what I care about is like actually getting the truth of the answer.
And so if I find out that, you know, a paper is based on something wrong, like, of course, I want to retract that.
[Kristin] (44:35 - 44:40)
And Kate, you've been pretty public with your story. Do you think that's made it easier for other researchers to do the right thing in the future?
[Kate] (44:41 - 45:36)
I hope so. That's sort of why I did it, you know, in the sense of I wanted to protect myself. But I also was hoping that people could use this as an example of like, you can survive this, everyone will make a mistake, mistakes are human nature.
But what matters most is what you do once you discover the mistake, right? For better for worse, I get pretty regularly emails from folks who say they're going through something similar, either they have to deal with the retraction for, you know, a totally honest reason, or they have a collaborator co author that they're a little bit nervous about how do they handle that. And so I'm happy to play that role of like a sounding board of just being like, you'll survive, just keep good documentation, be as transparent as you can.
And you know, you'll make it out the other side. And so I hope that it could be used as a good example of people can survive this. It's fine.
Mistakes are human. But it's how you respond that defines, you know, scientists or not.
[Kristin]
I love that.
[Regina] (45:37 - 45:43)
So that was Kate Laskowski, winner of the Senior Award. And now we have the Junior Award winner.
[Kristin] (45:43 - 46:04)
But first, let's take a short break. Welcome back to Normal Curves.
[Regina] (46:04 - 46:07)
And now we have the Junior Award winner, Alyssa Smith.
[Kristin] (46:07 - 46:14)
Welcome, Alyssa. We are so excited to have you here. And we would love it if you could start just by giving us a brief introduction.
[Alyssa] (46:15 - 46:44)
All right. Thank you so much for having me on here. I'm really excited to be here.
I recently graduated from Northeastern University's Network Science Institute with a PhD in, you guessed it, network science. I study the networks that people form on social media and how information flows through these networks and how attention is allocated on these networks. So I primarily identify as a computational social scientist or a network scientist, depending on who I'm talking to.
[Kristin] (46:44 - 46:46)
Great. And tell us where you are now, Alyssa.
[Alyssa] (46:46 - 46:56)
Yes, I am an assistant professor in the Department of Mathematics and Computer Science at the College of the Holy Cross, and I teach courses for the statistics and data science program.
[Kristin] (46:56 - 47:05)
Can you tell us about the paper that we're going to talk about, for which you won the Ctrl-Z Award, but at a level that someone who doesn't know network science could understand?
[Alyssa] (47:05 - 47:51)
Absolutely. I was interested in a kind of social media user that we collectively, my co-authors and I, called an attention broker. An attention broker is someone who redirects attention on social media.
So for example, if I were to repost something by Kristin and my followers were to see that repost, they go, oh wow, that's a great post. They flock to your account and then they start following you because of that repost. I am acting as an attention broker.
I'm kind of endorsing to my followers, hey, this is something that I think is worth broadcasting to you all. Maybe you should take a look. And if they do take a look, like what they see and decide to follow, I have influenced the allocation of attention on that social media platform in a very small but meaningful way.
[Regina] (47:51 - 48:19)
That's such a clever idea for a paper. So let's now get into what happened when you found the error, the thing that got you the Ctrl-Z Award. You were not yet finished with your PhD, but you were interviewing, you were on the job market.
And what was it, a week before the job talk, you noticed a mistake? So take us back to that moment. What happened?
How did you notice the error and what did you do?
[Alyssa] (48:19 - 49:08)
Okay. This was a very stressful time on multiple levels. So the paper was published in PNAS Nexus and I was on the job market.
This was, I think, late October, 2025. And I was preparing my job talk, which centered on this paper. When I was looking at a particular plot that was one of the plots to illustrate the trajectory of follower accumulation for both the reposted and the control accounts, I noticed that the error bars were smaller than I thought they should be.
And so I went back into the code to see, okay, did I make a simple mistake? How did I do this? Is there something that I'm missing that past me knew about, but current me doesn't because I forgot it?
And I realized that when I was computing, what I thought was the average, I had taken the sum and forgotten to divide by the count.
[Regina] (49:08 - 49:10)
You forgot the denominator.
[Alyssa] (49:10 - 49:12)
I forgot the denominator.
[Regina] (49:12 - 49:22)
Oh my goodness. Okay. First of all, congrats for you for recognizing that your confidence interval was too narrow, that your error bars were too small.
[Kristin] (49:22 - 49:30)
As a statistician, I have to say how much I love that, that the thing that you notice is that the error bars are wrong. Spoken like a true statistician.
[Regina] (49:31 - 49:45)
We love that. Okay. So you went back into your code then to try to figure out what happened, and that's when you saw the missing denominator.
And did your heart sink? This is pretty bad, right? What were you feeling?
[Alyssa] (49:45 - 50:49)
I was terrified on multiple levels. There was the, am I going to graduate level? There was the, am I going to be able to give a job talk level?
And there was the, is this paper still intact level? I spoke with my coauthors who I should name and thank because they were incredible supports on this whole ramshackle journey. John Green, Brooke Foucault-Wells, and David Lazar.
Brooke and David were my advisors during my PhD. And I talked to them and I said, Hey, I took the sum instead of the average. These plots are incorrect.
And they said, okay, this is how you file a correction with a journal. This is what you need to do. Here's how you write it up.
Here's what you say. So I wrote up the correction, I fixed the plots, I put the fixed plots in the job talk, and I sent the correction into the journal. And I thought that was the end of the story.
I was still incredibly embarrassed because, you know, we learn about denominators in grade school, but I wanted to make sure that I was doing the right thing and that I would not be bitten by this down the road if someone else were to find the error in the code.
[Kristin] (50:49 - 50:56)
You are probably not the first person to forget a denominator. So tell us, how did you then figure out that there was a bigger problem and tell us about that?
[Alyssa] (50:56 - 51:37)
Okay. So as someone who had never been through the corrections process, I was unaware that you were not supposed to wait months to receive a response to your correction request. So months had passed by the time I revisited this, because a lot had transpired, you know, finishing the dissertation, getting a job offer, and so on.
And I didn't realize that the correction should not be sitting in limbo for months. It turns out that it had been overlooked, which happens, you know, they're a very busy journal. When I was revisiting the correction, I looked at the plots I'd created again, and I realized that one of them, the spike in follower accumulation, was off by one.
[Regina] (51:37 - 51:38)
What do you mean off by one?
[Alyssa] (51:39 - 51:48)
By off by one, I mean that the spike should have been at t equals zero, which is the time at which the retweet happened on Twitter. But instead, it was at t equals negative one.
[Regina] (51:48 - 51:54)
Wow. So it's breaking the laws of physics, before the retweet happened. That's when you saw the spike. Okay.
[Alyssa] (51:55 - 52:01)
Yes. So unless everyone on Twitter is psychic, which I highly doubt, that should not be happening.
[Regina] (52:01 - 52:05)
Okay. So what went through your mind and your body then?
[Alyssa] (52:05 - 52:29)
A wave of fear, followed by a semi-solid belief, so about the consistency of Jell-O, that if I went back into the code, I would understand what happened, and it would be fine. When I dove back into the raw data, though, I found out that it was very much not fine. When I went into the most raw of the data that I had, the least processed data, if you will.
[Regina] (52:29 - 52:30)
And what did you see there?
[Alyssa] (52:30 - 53:15)
I saw two main problems. So I had two case studies in the paper. One of the case studies, the timestamps for the following events were simply not what they should have been, because there was an error in how I specified those timestamps in the code.
[Regina]
For just one of the case studies?
[Alyssa]
For one of the two, yes. And then I looked at the function that I had used to specify the timestamps in the first place.
These had to be very precise. So with the help of one of my colleagues, Brennan Klein, who is at the Network Science Institute, we looked into that function. We looked it up and down, pulled out its guts, and realized that the way that I had been specifying those timestamps when I was requesting the data was too imprecise to be useful for the analysis I was doing.
[Regina] (53:15 - 53:20)
So it was like an approximate time rather than an exact time?
[Alyssa] (53:20 - 53:38)
Yes. So imagine trying to cut something with a razor versus a really dull butter knife. I was using a very thick, egregiously dull butter knife, maybe even the handle of the butter knife to cut this when I thought I was using a razor, if that makes sense.
[Regina] (53:38 - 53:41)
J Yes, that's a great analogy, actually.
[Alyssa] (53:41 - 53:44)
I've had time to practice it.
[Regina] (53:44 - 53:54)
Okay, so you and your colleague then are going through the code, and you notice that you are pulling the data incorrectly.
[Alyssa] (53:54 - 54:02)
Yes. And because this was the Twitter API, and this was February 2026, there was no hope of getting the data again.
[Kristin] (54:02 - 54:05)
Explain API for those of us who are not as techie.
[Alyssa] (54:06 - 54:40)
Absolutely. An API is an application programming interface, which says if you send me a request in this particular format, I will give you data back in this specific format. The Twitter API, rest in peace, used to allow me to query the following events in a very specific way.
However, that particular way of asking the API for data is no longer working. And Elon Musk is charging a ton of money to use the Twitter API. So it was a no go in terms of getting the data back again.
[Kristin] (54:40 - 54:42)
So you could not redo the analysis?
[Alyssa] (54:42 - 54:45)
I could not redo the analysis on Twitter. Yes.
[Regina] (54:45 - 54:47)
Okay, now tell us what you're thinking at this point.
[Alyssa] (54:48 - 55:00)
I thought my career was over, and that I was a terrible scientist. I was worried that I would not graduate. I did not know what I was going to do for a non-trivial amount of time.
[Kristin] (55:01 - 55:03)
Wow, that must have been really, really stressful.
[Alyssa] (55:04 - 55:09)
Yeah, it was quite bad. I think that was by far the worst part of my PhD.
[Kristin] (55:09 - 55:10)
So what did you do?
[Alyssa] (55:11 - 56:30)
The first thing I did after freaking out, mourning my future as an academic, and deciding that I was a terrible scientist, was to decide to use the data we had at home, which was this blue sky data set that I had collected with collaborators, including Nicholas Landry at the University of Virginia, who actually nominated me for the Ctrl-Z award, and some other folks, including my advisor, Brooke Foucault-Wells, Ilya Amberg, and Sagar Kumar, who's a lab mate of mine.
We had collected this data set of blue sky data, which conveniently included the entire following network with exact timestamps for the following events. And we had released an anonymized version of this data set. So instead of having the specific account IDs, we gave everyone an integer ID, and we coarsened the timestamps to day-level precision, so that people could not be re-identified.
And that data set is publicly available on the social media archive at the University of Michigan, if you want to check that out for whatever reason. We had the non-anonymized version on the University of Virginia servers. And I decided to work with that data set because it had what I needed.
And basically, I redid the paper, but with better causal inference using the blue sky data, since we already had the data at home.
[Regina] (56:30 - 56:36)
You make it sound so tidy. I'm guessing that that was not a tidy process along the way at all.
[Alyssa] (56:37 - 56:49)
No, it was awful. I could not think clearly. I made a ton of mistakes in the early code because my brain was just not working anymore.
It was really rough.
[Kristin] (56:49 - 57:04)
Alyssa, I do think, though, that almost everyone in your situation would feel the same way. So kudos to you for just trucking through it and getting through it, because we'd all feel that way of, what is this going to do to our careers? And that's a really hard place to be.
[Alyssa] (57:04 - 57:34)
Thank you. I really appreciate that. It was rough, but I had a really good support structure around me.
So I think my advisors, as well as John Green, the third co-author, did a really good job of being forward-looking instead of focusing on what had happened, which I really appreciated. It was kind of like, how can we make this work as best we can, given that it's a terrible situation? How can we keep moving forward?
What can we learn from this? And I think I was really fortunate to have folks like that in my corner.
[Regina] (57:35 - 57:53)
I love that, because with advisors that would maybe be a little less supportive, I can imagine myself just wanting to drop out of the PhD program right there to say, bye guys, I'm going to go learn how to make coffee and do this for a living. But you kept going through it step by step.
[Alyssa] (57:54 - 58:04)
I think the step-by-step aspect was the most important. I think if I had spent very much time zooming out and thinking about the big picture of it all, I would not have been able to keep going.
[Regina] (58:04 - 58:26)
So Alyssa, I heard that you, when you were going through this process, found yourself searching Retraction Watch for stories of other scientists who had been in the same situation who had retracted their own work. Kind of like looking for role models, retraction role models. Did you find anything?
This makes complete sense to me. What did you find?
[Alyssa] (58:26 - 1:00:19)
Yes, I found a great many role models. The reason I was on Retraction Watch in the first place, the very first place, was to try to find someone who had been in a really similar boat where there were errors in the code. And I was hoping that someone was able to make a substantial correction, because what we were hoping for was, okay, we have this blue sky data, can we redo the analysis and submit a very expensive correction?
So originally, that is what I was looking for, was someone who was able to substitute new data into a paper and do a substantial correction and not have to go through the retraction process. But I found many people who went through the retraction process and were very brave about it. I still don't feel like I was brave, but I saw many people who were in what I would categorize as dissimilar boats, but same rough seas, watching out for pirates and what have you, giant squid, and getting through it.
So I saw people who were in very different circumstances, but all of whom were very honest about how unpleasant and stressful and demoralizing of an experience it was. And who also, I think, felt like they had done the right thing and would do it again. One situation that stood out to me in particular was someone named Mitch Brown, who I believe was a postdoc at the time that he was interviewed.
And he discovered some miscoded variables in some data from Qualtrics. And for whatever reason, the feelings that he described really resonated at the time. I might've cried when I read that blog post.
And the fact that he was working on a replacement article at the time was also very reassuring, because that was something that I was seeing as a distinct possibility in my future. So I think to sum up, lots of role models, no exact same boat, many people in different watercraft sailing the rough seas.
[Regina] (1:00:20 - 1:00:38)
I love that analogy, but it's true. They don't teach us this in graduate school. They teach us just how to prepare a submission, how to submit.
They don't tell us how to double check for errors like this and what to do if you found these honest errors. No one talks about this.
[Alyssa] (1:00:38 - 1:00:46)
They really don't. I think that is part of what made it so scary, was the unknown unknowns of the retraction process.
[Kristin] (1:00:47 - 1:01:00)
But Alyssa, I have to say, you are brave. So don't say that you weren't brave, because you were brave.
[Alyssa]
Thank you.
[Kristin]
I'm curious. So you ended up retracting the paper. Tell us a little bit about how the scientific community reacted.
And were you surprised by the reaction?
[Alyssa] (1:01:01 - 1:01:07)
I really did think that there would be some trolls or negative backlash, even when the award happened. But people were incredibly supportive.
[Kristin] (1:01:07 - 1:01:12)
So you already had a job offer when all of this went down. You must have been concerned about the job offer.
[Alyssa] (1:01:12 - 1:01:51)
Absolutely, yes. I was terrified that I was going to lose my job offer. So the retraction didn't come up during my job search because of how everything was timed.
But when we did eventually find out from the journal that we were going to have to retract the paper, I was terrified to tell my soon-to-be department chair. I was really worried I would lose the job. So I sent him that email.
I spent the day. I don't know if it was even a full day. Eric is very fast on email.
But Eric was incredibly supportive when I reached out. And he has continued to be a very supportive chair. So shout out to Eric Ruggieri, amazing department chair, very supportive, very reassuring, basically told me to keep my head up and keep chugging.
[Regina] (1:01:52 - 1:02:22)
This is not the picture of science and the scientific community that many people have. This sounds incredibly self-aware and supportive and human, maybe. It sounds like a very human process.
And I think a lot of people think of science and the scientific community. We just see the feuds. We don't see behind the scenes people helping out a young graduate student, new professor like yourself and saying, hey, it's going to be okay.
You're doing the right thing.
[Alyssa] (1:02:23 - 1:02:38)
Yes, people do science, but I think we forget a lot of the time that people do science. And sometimes the people who do science are unkind and vindictive. But I think more often than not, much more often than not, at least what I've experienced is that they can be incredibly kind and generous and thoughtful.
[Kristin] (1:02:39 - 1:03:00)
I think every scientist can probably put themselves in your shoes, any scientist who's ever coded, I should say, because we all know pages and pages and pages of code, errors happen, right? It's probably not uncommon. And how many just never get caught because nobody goes back and checks, right?
So I think there's a certain amount of people can put themselves in your shoes and have some empathy.
[Alyssa]
Absolutely.
[Regina] (1:03:01 - 1:03:24)
I have one question about this. So the replication that you did with the Blue Sky data, it ended up showing the same sorts of effects that you found in the original paper. How do you think things would have changed if you had done this work on the new data and then came up empty-handed, null result, couldn't replicate?
That would have been its own kind of terrifying.
[Alyssa] (1:03:24 - 1:03:49)
But I think that even if I had come up with null results in the redo of my paper, that would have still been informative. That still tells us something about how social media operates. And maybe it's because it's Blue Sky instead of Twitter.
Maybe there was never an effect all along. We can't know. But it's not no information.
It tells us something about the world. And I think that would hopefully still be valuable.
[Regina] (1:03:49 - 1:04:50)
That is fabulous. And, you know, Kristin and I being statisticians, that makes us so happy and warms our heart to hear you say, no, null results are interesting. There's still information here.
I love that you're thinking about the scientific process and data analysis in this really holistic way, not just searching for p-values so you can publish, right? You are looking at it from the big perspective. And there are so many scientists with decades more experience than you who don't have this great perspective that you have.
Thank you. So, these Ctrl-Z awards, I love them. I feel like it has helped reframe your experience maybe.
So, what do you think? If you were to have, you know, a magic wand where you could change science, do you think that it would help if we did more to make correcting mistakes something that we value rather than fear?
[Alyssa] (1:04:51 - 1:05:46)
Yes. I think one thing that I've learned from talking with collaborators who maybe haven't been through a retraction, but have been through like a very extensive correction process, that like correction, retraction, and replacement, even retraction, and tracking down an error and specifying it, that's all labor, unpleasant labor sometimes, very intellectually demanding labor. And I don't think it's labor that necessarily gets rewarded in a lot of the ways that we measure success in academia.
And I wonder if some sort of more formalized recognition for correcting an error, because that takes time away from progressing your research that you hope is actually correct and working on your teaching and doing everything else that is asked of you as an academic. I wonder if there could be more formal recognition for the amount of labor that goes into correcting these errors.
[Kristin] (1:05:47 - 1:06:01)
Yeah, I think that's a great comment. Right now, it's a lot of work. And I think that's part of the reason, not just the fear of how people will judge you, but also just the fact that it's a lot of work and it's not rewarded.
And a lot of things in academia, unfortunately, the incentives are not aligned well.
[Alyssa] (1:06:02 - 1:06:03)
Absolutely. Yes.
[Kristin] (1:06:03 - 1:06:08)
Alyssa, do you think that telling your story publicly has made it easier for other researchers to do the right thing?
[Alyssa] (1:06:08 - 1:07:00)
As someone who has just done work with causal inference, I don't think that I can necessarily claim a causal effect from me individually in a way that would be responsible. But I did actually have someone reach out to me recently to just talk through a decision that they had to make with respect to correcting the scientific record. I don't want to give any specifics or anything like that.
And it was a post hoc conversation. But having that conversation was wonderful because I got to connect with someone who had experienced something similar. And just we had the space to both talk through our thought processes and kind of reflect on what we had experienced in a space that felt like we already had this baseline of mutual understanding because we had both experienced this.
And I think the conversation was helpful for both of us. I was supposed to be the one giving advice, but I found it very helpful and reassuring as well.
[Regina] (1:07:00 - 1:07:30)
I think there's something about talking to people who have had a similar kind of experience. You are learning and you're giving help as well. So thinking ahead in the future as you perhaps mentor students of your own, how do you think that you would frame this differently to help students maybe proactively think about correcting mistakes or retractions?
[Alyssa] (1:07:30 - 1:08:29)
That's a really good question. I think that at a very three 30,000 foot level view, I would say that it's okay to talk about things that aren't going well. And it's okay to talk about making mistakes.
And it's good to acknowledge that you've made a mistake. And I think that's something that especially in graduate programs where you may not want to look vulnerable and this may go for all of academia as well. You don't want to look vulnerable.
You don't want to look like you're struggling. You may be looking happy on the surface, but actually barely treading water. And I think talking about the things that aren't going well, being honest about them is one thing that I would advise students to do and to find environments where you can actually be honest when things aren't going well.
I wish that I had reached out to some of the folks who were featured on Retraction Watch and kind of asked them how the process had gone for them just to get a sense of what the unknown unknowns were. I think the thing that I would emphasize to them there is that people can be incredibly generous when you do reach out to them.
[Kristin] (1:08:29 - 1:08:37)
Yeah, I imagine like you, Alyssa, that they would have actually enjoyed talking about the experience and sharing that experience with somebody else and the wisdom they learned from it. Yeah.
[Alyssa] (1:08:38 - 1:08:38)
Exactly. Yes.
[Kristin] (1:08:39 - 1:08:47)
So you teach some statistics. Out of this whole experience, are there things you'd be teaching in statistics now that you might not have having had this experience?
[Alyssa] (1:08:48 - 1:09:40)
So I am not actually a statistician by training by any stretch of the imagination. I am much more so on the data science side of the statistics and data science program. But I am qualified to offer coding advice, probably more since this trial by fire, and I hope that will be an adequate substitute.
So one of the things that I am about to start harping on in my intro to Python course is documentation. And one thing that I am a little bit sad that past me did not do was documenting what I was thinking as I was writing the code that was ultimately broken. So I would probably emphasize documenting what you're thinking and document as you go, even track the scripts that you're running, what it's outputting, where you're storing it, because I think we all have a tendency kind of to document our code after we've written it.
[Kristin] (1:09:40 - 1:09:41)
A hundred percent.
[Alyssa] (1:09:42 - 1:10:00)
And just kind of talk about, okay, this is the functionality, these are the variables and so on. But I do think that the history of the code is very important information, and I wish I had left myself some more primary sources in there. So I think documentation is something that I am going to be harping on even more so than I had planned to.
[Kristin] (1:10:01 - 1:10:07)
That's great. Students need to hear that because that is such an easy thing to not do well, because it takes time and we all are rushing.
[Regina] (1:10:08 - 1:10:42)
Alyssa, hearing your story, you almost did not catch these errors, right? It just so happened that you looked at the error bars and something clicked in your head and you said, wait a minute, and then digging into that, you found deeper problems. Do you have now a suspicion that there are so many errors out there that have never been caught?
Just based on your, I guess, your extrapolating from your N of 1. But how does that affect how you are looking at the published literature now?
[Alyssa] (1:10:43 - 1:11:41)
I think it makes me more cognizant of the fact that science is done by humans and humans are flawed. But I also think that humans have an to reflect on process and debrief and learn from past experiences. Because I know that there's been a lot of argument for maybe automating science more and maybe using large language models to just generate research with varying levels of human intervention.
And I hope that science continues to be done by people because people as individuals and as groups learn when we screw up, and I think also grow when we screw up. And I think that is one of the most valuable aspects of making a mistake as a scientist is that you grow and you become a better scientist. And maybe you don't become the best scientist ever, but you become a better version of yourself.
[Regina] (1:11:41 - 1:12:09)
I love that answer. And I think that you are going to be a fantastic researcher and fantastic mentor and even a spokesperson of sorts for this kind of research, your perspective about the humanity in science, you know, the humanness and what value it brings. I feel like people are looking for that more and more these days.
And you are able to articulate that so well.
[Alyssa]
Thank you.
[Kristin] (1:12:10 - 1:12:15)
Alyssa, I know you have a website where you're maybe going to talk about some of this. So we'd love to send our listeners there. What is your website?
[Alyssa] (1:12:15 - 1:12:23)
It's asmithh, so asmith with an extra h at the end, github.io. That's where you can find my online presence.
[Kristin] (1:12:24 - 1:12:53)
Thank you so much for sharing your story with us, Alyssa. This has been fabulous. Okay, Regina.
So now that we've heard from both of the Control Z award winners, let's go back and rate our claim. The claim for today was acknowledging and correcting your mistakes may have better outcomes than you expect. How do we rate claims on this podcast with our highly scientific one to five smooch rating scale, where one means little to no evidence for the claim and five means strong evidence for the claim.
So, Regina, how are you going to rate the claim today?
[Regina] (1:12:53 - 1:13:07)
I think we need to emphasize the may, may have better outcomes, but definitely five smooches. If this end of two stories that we've heard today are any indication, it sounds like it could go way better than expected. Yeah.
[Kristin] (1:13:07 - 1:14:00)
Admittedly, this is really anecdotal evidence that we have, but I totally believe this own up to your mistakes. You're going to be better off in the long run. So five smooches for me, but this is making me think, Regina, you know, someone could do a study and actually test this claim, right?
[Regina]
It's an empirical question.
[Kristin]
You could compare scientists who discover a serious problem and they own it like Kate and Alyssa did. We could compare them with some other scientists that we may have even talked about on this podcast before who find that they've published something incorrect, but they double down.
And then we could follow them out five or 10 years and measure how their careers went, some kind of professional reputation metric. And actually, Regina, I think this might even be kind of like a network science or a behavioral topic. Maybe one of our inaugural winners could run this study.
[Regina] (1:14:01 - 1:14:11)
I like it. I think this could be interesting. We need to quantify the effect that doubling down has.
How big is the effect size?
[Kristin] (1:14:11 - 1:14:35)
I want to mention the stories we heard today, they contrast so much with the story we heard in the last episode on statistical power in which we talked about those authors who doubled down on post-hoc power. It's such a contrast, but I know who comes out looking better. All right, Regina, I feel like the methodological morals write themselves for this episode, but what do you have?
[Regina] (1:14:35 - 1:14:41)
I'll go with this one. Science is self-correcting, which means scientists need to be self-correcting too.
[Kristin] (1:14:41 - 1:15:02)
Oh, I love that. Can I have two? Of course.
Okay. How about this one, which was mentioned by both of our guests today? Good documentation is boring--right up until it saves you.
But I also want to put a shout out to the Ctrl-Z awards and Retraction Watch. So if you want scientists to correct their mistakes, reward them for correcting their mistakes.
[Regina] (1:15:04 - 1:15:06)
It's all about the incentive system, isn't it?
[Kristin] (1:15:07 - 1:15:13)
Yeah, it is really. And we love the Ctrl-Z awards. We will interview next year's winners if we're still around by then.
[Regina] (1:15:14 - 1:15:29)
Please support us. Positive thinking, Kristin.
[Kristin]
Please support us so that we can be around next year.
[Regina]
This has been so much fun, Kristin. Thank you. Thank you to our guests.
[Kristin]
Yeah, our guests were both great. Thank you, Regina. Thanks everyone for listening.
RSS Feed