Statistical Power: When is it useful?
Statistical power is one of the most important ideas in statistics—and one of the most misunderstood. In this episode, we explain what statistical power really means, why researchers need to think about it before they collect their data, and what goes wrong when they calculate “post hoc power” after the study is over. Along the way, we hunt for UFOs, test Kristin’s psychic abilities, explore the ethics of underpowered studies, and discover why post hoc power is really just a p-value wearing a fake mustache. Then statistician and powerlifter Andrew Althouse joins us for the inside story of a years-long scientific feud over post hoc power, where we discuss what happens when being wrong about statistics turns into doubling down, angry letters to the editor, and debates about how science corrects itself.
Statistical topics
- Effect size
- False negatives
- P-values
- Post hoc power
- Power calculations
- Research ethics
- Sample size
- Scientific controversy
- Sensitivity analysis
- Statistical power
- Statistical significance
- Study design
- Type II error
Methodologic Morals
- “Big questions need studies powerful enough to answer them.”
- “Post hoc power is circular reasoning masquerading as new information.”
- “Know when to cut your losses. A mistake gets more expensive every time you double down.”
References
- Althouse AD, Chow ZR. Comment on “Post-hoc power: If you must, at least try to understand.” Ann Surg. 2019;270:e78-e79.
- Althouse AD. Post hoc power: Not empowering, just misleading. J Surg Res. 2021;259:A3-A6.
- Altman DG. Statistics and ethics in medical research: III. How large a sample? BMJ. 1980;281:1336-1338.
- Bababekov YJ, Hung YC, Hsu YT, et al. Is the power threshold of 0.8 applicable to surgical science?—empowering the underpowered study. J Surg Res. 2019;241:235-239.
- Bababekov YJ, Stapleton SM, Mueller JL, et al. A proposal to mitigate the consequences of type 2 error in surgical science. Ann Surg. 2018;267:621-622.
- Chang DC, Stapleton SM. Response: The proliferation and misinterpretation of “as safe as” statements in surgical science: A call for professional discourse to search for a solution. J Surg Res. 2021;259:A12-A15.
- Freiman JA, Chalmers TC, Smith H Jr, et al. The importance of beta, the type II error and sample size in the design and interpretation of the randomized control trial: Survey of 71 negative trials. N Engl J Med. 1978;299:690-694.
- Hoenig JM, Heisey DM. The abuse of power: The pervasive fallacy of power calculations for data analysis. Am Stat. 2001;55:19-24.
- Nuzzo RL. Statistical Power. PM&R. 2016;8(9):907-912. doi:10.1016/j.pmrj.2016.08.004
- Nuzzo RL. Post hoc Power. PM&R. 2021;13(4):422-424. doi:10.1002/pmrj.12476
Kristin and Regina’s online courses:
Demystifying Data: A Modern Approach to Statistical Understanding
Clinical Trials: Design, Strategy, and Analysis
Medical Statistics Certificate Program
Epidemiology and Clinical Research Graduate Certificate Program
Programs that we teach in:
Epidemiology and Clinical Research Graduate Certificate Program
Find us on:
Kristin - LinkedIn & Twitter/X
Regina - LinkedIn & ReginaNuzzo.com
- (00:00) - Intro
- (00:43) - What is Statistical Power?
- (03:22) - The UFO Telescope Analogy
- (06:19) - The Psychic Ability Experiment
- (10:36) - Planning Power Before You Study
- (12:46) - History and the Landmark 1978 Paper
- (17:42) - Enter Post-Hoc Power: The Problem
- (19:26) - Three Problems with Post-Hoc Power
- (24:33) - The Surgery Journal Disaster
- (29:59) - Guest Interview: Andrew Althouse
- (32:30) - The 2018 Paper and the Controversy Begins
- (39:54) - Doubling Down and the Cost of Being Right
- (53:03) - Rating the Claim
00:00 - Intro
00:43 - What is Statistical Power?
03:22 - The UFO Telescope Analogy
06:19 - The Psychic Ability Experiment
10:36 - Planning Power Before You Study
12:46 - History and the Landmark 1978 Paper
17:42 - Enter Post-Hoc Power: The Problem
19:26 - Three Problems with Post-Hoc Power
24:33 - The Surgery Journal Disaster
29:59 - Guest Interview: Andrew Althouse
32:30 - The 2018 Paper and the Controversy Begins
39:54 - Doubling Down and the Cost of Being Right
53:03 - Rating the Claim
[Kristin] (0:00 - 0:04)
You don't want to find yourself in an awkward situation where you haven't planned ahead.
[Regina] (0:05 - 0:11)
Then it's too late. You don't want to be caught unprepared and just same thing in statistics.
[Kristin] (0:17 - 0:26)
Welcome to Normal Curves. This is a podcast for anyone who wants to learn about scientific studies and the statistics behind them. I'm Kristin Sainani.
I'm a professor at Stanford University.
[Regina] (0:26 - 0:32)
And I'm Regina Nuzzo. I'm a professor at Gallaudet University and part-time lecturer at Stanford.
[Kristin] (0:33 - 0:37)
We are not medical doctors. We are PhDs, so nothing in this podcast should be construed as medical advice.
[Regina] (0:38 - 0:43)
Also, this podcast is separate from our day jobs at Stanford and Gallaudet University.
[Kristin] (0:43 - 0:51)
Regina, today we're going to unpack the concept of power.
[Regina]
Like Wonder Woman and Superman.
[Kristin]
Not exactly. Actually, statistical power
[Regina] (0:51 - 1:02)
Which is a little less exciting. But, Kristin, statistical power is one of the most important ideas in statistics.
[Kristin] (1:03 - 1:19)
It is. And to liven things up, we have some interesting stories to tell with this one. Later in the episode, we will talk with a researcher who has been involved in a juicy statistical power controversy.
And he also happens to be a power lifter. So I think he's perfect for this episode.
[Regina] (1:20 - 1:41)
Perfect. And today we will explain not only what vanilla statistical power is, but we are also going to dive into one of the most persistent statistical fallacies in science and medicine. And that is a different flavor of statistical power. Something called post-hoc power or observed power.
[Kristin] (1:42 - 1:48)
Regina, post-hoc power is one of those ideas that's been criticized by statisticians, shouting into the wind, for decades.
[Regina] (1:49 - 1:57)
Journal reviewers and editors, apparently they have not gotten the memo about this though. And some of them just won't let it go.
[Kristin] (1:57 - 2:04)
We'll see just how dug in people can get about this issue and what problems that brings with it.
[Regina] (2:04 - 2:20)
And for only the second time on this podcast, we are going to have a special guest interview, Andrew Althouse, who is a distinguished statistician at Medtronic. And he played a vital role in that saga.
[Kristin] (2:20 - 2:50)
And Regina, distinguished statistician is actually his job title. And I love that title. I want that title.
And I may have to switch jobs just to get that title. Our conversation with him was a lot of fun. So stay tuned for that.
And Regina, since we always have a claim that we rate at the end in these episodes, it's a little harder with a methodological episode, but I decided to make the claim today about post-hoc power. So here's the claim. Post-hoc power is never useful.
[Regina] (2:51 - 2:55)
Never useful. This is a bold claim, actually.
[Kristin] (2:55 - 3:22)
It is. And for any listeners who are already familiar with the concept, think about how you would rate that. And we'll return to this claim at the end of the episode. Regina, you and I switched off writing a statistics column for over a decade for the journal Physical Medicine and Rehabilitation.
And you wrote two articles for that column on statistical power. So why don't you start us off by explaining statistical power with one of your great analogies?
[Regina] (3:22 - 3:32)
Why, thank you, Kristin. So this is a fun one, not one I used in the columns. Statistical power, Kristin, is like looking for UFOs in the night sky.
[Kristin] (3:32 - 3:33)
Oh, UFOs. I love it.
[Regina] (3:34 - 3:39)
UFOs. So, Kristin, these UFOs do not come down to visit us.
[Kristin] (3:39 - 3:39)
Oh.
[Regina] (3:39 - 4:06)
Just want to be clear. They stay up in space and some are brightly glowing. Maybe these are the big motherships and some UFOs are dimmer.
Maybe they're the little scouting ships. And these little scouting ships are harder to see, of course, because they're dimmer. And in order to reliably see the dimmer UFOs, we need to spend more money to get a more powerful telescope.
[Kristin] (4:07 - 4:12)
I see where you're going with this analogy. Statistical power is like the power of the telescope.
[Regina] (4:12 - 4:46)
Excellent. And more powerful telescopes can detect dimmer objects better, but at the same time, they can also easily catch those big, glowing motherships. Of course. So, Kristin, you and I, let's say, are looking for UFOs, a side job, and we need to decide ahead of time what counts as worth detecting. We need to think about how dim of a UFO we want to be able to spot, how subtle of a UFO, if a UFO is really in the sky.
[Kristin] (4:47 - 4:54)
Right. We hope there's a UFO out there to find in this hypothetical scenario, but we need to remember there may not be.
[Regina] (4:54 - 5:17)
Okay. So, Kristin, let's say our goal is to spot the dim little scout ships, but we buy a cheap telescope that can only reliably spot something as bright as the moon. We don't have a lot of money.
So, this is a waste of our money, though, because we really don't have much of a chance of finding those dim ships that we're looking for.
[Kristin] (5:17 - 5:41)
Right. So, bringing it back to statistical power, in research, some effects are big and easy to detect, while others are much smaller. And if we want to detect those smaller effects, we need a study that's powerful enough to find them.
Otherwise, we risk doing a study that's kind of like the cheap telescope. It's unlikely to find the effects we are looking for, and so is essentially a waste of time and money.
[Regina] (5:41 - 6:02)
Right. Now, in our UFO world, we increase power by making the telescope bigger, and in the stats world, we increase power by making the study bigger. So, a bigger telescope means we gather more light into it, and a bigger study means we gather more observations, right? We increase the sample size.
[Kristin] (6:02 - 6:19)
Right. And increasing the sample size makes it more likely that we'll find the effects we're looking for. Regina, we do need to give the technical definition, so let me do that here.
Statistical power is the chance that our study will detect an effect of a certain size if it's actually there to be found.
And by detect an effect, we mean find an effect that is statistically significant, which is typically defined as a p-value less than 0.05.
[Regina] (6:19 - 6:39)
That technical definition, Kristin, is complicated and unsatisfying, so let's use a fun example to illustrate statistical power here.
[Kristin] (6:39 - 6:55)
Regina, maybe we can stick with the same fun example we used in our p-value episode. And by the way, the p-value episode is a really good episode to listen to before this one. And in that episode, we tested whether I am psychic, and spoiler alert, we found no evidence that I'm psychic.
[Regina] (6:56 - 7:29)
It was very sad, I will be honest with you. But Kristin, let's now imagine we're in a world in which you really are a psychic. But we don't know it, so we're designing an experiment to try to detect your psychic ability.
And we decide on this. I, in Washington, D.C., will flip a coin every day for a month, and you, in California, need to divine with your psychic abilities whether I got a head or a tail on each day. And after 30 days, one month, we count up how often you were right.
[Kristin] (7:30 - 8:06)
Oh, I want to try this now. And report back. All right, Regina, let's say for illustration that I am psychic, but my psychic ability is kind of like a dim scouting UFO ship.
It exists, but it's not that bright. I'm only slightly psychic. I don't get the right answer every time, just a little more often than chance would predict.
So a normal person guesses right 50% of the time, and let's suppose that I get it right 60% of the time. That means we would expect a normal person to, in our experiment, get it right 15 days out of the 30. But for me, it would be about 18 days out of the 30.
[Regina] (8:06 - 9:21)
Right. Kristin, your psychic abilities here are so small and puny that it's going to be hard for us to statistically detect your ability with our experiment that is just one month long. We can do the calculations and see that you would need to get at least 20 days right out of 30 for the results to be statistically significant, which would mean that it would have a p-value of less than 0.05. And that significant p-value we want because we need that in order to publish and declare your psychic abilities to the world. But Kristin, we can calculate how likely you are to clear that bar of 20 correct days out of 30. And with your paltry 60% hit rate, you'd expect, like you said, to get only 18 right, not 20. And we do the math and we figure out that you have just a 29% chance of clearing that statistical significance bar of getting at least 20 right.
So even though you really do have psychic abilities, genuine abilities, well, at least in this universe, chances are our experiment will miss it and we'll end up with a null result.
[Kristin] (9:21 - 9:46)
And that 29%, that is the statistical power of our study. Put another way, if we repeated our month-long experiment 100 times, we would detect my true psychic ability only about 29 times out of 100. The other 71 times we would miss it completely.
A power of 29% is just too low. It's not worth running an experiment that's more likely than not to miss my psychic ability.
[Regina] (9:46 - 10:08)
That would be a waste of time and money. So we need a stronger telescope. We can't make your psychic ability stronger, just like we can't make our UFO brighter.
It would be nice, but it is what it is. We can't do that. But we can make our study more powerful enough to detect it by collecting data for more than 30 days.
[Kristin] (10:09 - 10:35)
Right. And this is why we need to think about statistical power before we do this study so that we don't end up doing an underpowered and possibly useless study. During study planning, we have to decide how much power we want, and it should be a lot higher than 29%.
Typically, researchers aim for 80% power, although that's just by convention. And then once we know the power we want, we work backwards to figure out the sample size needed to achieve that power.
[Regina] (10:36 - 10:42)
And to do that, of course, we can use math or statistical software or computer simulations.
[Kristin] (10:42 - 10:56)
Right. And for our psychic ability experiment, I can work out that to achieve 80% power, we'd need to collect 158 days of data. In other words, we would need to run the study for over five months rather than just one month.
[Regina] (10:56 - 11:11)
Now, Kristin, this calculation, of course, was easy here because in our imaginary world, we knew the true extent of your psychic ability, your 60% hit rate. But in a real study, we would not know that.
[Kristin] (11:11 - 11:32)
Exactly. If we knew that already, we would not need to do the study. Right.
So when we're planning our study, we have to choose an effect size to use in our power calculations. We can either base that on previous research or we can choose the smallest effect we'd actually care about detecting. And Regina, this is part of the art of sample size and power calculations.
[Regina] (11:32 - 11:46)
That's a great way of putting it. Art. Yes.
Art. So these calculations should be a standard part of planning a study. Grant reviewers and journal editors and peer reviewers usually expect to see them.
[Kristin] (11:46 - 12:04)
And Regina, we'll put more details about power calculations and the ingredients that go into them in our show notes for those who want more technical details. Right. I want to take a minute now, though, and go back in time and talk about the history of statistical power because power calculations were not always standard in research.
[Regina] (12:05 - 12:46)
Kristin, I think we need to maybe go back to the 1930s, 40s, 50s because that's how far statistical power goes back. And it kind of started with two statisticians named Neyman and Pearson. And students and listeners might remember those names from our p-value episode because they were big shots in early stats history.
And they're the ones that basically took the ideas that had just been floating around in statistical theory and gave them names like statistical power. And they sort of formalized the whole framework that we use today. But everyone should go listen to the p-values episode for more history.
[Kristin] (12:46 - 13:33)
That's a great episode to listen to.
Today, Regina, most grant applications ask about statistical power, but that wasn't always true. John Bland, who is a well-known medical statistician, he famously said that when he started his career back in 1972, quote, little was heard of power calculations. So they were not routinely used 50 years ago.
This started to change in the late 1970s with a landmark paper in the New England Journal of Medicine. And Regina, do you know the paper I'm talking about?
[Regina]
Yeah, but you might need to remind me of some of the details.
[Kristin]
You don't have it memorized?
[Regina]
I don't have it memorized, no.
[Kristin]
Okay.
The authors looked at 71 randomized trials that had all found no significant difference between treatment and control. And then they tried to work out something important. Did these studies actually have enough statistical power?
[Regina] (13:34 - 13:43)
Of course, it was too late to do a power calculation before the trials were run because this was after they were run and they did not have access to a time machine.
[Kristin] (13:43 - 13:58)
Right. But they asked a hypothetical question that could have been asked before the trials ever began, which is what statistical power did the study have to detect a 25% benefit in the treatment group compared with the control group?
[Regina] (13:58 - 14:11)
Right. 25%. It's kind of arbitrary, but the idea is the authors thought that an effect of this size was where it started to be large enough that we definitely wouldn't want to miss it in our results.
[Kristin] (14:11 - 14:46)
Exactly. And what they found is that only five of those 71 studies had 80% power or more to detect an effect that big, meaning only five studies were big enough to provide adequate power.
[Regina]
Which is shockingly low.
[Kristin]
Very low, yes. And they even redid the calculations using a bigger effect size, a 50% treatment benefit, which is of course easier to detect like a brighter UFO, even with an effect this big, only 22 out of 71 or about a third of the studies had at least 80% power.
[Regina] (14:47 - 15:11)
Kristin, if I remember correctly, this paper was a bit of a call to action and it's still actually a classic because it exposed the fact that many randomized trials were using sample sizes that were just too small. So these randomized trials were actually not likely to be able to detect effects that were clinically meaningful. That's bad.
[Kristin] (15:13 - 15:52)
That is bad. And there was another paper that was also kind of a famous call to action that you're probably familiar with.
This was Doug Altman's 1991 paper on power in the British Medical Journal. And he argued that power is not just a statistical issue, it's actually an ethical one. If people volunteer for clinical trials, we owe it to them to conduct studies that have a genuine chance of answering the scientific question.
So over the next decade after this paper, that way of thinking became standard. And by the late 1990s, leading medical journals increasingly expected scientists to perform sample size calculations ahead of doing their studies.
[Regina] (15:53 - 15:57)
They expected them, but that doesn't mean the scientists always did them very well.
[Kristin] (15:58 - 16:06)
Definitely not. We read, Regina, lots of papers with power calculations that we think are complete nonsense.
[Regina] (16:06 - 16:06)
They're kind of pro forma, right?
[Kristin] (16:07 - 16:42)
Yes. It's become a checkbox. And I think people just often invent calculations, throw something in just to check the box off, or just come up with numbers that give the sample size they already had in mind. And really, these calculations are not exact.
They require a lot of assumptions. That's why I said it's kind of an art. They require assumptions about like, what effect size do we think is important?
And what variability do we expect in the data? And I really like the way Darren Dahly, a statistician at University College Cork, put it in a recent talk. He called it sample size theater.
[Regina] (16:43 - 16:45)
Which cracks me up because I'm picturing puppets.
[Kristin] (16:48 - 17:07)
Well, I think that's a really good visual, Regina, because many sample size calculations that I see in papers are kind of at about the level of children's theater in terms of sophistication. And Regina, bad power calculations kind of help set the stage for what we are going to focus on next, which is the problem of post-hoc power.
[Regina] (17:08 - 17:41)
And we will get to that and to describing our juicy scientific feud after the break. Welcome back to Normal Curves. Today we're talking about statistical power.
And we were about to talk about the concept of post-hoc power and a fascinating case study in a surgery journal.
[Kristin] (17:42 - 17:53)
So far, we've talked about what power is supposed to be. Before you run the study, you pick an effect size that you care about and calculate how many participants you need to reliably detect that effect if it's there.
[Regina] (17:53 - 17:57)
Right. Statistical power is essentially a planning tool.
[Kristin] (17:58 - 18:02)
It's like asking, how much gas do we need before we leave on this road trip?
[Regina] (18:02 - 18:13)
Or Kristin, do you mind if I bring this section here? Oh, go for it. It is like asking before you go on a date, do I need to pack prophylactics?
[Kristin] (18:15 - 18:21)
Yeah. You don't want to find yourself in an awkward situation where you haven't planned ahead. Then it's too late.
[Regina] (18:22 - 18:34)
You don't want to be caught unprepared. And just same thing in statistics. You don't want to be caught unprepared in your study because once you've collected your sample, then it's too late.
[Kristin] (18:36 - 18:50)
Thank you, Regina, for bringing sex into this episode. You're welcome. But there's a weird perversion of statistical power that is pervasive in the literature.
And that brings us to post-hoc power, which is also sometimes called observed power.
[Regina] (18:50 - 19:01)
Right. And post-hoc power is when people run their study, get their results, and then calculate power afterwards based on those results, which is not the right order.
[Kristin] (19:02 - 19:10)
Exactly. Basically, researchers perform a power calculation using the study's sample size and the exact effect size observed in the study.
[Regina] (19:11 - 19:25)
Yeah. Which may not seem so horrible on the face of it. So Kristin, you and I will try to explain why it actually is kind of horrible.
And there are lots of problems. We're just going to focus on three of them. So Kristin, start with problem number one.
[Kristin] (19:26 - 19:30)
Right. The first issue is that post-hoc power does not actually tell us anything useful.
[Regina] (19:32 - 20:00)
That is it, pretty much right there. The same ingredients that go into your power calculation are the exact same ingredients that go into calculating a p-value. So it means that your post-hoc power is just another way of stating the p-value.
The post-hoc power is just recycling the same information as what you got in your study. It feels like you've learned something new, but you really haven't. It's all the same stuff.
[Kristin] (20:01 - 20:04)
Post-hoc power is the p-value in disguise.
[Regina] (20:04 - 20:11)
Oh, I like this. I'm imagining a p-value wearing a fake mustache and those fake glasses.
[Kristin] (20:12 - 20:16)
I'm picturing this p-value wearing the glasses and the mustache as part of our puppet show now.
[Regina] (20:17 - 20:19)
I think we've got a theme.
[Kristin] (20:19 - 20:38)
So Regina, I don't think a lot of people realize this, but there is a one-to-one relationship between the p-value and post-hoc power for a given significance threshold and statistical test. So if you tell me the p-value without me knowing anything else about the study, I can tell you the post-hoc power. It's almost like I'm psychic.
[Regina] (20:38 - 20:40)
Yeah, but it's really just math.
[Kristin] (20:40 - 21:16)
Let me give some examples. For these, I'm assuming the most standard statistical test and significance threshold of less than 0.05. So if your study finds a p-value of 0.8, highly non-significant, your post-hoc power will be 6%. If your p-value is 0.5, your post-hoc power is 10%. If your p-value is 0.2, post-hoc power is 25%. If your p-value is 0.05, right at the significance threshold, post-hoc power is 50%. And if your p-value is 0.005, highly significant, your power is 80%. It's always the same. You don't learn anything new.
[Regina] (21:16 - 21:50)
This is really profound, Kristin, and it completely blew my mind once I realized this, right? Your p-value needs to be 0.005 before your power is 80%. People do not realize this, but then once you say it and you think about it and it really gets in, it becomes obvious and it really changes the way you see the universe, the way you think about post-hoc power.
And we will put more details in the show notes for people who want to go deeper into all of those interconnections.
[Kristin] (21:50 - 21:59)
The second problem with post-hoc power is more a philosophical sticking point that maybe only statisticians like us fuss over, but it is important.
[Regina] (21:59 - 22:27)
Right. It is important, but it is kind of fussy. The problem is that statistical power is a probability.
It's a probability of something happening in the future. It's a before-the-fact concept. It's the probability before you run a study that you will detect a true effect of a certain size if one exists.
Once the study's over, though, it's already happened. Things that have happened do not have a probability associated with them.
[Kristin] (22:27 - 22:33)
Right. After the study is done, it either got a significant result or it didn't. There is no probability.
[Regina] (22:34 - 23:10)
I think about it with coin flips, right? Before you flip a coin, there's a 50% chance it'll land heads. But then you flip it.
After you flip it, it's done. It either landed heads or it didn't. There's no longer a probability attached to that outcome.
And it's the same way. Statistical power only has meaning before a study is conducted. Once the study is complete, the meaningful questions are just about what the data tell us, not about the probability that we would have found an effect if we had a time machine.
And it just doesn't make any sense.
[Kristin] (23:10 - 23:13)
Maybe the UFOs have a time machine that we can use.
[Regina] (23:14 - 23:16)
And it's all coming together.
[Kristin] (23:17 - 23:49)
All right. The third problem with post-hoc power that I want to talk about today is that people often use it as empirical evidence to try to support arguments that it can't actually support. For example, post-hoc power often comes up when a study finds a non-significant result.
People find a non-significant result. They then calculate post-hoc power, which, by the way, will always come out to be below 50% when the result is non-significant. But then they say, look, I found low power, so the reason I failed to find anything is because my study didn't have enough power.
[Regina] (23:50 - 24:03)
They're trying to use it as a get-out-of-jail-free card. The implication is that, OK, well, if they just had a bigger sample, they sure would have found the significant effect. Uh-huh.
It's not their problem.
[Kristin] (24:04 - 24:21)
Right. But, of course, the logic here, it's completely circular, right? The low post-hoc power is not independent empirical evidence that these studies were underpowered.
It's a mathematical consequence of choosing studies with non-significant results in the first place.
[Regina] (24:22 - 24:33)
It's a subtle point, but it's this circular logic. So that leads us, I think, Kristin, to the case study in the Surgery Journal, which basically did something exactly like this.
[Kristin] (24:33 - 25:00)
Yes. This was a 2019 paper in the Journal of Surgical Research. The authors reviewed over 400 papers from top surgical journals published between 2012 and 2016, and then they narrowed that pile down to 69 studies that reported non-significant results for the primary outcome. They then calculated post-hoc power for those 69 studies.
And, Regina, do you want to take a wild guess here at what they found?
[Regina] (25:01 - 25:06)
Well, they had to find low power in all 69 non-significant studies.
[Kristin] (25:06 - 25:13)
It's like you're psychic. Yes. The median post-hoc power among those studies was only 16 percent.
[Regina] (25:13 - 25:21)
Which is low. Right. And exactly what we would expect, because, again, all the studies were non-significant.
[Kristin] (25:21 - 26:03)
Exactly. Since they only selected non-significant papers, post-hoc power is mathematically guaranteed to be low. But here's what kind of blows my mind, Regina. They put all this time and energy into this study, right?
They read all these papers from the literature. They went through them in great detail, extracted all these numbers, calculated things. And it was a complete waste of resources and effort because the catch is they did not need to do any study at all.
If I tell you that I have 69 non-significant studies like I just did, we don't need to do anything further. We already know mathematically, without ever reading a single one of these papers, that the median post-hoc power will be considerably less than 50 percent.
[Regina] (26:03 - 26:15)
Right. So this study is non-sensical, basically. Or I guess, Kristin, it's overly-sensical, maybe, because it tells us absolutely nothing.
[Kristin] (26:15 - 26:48)
Exactly. So here we have this study that didn't need to be done. But it gets worse than that because now they are acting as if they have independent empirical evidence that they can use to make an argument. And what they did here is to say, look, at all these studies in the surgical research, whoops, the median power was only 16 percent.
So that's evidence that maybe we shouldn't be targeting 80 percent power. Maybe that's just an unrealistic expectation for surgical research. And, of course, 80 percent power is kind of the convention that most people use.
[Regina] (26:48 - 27:00)
Right, right. So this is actually a reasonable thing to discuss, right, whether 80 percent is always the right planning target. But the problem is it's just completely unrelated to what they did in the paper.
[Kristin] (27:00 - 27:09)
Right, right. Yes, I mean, 80 percent is a convention. It is a rule of thumb.
And we probably shouldn't just arbitrarily always be choosing it. We can have that discussion. But it's completely unrelated to what they did in the study.
[Regina] (27:09 - 27:34)
So them saying this is a little like me going into my university school records and I look at all the students with a low GPA and only those. And then I come out and conclude that our university standards for all students are just too high. Well, of course, because I only looked at the ones with a low GPA.
It's like the conclusion is baked into the selection process.
[Kristin] (27:35 - 27:40)
Exactly. Their post hoc power calculations provide no evidence to support the argument that they're trying to make.
[Regina] (27:40 - 27:46)
Kristin, since post hoc power is the wrong tool here, let's talk about what they could have done that would have been useful.
[Kristin] (27:46 - 28:40)
Sure. I mean, they could have done what the New England Journal of Medicine, that 1978 study that we talked about earlier, what they did. They could have calculated how much power these studies had to detect different clinically meaningful effect sizes, such as a 25 percent benefit or a 50 percent benefit.
And this is a little subtle, Regina. So I just want to clarify for listeners who might be thinking like, wait, didn't the New England Journal of Medicine paper do the same thing? Sixty nine non-significant studies, seventy one non-significant studies calculate power.
But no, what they did in that New England Journal paper was not the same. They did not use post hoc power. They did not plug into the power calculations effects found in those studies.
Rather, they did calculations that could have been done before the study was run. They said, you know, if the effect size was a 25 percent benefit or 50 percent benefit, what would the power have been? It's subtle, but it's totally different.
So they could have done that here.
[Regina] (28:41 - 29:08)
Or Kristin, they could have done something that we call a sensitivity analysis. And that is basically asking, OK, given my actual sample size, what is the smallest effect that my study can reliably detect? And is that effect worth caring about?
That's like the UFOs, right? What is the dimmest UFO I can see given the telescope that I already bought? And do I even care about that UFO?
[Kristin] (29:09 - 29:38)
Yes. And we'll put more details about sensitivity analyses in the show notes for anybody who wants to go deeper. All right, Regina, I think we have set this up well.
And now it's time to get to the behind the scenes juicy science story, because this paper turns out to be part of a multi-year scientific feud. And this is tabloid worthy stuff for scientists. And one of the most prominent critics at the center of it all was our upcoming guest, distinguished statistician Andrew Althouse.
[Regina] (29:38 - 29:42)
And we will play our interview with Andrew after the break.
[Kristin] (29:59 - 30:09)
Welcome, Andrew. Before we jump into the case, can you tell us a little bit about your background? And by the way, I love your title.
You are a distinguished statistician.
[Andrew] (30:09 - 31:14)
Yeah, I have to admit, I laughed when I moved from the academic side of things to the industry side. On the industry side, I got hired as a senior principal statistician. And then when it came time to promote me, the title of distinguished statistician really made me laugh.
I mean, I work from home by sweatpants. The word distinguished doesn't really feel like it should apply. But anyway, how did I get there?
I've been working as a collaborative statistician in medicine since graduate school. I got a master's degree in applied statistics and a doctoral degree in epidemiology, both from the University of Pittsburgh. And after that, I joined the faculty at the University of Pittsburgh School of Medicine from 2013 to 2022, working with a number of different researchers and physicians.
And in 2022, I got an opportunity to test the waters of industry, the big bad wolf of industry. I'd kind of seen what it was like working on the academic side. And I thought it'd be interesting to see what was it like to work in clinical research and particularly clinical trials on the industry side.
So for the past four years, I've worked for Medtronic, a company that's a medical device maker, best known really for pacemakers. I work primarily on artificial heart valves.
[Kristin] (31:15 - 31:25)
Oh, cool. And Andrew, I've also read that you are a power lifter and can lift like 600 pounds. And I feel that that is relevant to this episode because you know, this is an episode on power.
[Andrew] (31:26 - 31:37)
Yeah, it's been a hobby for a while. I played football when I was younger, and I've kept lifting weights. And I mean, right now, it's really just a way to get the aggression out, especially when we start talking about episodes like this, you got to have some place to let it out.
That's mine.
[Regina] (31:38 - 31:49)
All right. All right. Well, Andrew, speaking of aggression, you did play a major role in the controversy that we're talking about today.
So how did you get involved in it in the first place? What was the start of the saga?
[Andrew] (31:50 - 32:18)
Oh boy, this takes me back. You know, Twitter used to be a place where we really had a lot of active conversations about science and statistics. And, you know, it used to be a really vibrant place.
Unfortunately, that's kind of changed the last few years. But how I got involved is really through Twitter. I started seeing this article coming up a lot on Twitter.
And you know, some of my physician colleagues, they didn't think something was quite right about this. That's kind of what drew my attention. There were people who were like, hey, wait a minute, this doesn't look right.
And I looked at it and I agreed. I didn't really think it looked right.
[Kristin] (32:18 - 32:29)
So Andrew, the same group, we already introduced the 2019 paper, but you're going back a little further. And they actually started this dialogue back with a paper in 2018. So we're going to start with that.
[Andrew] (32:30 - 32:48)
Yeah. That first paper that this group wrote was called A Proposal to Mitigate the Consequences of Type 2 Error in Surgical Science. And I was like, all right, you know, sounds like a noble enough idea.
With statistical power, we do worry about type 2 error. We don't want to falsely conclude that there's no effects in a particular situation just because the study was too small.
[Kristin] (32:49 - 33:01)
Right. Just as a reminder, type 2 error, those are the false negatives. And the authors here were worried about the situation where people get a non-significant result and they mistakenly conclude that that means that the treatment does not work.
[Andrew] (33:02 - 33:28)
This is actually something we statisticians complain about when people falsely say, oh, non-significant, no difference between groups. So sounds like they've got kind of a noble goal here. Where these folks went wrong here in all this is their proposed solution, where they recommended that any time you have a non-significant difference, what you need to do is a post-hoc power calculation based on the effect size that you did see.
That's really where they kind of went off the rails here.
[Regina] (33:29 - 33:43)
Right. And Kristin and I talked earlier in this episode about what post-hoc power is and why it's fundamentally a flawed concept. But I'm curious, Andrew, how would you explain it to a non-technical audience?
[Andrew] (33:43 - 35:11)
Yeah. It's that it's a self-fulfilling thing. It's a circular problem that once the study's done, calculating a power number based on the observed effect size, it's just the p-value from a different angle.
It doesn't actually give you any new information. And furthermore, if the effect was non-significant by truism here, the post-hoc power will appear to be low, but it doesn't actually mean that the study really was underpowered. So I'll use this really extreme kind of cartoon example to make a point.
If I ran a clinical trial of 100,000 people, huge trial, 50,000 people get treatment A, 50,000 people get treatment B. And I see that 50.1% of the people in group A have a response to the treatment and 50.0% of the people in group B have a response to the treatment. So I say, oh, non-significant difference.
I'll do my post-hoc power calculation. And I say, oh, look, it's just low power. I didn't have enough power to detect the effect.
But obviously, a trial of 100,000 people is huge. It's plenty large to detect any reasonable effect of interest. A trial of 100,000 people with a 50% base response rate has about 88% power, even if the true effect size is 1%.
If it was a 1% increase from that 50%, that trial has 88% power. So this trial isn't underpowered to detect any real effect size we would care about. But in the little cartoon example I just described, if there's just bang on no difference between the groups, you do your post-hoc power calculation and say, oh, it's just underpowered.
[Kristin] (35:11 - 35:17)
That is a fantastic example, Andrew. Yeah, because that's obviously a well-powered study, but maybe there's just no effect.
[Andrew] (35:18 - 35:43)
Yeah. And that's the problem here. If you take this premise to its logical conclusion, the only two possible conclusions of any study is that one treatment's better than the other, or the study's underpowered.
There's no column C for actually there was no difference between the two treatments. And that leaves out probably the most common unfortunate occurrence in trials, which is that maybe in a lot of cases, there's actually not much of a difference between the two treatment groups.
[Regina] (35:44 - 35:54)
Okay. So the authors had good intentions, but the paper was logically incoherent, right? So what did you do, Andrew, when you first read it?
[Andrew] (35:54 - 36:36)
I wrote a letter to the editor. I was young and had energy and the time for this sort of scrap. And it was kind of funny, though.
I wasn't the only one who did this. There were at least seven letters to the editor about this paper, which when you think about it, these weren't coordinated letters. We were all individual people at different universities and different places.
For kind of an obscure article in a mid-level journal, this was like a pretty mid-level clinical specialty journal. It's not a statistics journal. It's not one that regularly would publish statistics papers.
Normally, papers like this basically get ignored, to be honest. So to see that much of a response was like, oh, wow, they really hit a button here.
[Kristin] (36:36 - 36:50)
Yeah, that says a lot that seven different people independently felt the need to point out the fallacy in this paper. Maybe it got some exposure on Twitter, as you said, but that is a lot of criticism. So what happened?
Did the journal respond in any way?
[Andrew] (36:51 - 37:24)
They did. And this is where I think it really starts to become an interesting conversation, not just about the controversy itself, but about science and journals and, you know, how should scientific discourse go? So the journal, to their credit, they accepted the letters.
They published the letters. I'm not really sure whether the journal editors, to be honest, understood the issue very well. I think it was just like, wow, we got a bunch of letters from people.
I guess we'll just publish the letters. In their mind, I think they're kind of letting the discourse play out. We'll publish the letters, we'll publish the author's responses, and we'll just kind of let people try to sort out what the right answer is.
[Kristin] (37:24 - 37:28)
And how did the authors respond in their response then?
[Andrew] (37:28 - 37:46)
Oh, boy. Not only did they not really back off or even really concede that anyone had made a legitimate point, they actually like doubled down. They went even harder.
So actually, what ended up happening is their response was to write that 2019 paper that I think you've already alluded to a little bit earlier on the episode.
[Kristin] (37:47 - 38:15)
Yeah, I mean, doubling down actually is very common. Authors often double down when they are criticized, and I guess it's human nature. In our p-value episode, I talked about my own criticism of a statistical method called magnitude-based inference, and I experienced something very similar.
The creators of that method absolutely doubled down, and I think actually the original creator is still doubling down eight years later, but at this point he’s kind of shouting into the wind. It hasn't really served him.
[Regina] (38:16 - 38:37)
We've talked about this a lot on this podcast before. Krista and I are always very impressed when authors own their mistakes, right? In the red dress effect episode, we followed a research team over 10 years, and we saw how their methods actually evolved and matured.
So it is possible, just not common.
[Kristin] (38:38 - 38:48)
Yeah, and I guess, Andrew, as you said, this leads us up to the 2019 Journal of Surgical Research paper that we already set up for listeners before the break, and this was the same research group.
[Andrew] (38:48 - 39:54)
Yeah, same group, I think, continuing their train of thought to what they thought must have been the logical conclusion. So they've already talked about how what you needed to do is do post-hoc power calculations, and now they did this literature review. They collected a whole bunch of clinical trials and surgical journals with non-significant results.
They did post-hoc power calculations for all of them. As we've been explaining, of course, they found that all of them had low post-hoc power, and then they started to draw some kind of troubling conclusions. This is why this actually bothered me so much.
They basically imply that what this means is that in surgery, it just must be impossible to do an adequately powered trial. They talk about how they're challenging the threshold of 0.8 for power. They use this phrase, empowering the underpowered study, and it almost veers towards promoting this idea that you don't even have to plan at all a sample size.
Just do a study with whatever people you get, and then publish the results. I don't want people to take away the conclusion that, oh, it's hard to do a proper power calculation, so you just don't have to do it, because that's really a troubling direction that this could go as well.
[Regina] (39:54 - 39:57)
Right. Anarchy is not the solution here. Yeah.
[Andrew] (39:57 - 39:57)
Yeah.
[Regina] (39:57 - 40:13)
Do not do that. Okay. So we, of course, can debate whether that target threshold for power should be 80% or not.
That is an open question, but that has nothing to do with post-hoc power, and that's where they got confused.
[Andrew] (40:14 - 40:23)
Yeah. I think that's a really important point to hit. There's an actual debate to be had about what's the right way to do a power calculation for a study.
They're not solving that problem.
[Regina] (40:23 - 40:25)
They are not solving that problem.
[Andrew] (40:26 - 41:01)
They also got so defensive in such weird ways in this paper. They drew parallels to challenging the voting age and things like that. Like they're these brave social crusaders who are challenging the concept of statistical power.
It's like, come on, guys. You don't even really understand what the problem you're trying to solve here is. It almost felt like a deflection mechanism.
Instead of addressing the issue on the merits, because we've been criticized by real statisticians who have real experience in this field, we're just going to change the narrative to this, like, we're victims of persecution and mean people on social media kind of thing.
[Regina] (41:01 - 41:11)
So, basically, they tripled down and retreated to the position of a noble martyr, trying to change science.
[Andrew] (41:11 - 41:39)
Yeah. It was a little bit like a delusions of grandeur, right? Like, come on, guys.
There's a lot of smart people that have done clinical research for a long time. I think they have an understanding what statistical power is. There's a certain amount of hubris in believing that you've stumbled on that thing that nobody else has ever figured out of like, oh, gee, we did some research and figured out that it's actually impossible to do clinical trials with statistical power calculations in surgery.
At some point, don't you have to think, maybe we're not understanding this issue properly?
[Regina] (41:39 - 41:44)
Okay. So, you wrote a letter to the editor then, and how did the authors respond?
[Andrew] (41:44 - 42:34)
Yeah. So, now we're into another exchange of letters. We said earlier that there was the first exchange of letters.
Now we're at another journal with a different editorial staff, and we're back to letters to the editor. And again, same thing happened. Actually, the journal editor called me, which was a little bit of a surprise.
And he was really nice, but he was also kind of patronizing. Again, I didn't get the impression that he really understood the merits of the issue here. He just framed it as, all right, well, we'll publish your letter, we'll do letters back and forth, and we'll let the readers kind of figure it out.
That's kind of a frustrating thing about academic sciences. I don't want to act like, oh, you just need to stifle debate. Sometimes a person is right and another person is wrong, and the answer isn't just, oh, we'll just publish both sides and let people figure it out.
[Kristin] (42:34 - 42:38)
Yeah, that's too bad that the journal editor didn't recognize that the paper was just wrong.
[Andrew] (42:39 - 43:18)
It's tricky. I do want to be sympathetic. The journal editor probably doesn't want to totally gatekeep, but that's actually kind of what a journal editor is supposed to do, right?
We don't need journal editors if the answer is just everyone can publish anything at any time. A journal is supposed to be some kind of assurance of quality, right? We could all publish our own blog posts on the internet these days.
Nobody needs journals anymore just to get thoughts out there. I'm a little troubled by the idea that the journal could just throw up their hands and act like they're not responsible at all for quality control here and be like, well, we'll just publish it all and let the readers decide. Now you do have some responsibility that what's being published in your journal actually meets some standard of quality.
[Regina] (43:19 - 43:30)
Quality control, I think, is the right word to be bringing in here because it seems like this journal made a debate just a public forum instead of realizing, no, there is right and wrong.
[Andrew] (43:31 - 43:32)
Yeah, I agree.
[Kristin] (43:32 - 43:47)
Yeah, and the authors actually got a chance to respond in this debate. So I read their response. It was very interesting.
They seem to have two main arguments. And the first one was that they insist that being redundant is not the same as being wrong.
[Regina] (43:49 - 43:54)
They really do not want to be called wrong, do they? I think that's the problem here, yeah.
[Kristin] (43:54 - 44:15)
But it's a silly argument. As we've talked about, the issue really isn't the redundancy. It's that they are making a completely circular argument, right?
They use the non-significant result to get low posthoc power and then turn around and use that low power as an excuse for the non-significant result. Their second main argument, though, is an even more interesting one because it seems to be, basically, people were mean to me.
[Andrew] (44:16 - 46:15)
Yeah. I mean, I felt like by the end, they weren't even really trying to address the statistical substance of the critique anymore, like by their last response letter. You know, there's some really funny phrases that I kind of laughed at in their last response letter.
They used the phrase in the title, a call for professional discourse to search for a solution. At this point, they have retreated. They're not even trying to defend the actual statistical substance anymore.
Now, it's just kind of this people are being mean to us on Twitter feeling. I kind of laughed because they explicitly thanked one of the people for their professionalism. And I thought that was pretty clearly an intentional kind of dig at some of the other people that responded.
And as far as I could tell, that was because some of us had talked on Twitter and others hadn't. And I think they have this kind of old school attitude of like, well, the right way to critique science is to write a letter to the editor. You don't go talk about it on social media.
You don't go tweet about it. You just write a letter to the editor and wait for the editor to publish your letter. And we do it back and forth.
Like, come on, this isn't 1850. We have more modern ways of communication. And look, I understand, like, you shouldn't be overtly mean.
Nobody should ever be overtly mean. I am not saying that you get a license because someone is wrong in a journal to be mean to them. But you also don't owe them a tone of deference.
Like, you are allowed, if you think they are wrong, to say that you think they are wrong. And you're allowed to use the word wrong. You don't have to find some alternative, nicer way to say that.
You can still say that you think they're wrong. And that's something that I actually grew very passionate about through all this was this, like, frustrated kind of backhanded way of trying to tamp down any criticism by going after people's tone instead of the merits of the arguments. This is like a little microcosm of something that I think was actually becoming pretty common around that time, which was anytime someone published a controversial and a paper in any way, if they got a lot of pushback for it, just start talking about the tone of the people criticizing you instead of the arguments.
[Kristin] (46:15 - 46:28)
Yeah, Andrew, the exact same thing happened to me with Magnitude-Based Inference. Their argument against my statistical arguments was basically I was being mean to them. And as you said, it's a common way of diverting attention by complaining about the messenger.
[Andrew] (46:28 - 46:39)
And again, it's a deflection technique. It's like you're trying to play this, like, get out of jail free card, like get out of responding to criticism card by saying that the people criticizing you are being unprofessional the way they're doing it.
[Regina] (46:39 - 46:53)
And this really brings up the whole issue of what is the proper way to debate about science in a modern age? So, Andrew, do you think there is a role for scientific debate in social media?
[Andrew] (46:54 - 47:10)
Oh, man, I have mixed feelings on this after this whole episode. I think that people, if they feel strongly about something, I think that social media is a valid place to express that. And I think it can give kind of rapid feedback.
Is it effective? After this, I'm not really sure.
[Kristin] (47:12 - 47:19)
I imagine so. And, you know, Andrew, you ended up kind of leading the charge on this criticism. How was that experience?
[Andrew] (47:19 - 48:26)
I don't always feel like I was the most eloquent messenger here. If I could go back, I'm not sure that I did everything perfectly. I would get really almost like frustrated with myself because I felt like I should somehow be able to get it across to this group of people that what they were promoting was misguided.
And I was just like, how am I not getting it across to them? If I explain it another different way, maybe they'll get it. And so that was really frustrating because, like I said, you know, yeah, social media could get a little messy.
I definitely probably let myself get a little heated with my frustration at times. And I get that to some people, there's a very traditional way of you write a letter to the editor, you be calm and professional and you just let it work itself out. But I mean, look, I ended up dedicating a weird amount of time to this.
And the weird thing when I look back on it is like, I didn't get anything out of this. I wasn't going to get promoted by writing these to the editor. Letters to the editor are really not encouraged or rewarded in academia.
So when I reflect a little bit, I almost feel like I did it for nothing. I just did it because I was mad that someone was wrong on the internet.
[Regina] (48:26 - 48:36)
Yeah. I feel like we're basically relying on what, indignation as an incentive. Maybe that's not always the best motivator.
[Andrew] (48:36 - 48:58)
That was the only real motivation. There was no personal gain from it. Actually, if anything, I was worried that I was going to get in trouble, that some of these comments about tone would come back to bite me, that my boss or my division chief would be like, Hey, Andrew, why are you complaining so much about this?
So yeah, I think that through that experience, I come away a bit jaded about it all. Right.
[Regina] (48:58 - 49:03)
So do you think that this paper ultimately should have been retracted?
[Andrew] (49:03 - 49:31)
Yeah. I actually think any of these papers, once they started to see the amount of criticism from the statisticians who were pointing out that the underlying premise was flawed, I think the journal probably should have just retracted the paper and said, you know what? We made a mistake.
One other thing, we've talked a little bit about the philosophy here, is that some people would say, well, you should only ever retract papers if they're proven to be fraudulent. I was like, no, if a paper has something wrong about it, it can be retracted. It doesn't have to be fraud to warrant a retraction.
[Regina] (49:31 - 49:50)
It's that quality you're talking about. It's not just fraud. Here, it's a common misperception, right?
But it's just not true. That's the problem. And so why would it be important in your mind to retract this then?
Why did you spend all this time on this? Why does this matter?
[Andrew] (49:50 - 50:59)
So like you said, the real worry that I had, besides just being mad about someone being wrong on the internet, is that I felt like the logical conclusion of their work that they were trying to promote is actually that you didn't need to do proper sample size planning in surgery studies. For some reason, surgery studies are so special that they don't need proper sample size planning and power calculations, which is actually really troublesome because I think of any place that you want to have really appropriate sample size planning and power calculations, it would be something like surgical studies where patients are going to undergo very invasive procedures. When you think about the ethical imperative of doing clinical trials, I think Doug Altman talked about this.
Some other real giants of statistics have talked about this. An ethical imperative of a clinical trial is that the trial actually can answer the question that it's setting out to answer. And an appropriate power calculation and sample size is part of that.
You need to know that you're recruiting the right number of patients that the trial actually can answer the question. Otherwise, you shouldn't do the trial. So why did I get so worked up about this?
I felt like they were actually leading people to the exact wrong conclusion, which is just it's too hard to do power calculations in surgery studies, so don't do them.
[Kristin] (50:59 - 51:07)
Has anything happened with this case? This was back in 2019. Did it just get left in the journal as a debate or what happened?
[Andrew] (51:08 - 51:38)
It trailed off. Everyone ran out of steam for the debate. I never followed up to see if this group wrote yet another paper.
I'm not sure what else they would have written about by that point. I think that the point was kind of played out. So honestly, I have to admit, I haven't thought about this in a couple of years.
It was like a weird chapter in my life, and I just kind of left it and moved on. I still feel very passionately about the issue. I still care, but I was no longer really part of a public debate about the issue because it seemed like that had kind of died down.
[Kristin] (51:38 - 51:43)
Yeah, it gets tiring after a while being in these debates. It's a lot of stress, and so you want to put it behind eventually.
[Andrew] (51:43 - 52:11)
And that actually, though, to circle back to Regina's question about science being self-correcting, unfortunately, like we said, that's one of the downsides of this is that sometimes for science to be self-correcting, it relies on people to have the energy to try to correct things. And the incentives aren't always very well aligned to do that. Like I said, these are the papers that are the least rewarded when it comes to promotion files, tenure files.
Nobody gets a grant because they wrote a nice paper pointing out that someone else was wrong.
[Kristin] (52:11 - 52:24)
Absolutely. Well, Andrew, someone sent me a screenshot from PubMed yesterday on magnitude based inference, and there was clearly a peak in 2018, and then it has way dropped down. So maybe in the end, the truth wins out.
[Andrew] (52:24 - 52:26)
Hopefully, yeah.
[Kristin] (52:26 - 53:03)
All right. I think we are ready to wrap up the episode. And Andrew, we'd love for you to join us in our wrap up.
And we always do two things in the wrap up, which is to rate the claim for the episode and then to do methodologic morals. And so the claim for the episode today, because it's a methodologic episode, was that post hoc power is never useful. And we rate the strength of evidence for claims on this podcast with our highly scientific 1 to 5 smooch rating scale, where 1 means little to no evidence for the claim and 5 means strong evidence for the claim.
So Regina, I'm going to let you start here. How are you rating that today?
[Regina] (53:04 - 53:21)
Well, I noticed, again, you phrased this in the negative. So post hoc power is never useful. So if I give it 5 smooches, that means I believe that post hoc power is never useful.
Kristin, I think I'm going to go with 5 smooches. Post hoc power is never useful, period.
[Kristin] (53:21 - 53:32)
Yeah. You know, Regina, I think we've kind of made our case on this episode, so I'm also going 5 smooches on this one. Do you like how I got us something we could rate as 5 smooches?
[Regina]
Yes, I know.
[Kristin]
How about you, Andrew?
[Andrew] (53:32 - 53:49)
So I'll go with 5, but I want to make that really clear qualifier that we're talking about post hoc power based on the observed effect size. And I know you guys talked about this a little bit earlier in the episode, but if you're doing a power calculation after a study, there's actually a more appropriate way to do it. But post hoc power done this way, never useful.
[Kristin] (53:49 - 54:16)
Yeah. Thanks for that clarification. Post hoc could just mean in general, but we've been using post hoc power and observed power interchangeably here.
Yeah. All right. So I think 3 out of 3 statisticians agree.
This is kind of like 3 out of 3 dentists agree that you want to floss every day. So I think we have consensus. How about methodologic morals?
This is where we give a little takeaway from the episode, kind of like an Aesop's fable moral.
[Andrew] (54:17 - 54:24)
For me, I'd say big questions need studies powerful enough to answer them. You know, you have to do a proper power calculation before a study.
[Kristin]
I love that one. How about you, Regina?
[Regina] (54:24 - 54:40)
Okay, Kristin, I'm going to do my methodologic moral based on this whole idea of circular reasoning with post hoc power.
Post hoc power is circular reasoning masquerading as new information. What about you?
[Kristin] (54:40 - 54:45)
Know when to cut your losses. A mistake gets more expensive every time you double down.
[Regina] (54:45 - 54:50)
Oh, I love that one. Thanks so much, Andrew. This was terrific.
Fascinating.
[Andrew] (54:51 - 54:54)
Yeah. Thanks so much for having me. I appreciate it.
It was a fun walk down memory lane.
[Regina] (54:55 - 54:58)
Thanks, Andrew. Thanks, Kristin. Thanks, everyone, for listening.