Psychic or pure probability?
Authors: Lydia M McGee and Richard G E McGee
Published online: 6 December 2010
To the Editor: Paul, the 2-year-old octopus of Oberhausen Sea Life Aquarium, Germany, gained celebrity status over the course of the 2010 International Federation of Association Football (FIFA) World Cup by predicting winners with surprising accuracy. In order to make a prediction, Paul was offered food from two separate containers, each one featuring the respective team’s flag. Whichever container Paul chose to eat from first was deemed to be the predicted winner. During the World Cup, Paul correctly predicted all of Germany’s results, as well as the eventual winner, Spain. In other words, this cephalopod made correct predictions eight times in a row.

Suppose that before the beginning of the World Cup we hypothesised that Paul would correctly predict the results of eight football matches, including the grand final (H1). The null hypothesis (H0) would have been that Paul could not correctly make these predictions — that is, he was not psychic. Assuming a probability of 0.5 of correctly predicting a result (ie, a 50 : 50 chance), then the probability of predicting eight games in a row is 1 in 256 (1/28), or about 0.004. As Paul did correctly make these predictions, there is strong evidence to reject the null hypothesis. Using an exact 95% confidence interval to generate a prediction interval, the best we can say about the probability that Paul was psychic is that it is > 63% and ≤ 100%.
Therefore, should we assume Paul was psychic based on P = 0.004, or has something else happened? The most likely explanation is that a type 1 error occurred — we rejected the null hypothesis when it was true, because a P value (significance level) of 0.004 still allows a chance finding of a statistical difference to occur in 0.4% of tests. Interestingly, in clinical practice we often accept P values that indicate less significant results than this (ie, P > 0.004 but < 0.05), so type 1 errors may occur more often than we realise.
Of course, there are several problems with our analysis. First, we were already aware of the outcome when we conducted the analysis, so our probability for each successful prediction should have been 1 and not 0.5. This is an example of post-hoc probability analysis. Second, it was not an ideal scientific experiment because there was no control group, only one test was performed per match, there may have been differences in food preparation, and so on. Finally, it is not wise to conduct statistical analysis on implausible events, as this increases the probability of type 1 errors.
We do not believe that Paul had psychic powers, but his predictions do serve as a good example that type 1 errors can never be ruled out, even with highly significant results.
Interpreting Australian Stillbirth Rate Trends: Implications for Surveillance and Continuous Quality Improvement
Aleena M. Wojcieszek, Kirstine Sketcher-Baker, Christine Andrews, Michael Coory, Imogen Kettle, Melissa Malivoire, David Ellwood, Vicki Flenady
Paracetamol in Pregnancy: Uncertain Evidence, Certain Consequences
David J. Tunnicliffe, Miranda Cumpston, Debra Kennedy, Margie Danchin, Armando Teixeira-Pinto
Fatty Liver Disease in Australia: A Narrative Review on the Epidemiology, Natural History, Prognostication and Management in People With Metabolic Dysfunction
Karl Vaz, Daniel Clayton-Chubb, William W. Kemp, Stuart K. Roberts, Ammar Majeed
Birth prevalence, clinical sequelae, and management of congenital cytomegalovirus infections in Australia, 1999–2023: a national prospective study
Ece Egilmezer, Suzy M Teutsch, Carlos Nunez, Stuart T Hamilton, Adam W Bartlett, Pamela Palasanthiran, Elizabeth J Elliott, William D Rawlinson
The number of cancer‐related deaths that could be attributable to spatial disparities in survival in Australia, 2010–2019: a retrospective population‐based cohort study
Charlotte K Bainomugisa, Jessica Cameron, Paramita Dasgupta, Peter Baade
Mandatory research projects during medical specialist training in Australia and New Zealand
Paulina Stehlik, Caitlin Brandenburg, David A Henry