Topics

Statistics

Endocrinology Letters 20 November 2017 Free

Cortisone injections for tennis elbow should be an “avoid”, rather than a recommended procedure

To the Editor: We are strong supporters of Choosing Wisely, which promotes appropriate use of medical procedures and evidence-based medicine. We bring to your attention an example of a recommendation published in the 2017 edition of the Australian Therapeutic Guidelines for rheumatology,1 which is contrary to level 1 evidence (ie, multiple randomised control trials) and the Choosing Wisely ethos. The guidelines suggest that local corticosteroid injections may be considered for lateral epicondylitis (tennis elbow) and repeated if needed. The recommendation uses the less than prudent justification: “local corticosteroid injection can provide pain relief for 6–12 weeks”.1 There are now at least five high quality randomised control trials of corticosteroid injection for tennis elbow with 6 or more months follow-up, and collectively they show harm of corticosteroid compared with placebo injection or conservative treatment for time periods greater than 3 months. We reference three of these trials,2-4 and others show consistent results. There are no high quality published trials showing benefit of corticosteroid over placebo injection at time periods greater than 3 months, and one review, in fact, showed an association of poorer long term outcome with repeated injections.5 It is not reasonable, nor should it be good clinical practice, to justify a possible medium term harm by reference to a much shorter term benefit. Based on current evidence, corticosteroid injection for tennis elbow should become a Choosing Wisely “avoid” procedure. Practice guidelines such as the Australian Therapeutic Guidelines for rheumatology ought to more carefully consider level 1 evidence to avoid supporting a prevailing traditional treatment option that is not evidence-based. In treatments with potential benefits and harms that have been tested by randomised control trials, recommendations should only support those treatments with a high quality trial evidence of benefits outweighing harms.

John W Orchard · Bill Vicenzino

Tennis Elbow

Understanding statistical hypothesis tests and power

Medical researchers often attempt to understand whether a risk factor is involved in the aetiology of disease or whether an intervention reduces disease; this has been the subject of another article in this series.1 We typically do this by proposing scientific hypotheses from which testable statistical hypotheses can be developed. For example, our scientific hypothesis could be that people with irritable bowel syndrome (IBS) have higher risk of depression compared with those who do not have IBS. A small study might show that one in ten in the control group have depression, compared with two in ten in the IBS group. This might reflect a real difference in the rate of depression between IBS and non-IBS populations, or the difference of one person between groups might simply be the play of chance. An empirical approach to answering this question might be to repeat the study multiple times, with larger numbers, and to look at the consistency of the results. Unfortunately, this is an expensive and time-consuming solution. In the early 20th century, statisticians proposed different solutions. Fisher (1890–1962) suggested that researchers should start with a statistical hypothesis, commonly in the null form (no difference in incidence of depression in patients with IBS compared with those without depression); the compatibility between the observed data and this null hypothesis could be measured using a P value.2,3 If this was high, one would tend to accept the explanation that the data were consistent with the hypothesis. If it was low, one would tend to reject the hypothesis; no a priori threshold was specified for the P value and it was seen as part of a continuum, with the reader deciding where to set the threshold and how to incorporate other information.3 Neyman and Pearson sought to remove this subjectivity and focus on decision making, suggesting that one should also state an alternate hypothesis — for example, that there is an increased risk of depression in patients with IBS compared with those without depression. The discrepancy between the observed data and what would be expected under the null hypothesis would be calculated (the test statistic) and compared with an a priori threshold. If the test statistic fell below this threshold, the data were judged to be compatible with the null hypothesis. If the test statistic was above the threshold, the null hypothesis would be rejected. These opposing statisticians became embroiled in a long running and bitter feud,4 and we have been left with a legacy that reflects a combination of the two approaches — we frame a null and alternate hypothesis (Neyman and Pearson) but calculate a P value (Fisher). Even recently, the American Statistical Association felt the need to release a statement on the purpose and uses of P values.5 Many clinicians find the null hypothesis counterintuitive, and it is worth exploring why it was adopted. First, it is the simplest hypothesis to test; for example, that any difference between IBS and non-IBS populations in stress observed in our sample is due to chance. Second, it is more efficient to disprove a hypothesis than to prove it; as the philosopher of science Karl Popper argued, falsifiability is a more reliable criterion of truth than verifiability.3,4 It is important to note that rejecting the null hypothesis is not actually specific evidence for the alternate hypothesis, although it is customary to take it as such. Understanding P values Consider a clinician who observes that patients with IBS have higher stress scores on the Depression Anxiety Stress Scale6 than patients without IBS. The physician observes that individuals with IBS have a higher mean (x̄) score (mean, 15; SD, 10; n = 20) than non-IBS individuals (mean, 9; SD, 8; n = 42). The scientific hypothesis is that patients with IBS have higher stress scores than those without IBS. The corresponding statistical (null) hypothesis is that there is no difference in stress scores on average between those with and without IBS and the difference observed is simply due to chance. As noted above, the empirical approach would be to repeat this experiment hundreds of times and look for consistency. The beauty of statistics is that it saves us the time and cost of doing this. The P value estimates how likely we are to see this difference, or whether there was no real difference in mean stress scores in the population, or one more extreme, if it was simply the play of chance, without us having to actually do the repeated experiments. One way of viewing a hypothesis test is as an evaluation of the signal-to-noise ratio. In our example, the signal is the difference between IBS and non-IBS means (15 − 9 = 6 points) and the noise is the uncertainty around that difference between means (reflected through the standard error). In this instance, the appropriate test statistic is an unpaired t test, which is appropriate for the comparison of means of two independent samples of individuals (although there are many different tests appropriate to different scenarios). The t test is calculated as: As the numerator (difference between means) increases relative to the denominator (standard error of the difference), we could say that the signal-to-noise ratio increases, increasingly arguing against the null hypothesis and therefore supporting the alternate hypothesis. However, at what point does the ratio become large enough for the clinician to conclude that stress is a feature of IBS? The t value here (2.42) corresponds to a P value of 0.02. The interpretation of this P value is that if we were to repeat the experiment 100 times, and there were truly no difference in the population means, we would expect that two times out of 100 we would see a difference this large or larger. Fisher would leave it to us to decide whether this was rare enough for us to accept our hypothesis of an association, while Neyman and Pearson would say that anything less than 0.05 (five times out of 100), for example, is rare enough to reject the null hypothesis. Flaws in the argument Understanding the P value this way makes it clear that we are dealing with probabilities not certainties, so errors are possible. One can falsely reject the null hypothesis, concluding that IBS is associated with higher stress levels when it is not. The chance of this happening is another way to understand the P value and is also known as a type I error. The converse is the probability that there is no real difference between the IBS and non-IBS populations. The opposite error is also possible: accepting the null hypothesis of no association when there actually is an association — this is known as a type II error. The converse is the likelihood that there is a difference between IBS and non-IBS groups, which we rightly conclude by rejecting the null hypothesis — this is also called the power of the study. The likelihood of making these type I and II errors is partly influenced by our sample size, which is why epidemiologists and statisticians give such importance to calculating a sample size before a study begins. Let us return to our example of the possible association between IBS and stress. After the data have been collected, the hypothesis testing process involves calculating a test statistic. In this case, the unpaired t test is calculated as: Any quantity in the equation that makes t larger will lead to a smaller P value and increase the likelihood of rejecting the null hypothesis. This means that either the numerator becomes larger or the denominator becomes smaller. The denominator becomes smaller if either the standard deviation (s) is small or the sample size (n) is large. We do not have direct control over standard deviation but we can directly control sample size. Therefore, we might falsely accept the null hypothesis (type II error) either because the standard deviation is large (eg, there is a high degree of measurement error) or because the sample size is small. A red flag for a possible type II error is a clinically meaningful effect size (numerator) but a P value just above our predefined level of statistical significance. Good practice in designing medical studies is to calculate what sample size is required to achieve adequate statistical power (typically ≥ 0.8) at a given level of statistical significance (typically < 0.05), but only if the effect size is large enough to be clinically meaningful. Similarly, it is good practice to report either the a priori sample size calculation or the statistical power available for the sample size achieved. Conclusion Statistical hypothesis tests require clinical and biological science in their construction and clinical interpretation in applying their results (Box). While they can be viewed as a “black box”, with little or no understanding of their statistical basis, understanding their construction and origins yields important insights into the design of medical studies, including why a prior sample size calculation is important. Box – The role of statistical hypothesis tests in empirical studies

Michael P Jones · Alissa Beath · Christopher Oldmeadow · John R Attia

16 01022

Deconfounding confounding part 2: using directed acyclic graphs (DAGs)

In the previous article on confounding in this series,1 we presented the traditional explanation of a confounder. Over the past few decades, it has become clear that this definition has many limitations. For example, confounding can be induced by a network of variables rather than just a single variable, and adjusting for potential confounders can paradoxically increase confounding. What are directed acyclic graphs? One of the few true innovations in epidemiological methods has been the emergence of directed acyclic graphs (DAGs) to identify confounding. This development began in the early 1990s with work by Pearl and Robins based on formal logic and machine learning.2-4 DAGs are a formal system of mapping variables and the direction of causal relationships among them. “Directed” refers to arrows indicating the direction of causality between variables, and “acyclic” means that it should not be possible to start from any one variable and follow a series of arrows back to the original variable. Entire books are devoted to this method,2-4 but just a few highlights are sufficient to help clinicians understand confounding.5 In the example of smoking (exposure) and dementia (outcome) that we used in the previous article, we postulated that alcohol might be a confounder.1 These relationships are illustrated using a simple DAG in Box 1. Drinking alcohol may increase the risk of smoking — hence the arrow pointing from alcohol to smoking — and may also increase the risk of developing dementia — hence the arrow pointing from alcohol to dementia — but alcohol is not an intermediate between smoking and dementia. Smoking leads to cardiovascular disease (CVD), which also increases the risk of dementia; in this case, CVD is an intermediate between smoking and dementia. Alcohol is likewise causally related to CVD. If we observe a relationship between smoking and dementia, it is therefore not clear whether this is valid or whether it is spurious because of the effect of the other variables. How to read a DAG The steps to take in interpreting a DAG are as follows: Remove all arrows emanating from the exposure of interest. Look for any remaining path, called a “backdoor path”, that links the outcome to the exposure. A path is defined by successive arrows regardless of the direction of the arrowheads. A backdoor path is “closed” if any variable on that path is a “collider”, which is a variable with two arrowheads pointing into it from other variables on the same path. If a backdoor path can be traced and is “open” (ie, it does not include a collider), then adjusting for any variable on that path will close the backdoor path and remove confounding. Adjusting for a collider will reopen the backdoor path and increase confounding. In our example (Box 1), when we remove the arrows pointing from the exposure (smoking), we can no longer trace a path from dementia through CVD to smoking. This means that CVD is not a confounder but a mediator. Adjusting for CVD could therefore remove part of the effect of smoking that we are trying to detect; this is called over-adjustment bias.6 However, we can still trace a path from dementia through alcohol to smoking. This path does not include a collider and is thus open. By adjusting for alcohol, we can close this backdoor path and remove confounding due to alcohol. Note that another backdoor path exists from dementia through CVD to alcohol and then smoking. Adjusting for alcohol has therefore closed two backdoor paths at once. Using DAGs may seem like a convoluted process, with little benefit, compared with the old definition of confounding we presented in the previous article.1 This is certainly true when models are as simple as this one. But what happens when we want to tackle more complicated models? Let us assume we add the variables of socio-economic status (SES), diet, sex and age to the model, as shown in Box 2. How does the old definition of confounding apply here? Should we just adjust for everything? The DAG helps us sort out these relationships and decide on covariates to include in regression models. Open backdoor paths in this model include: Dementia – CVD – diet – SES – smoking Dementia – alcohol – SES – smoking Dementia – CVD – age – sex – smoking We should therefore adjust for a variable on these paths to remove confounding (eg, adjusting for SES closes the first two backdoor paths and adjusting for age closes the third backdoor path). Note that CVD remains a mediator. Another backdoor path that can be traced is: Dementia – CVD – age – sex – alcohol – SES – smoking This is already a closed path because alcohol is a collider on this path (ie, two arrows point into alcohol along this path, from sex and SES). Note that alcohol is a collider on this path but not necessarily on other paths. As this path is already closed, we do not need to adjust for any variables on it; indeed, if we were to adjust for alcohol (the collider), we would reopen a route for confounding! It takes reading and practice to become familiar with these rules, but they are extremely powerful in teasing out complex causal pathways. The astute reader will realise that trying to reduce confounding by adjusting for one variable along an open backdoor path could increase confounding if that variable is also a collider on another backdoor path. For example, as we saw above, adjusting for alcohol could close the backdoor path of dementia – alcohol – smoking, but it could also reopen the closed path of dementia – CVD – diet – SES – alcohol – sex – smoking, because alcohol is a collider on this pathway (Box 2). A tool to help readers learn to draw and interpret DAGs is a free software program called DAGitty (http://www.dagitty.net/dags.html).7 This program is reasonably intuitive, flexible and fast to learn. It allows the user to easily draw DAGs and, as a bonus, “reads” the DAG to provide the minimum set of variables for which it is essential to adjust to remove confounding. Other similar programs are also available, such as TETRAD (http://www.phil.cmu.edu/projects/tetrad), DAG (https://epi.dife.de/dag), and dagR, a set of functions for the statistical software R. The power of DAGs DAGs are powerful in that they lead us to several observations: Identifying confounders depends on the underlying causal model that is assumed. Confounding can be due to a network of variables, not just a single variable. Assumptions (model) must be drawn before conclusions (analysis) are drawn. The relationship between variables can be specified in many different ways (ie, the direction of the arrows can influence decisions made for analysis). All backdoor paths must ultimately have one variable with an arrowhead leading into the exposure; this means that confounding in complex webs of causation can be analysed by looking for the few variables that have arrows pointing to the exposure. Many routes of confounding can potentially be closed by adjusting for just one or two variables. Adjusting for confounding It could be argued that one should just adjust for all potential variables — the so-called kitchen sink approach — rather than taking any chances on specifying a potentially incorrect model. Although often appealing because of its simplicity, there are two main reasons why this approach is not recommended. The first is inefficiency: every variable that is added to the model uses degrees of freedom, which are the currency of power. This is particularly a problem for small studies, where problems such as over-fitting due to data sparsity and collinearity may arise.8 Adjusting for variables that are not confounders wastes power and may reduce the ability to detect an association. (There is a separate argument for including variables on the basis that they help explain the outcome and hence increase power; this can be thought of as “mopping up” some of the variance in the outcome so that there is more power to detect the contribution of the exposure, but this consideration is separate from confounding.) The second reason is bias: as we have seen, adjusting for a variable that is a collider reopens a backdoor path and increases the potential for confounding. Ideally, every study should make explicit the causal model behind the analysis. Residual confounding Unfortunately, even accurate specification of the causal model and expert use of DAGs do not completely remove confounding. This situation is called residual confounding, which occurs for three reasons. First, longitudinal data may not be available. With cross-sectional data, we can never be sure about the direction of causality (ie, the “chicken and egg” problem). Second, as we are not able to measure most variables perfectly, even adjusting for a variable cannot fully remove its effect. If we think of confounding as a flow of water that travels along the backdoor path, our inability to accurately measure SES, for example, means that we are unable to fully turn off the tap at that point. This is an argument for possibly adjusting for multiple variables along a backdoor path, so that the flow of confounding is reduced at multiple points instead of one point only. Third, in any observational study, we can never be sure that we have included and measured all the relevant potential confounders. What other variables have we not thought of, and not included on the diagram, that could create a backdoor path between our outcome and our exposure? As we saw in the previous article,1 the ultimate solution to confounding is a randomised controlled trial. When this is possible, it means that all known and unknown confounders are evenly balanced across the arms of the trial, thus removing their ability to affect the outcome. Box 1 – Simple causal diagram of smoking and dementia, with effect of alcohol (potential confounder) and cardiovascular disease (mediator)* People who drink alcohol are also more likely to smoke; alcohol may affect risk of dementia and CVD; and CVD may influence risk of dementia (eg, through subclinical infarcts). Green arrows radiating from the exposure are ignored when reading the pathways for potential confounding. CVD = cardiovascular disease. *Figure originally drawn with DAGitty (http://www.dagitty.net/dags.html). Box 2 – Causal diagram of smoking and dementia, with effect of alcohol (potential confounder) and cardiovascular disease (mediator), plus sex, age, socio-economic status and diet* Sex influences alcohol consumption, smoking and age (men are more likely than women to drink alcohol and smoke, and women live longer than men); SES influences alcohol consumption, smoking and diet; diet influences CVD; and age influences risk of CVD and dementia. As in the simple model in , we have to adjust for alcohol to close the dementia – alcohol – smoking path. However, doing this reopens the dementia – CVD – diet – SES – alcohol – sex – smoking path, because alcohol is a collider on this path. So we have to also adjust for sex or SES, or both, to reclose this path. Adjusting for sex, alcohol and SES would therefore be sufficient to remove confounding in this analysis of smoking and dementia. CVD = cardiovascular disease. SES = socio-economic status. *Figure originally drawn with DAGitty (http://www.dagitty.net/dags.html).

John R Attia · Christopher Oldmeadow · Elizabeth G Holliday · Michael P Jones

16 01167
Statistics Research 17 February 2014 Free

Trends in chlamydia positivity among heterosexual patients from the Victorian Primary Care Network for Sentinel Surveillance, 2007–2011

Victorian data show a concerning increase in chlamydia positivity over time, particularly in young women

Megan S C Lim PhD, BBiomedSci(Hons) · Carol El-Hayek BSc, MEpi · Jane L Goller RN, MPH, MHlthSci · Christopher K Fairley PhD · Phuong L T Nguyen BBiomedSci(Hons) · Rochelle A Hamilton MHlthSci · Dorothy J Henning GD-ADOLHW, MNurs · Kathleen M McNamee MB BS, FRACGP, MEpi · Margaret E Hellard FRACP, FAFPHM, PhD · Mark A Stoove BAAppSci(Hons), PhD

13 10108
Women's health Letters 3 February 2014 Free

Unexplained variation in hospital caesarean section rates

To the Editor: Lee and colleagues assessed the hospital caesarean section (CS) rates among hospitals in New South Wales according to the Robson 10-group classification. They identified the 20th centile CS rate for each group.1 These data could be used in Australian maternity services to reduce CS rates and potentially increase the quality of care. Monitoring the quality of care in maternity services has been difficult ...

Rick J Fielke

13 11200

Subscribe to MJA email alerts

No spam, you can unsubscribe anytime you want.

By providing your information, you agree to our Terms of Use and our Privacy Policy.

Thanks for Subscribing! Tell us more

Your email updates will use your name.

Good one! Your updates are coming

Thank you for subscribing to the MJA email alerts. Receive the latest content in your inbox.