Topics

Research design

Deconfounding confounding part 1: traditional explanations

The first article of this series1 presented a framework to assist in judging the presence of bias: selection bias, or systematic error in how participants are identified or selected; measurement bias, or systematic error in how variables are measured; and analytical bias — also known as confounding — or systematic error in the measure of association or conclusion about causation, due to improper or incomplete analysis. Selection and measurement bias should be managed pre-emptively by good design before the start of the study, but can be detected post hoc by critical appraisal. No statistical method removes the effect of selection or measurement bias post hoc, although there are methods that allow us to model different degrees of bias and evaluate the effect on the measure of association.1 Confounding is slightly different in that it can be adjusted for in the analysis, as long as its sources are understood and measured without too much error. What is confounding? A confounder has been traditionally defined as a variable associated with both the exposure and outcome of interest without being an intermediate on the causal pathway between them, which causes a spurious or distorted estimate of the exposure–outcome association. This may be conceptually difficult to understand in the abstract, so a concrete example is useful. At the beginning of this series,1 we used the example of smoking (exposure) and dementia (outcome), and we postulated that alcohol may be a confounder. In this sense, alcohol is associated with smoking, that is, people who drink also tend to smoke, and alcohol may independently contribute to the risk of dementia. We may decide to study 100 people who smoke and 100 people who do not smoke and follow them over many years for the development of dementia (Box 1). In our study, 30% of people who smoke and 18% of those who do not smoke develop dementia. The relative risk (RR) of developing dementia is therefore 30%/18% = 1.7. However, based on our clinical knowledge, we have identified alcohol as a potential confounder and we want to adjust for this variable in estimating the effect of smoking on dementia. We may perform this adjustment either by including alcohol as a covariate in a regression model or by stratifying on this variable. In this article, we will do the latter because it makes the relationship between the variables more obvious. Stratifying the sample into drinking and non-drinking groups means that we remove the effect of drinking from our analysis. One way to understand this is to consider that within the strata of drinking status, everyone has the same level of drinking; thus, there is no variation in this variable and no potential to influence the outcome. This is a simplification, but it is useful for demonstration purposes. The relationship between smoking and dementia stratified by drinking status is shown in Box 2. The RR of dementia with smoking is one in each stratum of drinking status. How is it that combining two groups that individually show no association between smoking and dementia yields an overall group that shows an association? Summing the frequencies of the respective cells across the two 2 × 2 tables in the strata (Box 2) yields the same numbers as the overall 2 × 2 table (Box 1). To understand how confounding works in this example, we need to see two relationships: Drinking is associated with smoking: in Box 2, (25 + 25)/(25 + 25 + 10 + 10) = 50/70 ∼ 70% of people who drink also smoke, and (5 + 45)/(5 + 45 + 8 + 72) = 50/130 ∼ 40% of people who do not drink smoke. Drinking is associated with dementia: the risk of dementia in people who do not drink (Box 2) is 10%, compared with 50% in people who drink (Box 2). Drinking status is, therefore, a confounder of the relationship between smoking and drinking. Another way of understanding this is to recognise that in the overall table, what is labelled as the smoking group is actually a drinking group, and it is the drinking that is responsible for the development of dementia. We may see this by rearranging the 2 × 2 tables (Box 2) by strata of smoking, and looking at the relationship between drinking and dementia (Box 3). These tables show that the RR of dementia with drinking is 50%/10% = 5, regardless of whether smoking is present or absent, that is, RR = 5 in both strata (Box 3). In this case, the effect of smoking was completely confounded by alcohol, and adjusting for alcohol reduced the effect of smoking to zero. In reality, many associations are only partially confounded, and the effect of smoking may have been reduced after adjustment for alcohol rather than removed. In other cases, confounding may be so extreme that the effect size is reversed after adjustment for the confounding variable — this is called Simpson’s paradox. Why should a confounder not be an intermediate? The last part of the definition of a confounder is that it should not be an intermediate between the exposure and the outcome, that is, the confounder should not be caused by the exposure. In the example, we could postulate that smoking led to drinking and that drinking is what caused the increased risk of dementia. In this case, adjusting for drinking would remove precisely the effect we were trying to see. The effect of smoking on dementia may be entirely mediated through drinking, in which case, adjusting for drinking would remove the effect we are trying to measure. On the other hand, the effect of smoking on dementia may be mediated partially through drinking and partially through other mechanisms, and therefore, adjusting for drinking would remove part of the effect of smoking. We may then speak about the total effect of smoking on dementia, which includes all pathways, and then tease out the direct effect (effect of smoking directly on dementia) or multiple indirect effects (effect of smoking on dementia via drinking or other intermediates). The ultimate solution to confounding In an observational study, we cannot ensure that all potential confounders are identified and accurately measured, and hence there is always the possibility of residual confounding. Some authors have suggested that the choice of a control exposure (which should have no association with the outcome) or a control outcome (which should have no association with the exposure) may be used to shed light on the possibility of residual confounding.2 However, using a randomised controlled trial (RCT) is the only way we can ensure that confounding is handled definitively. This is exemplified by the debate over hormone replacement therapy, where results from over 30 years of data from hundreds of observational studies — collected in different settings and analysed with adjustment for different potential confounders — were overturned by one large RCT.3 Why is the RCT so powerful? By randomly assigning a sufficiently large number of people to two (or more) groups, we achieve an even distribution of all known — and more importantly, unknown — confounders across the trial arms. Any difference in outcome between the two groups is due to the only difference between them: the intervention assigned. Nevertheless, this does not mean that results from all RCTs can be believed; between the time of randomisation (when all potential confounders are evenly balanced) and the time of analysis (when outcomes are measured), there are many opportunities for that even balance of confounders to be upset and we must use our critical appraisal skills to evaluate the validity of the trial.4 Box 1 – Association of smoking and dementia Dementia Risk of outcome Present Absent Smoking 30 70 30% Non-smoking 18 82 18% Box 2 – Association of smoking and dementia stratified by non-drinking and drinking Dementia Risk of outcome Present Absent Non-drinking Smoking 5 45 10% Non-smoking 8 72 10% Drinking Smoking 25 25 50% Non-smoking 10 10 50% Box 3 – Association between drinking and dementia stratified by non-smoking and smoking Dementia Risk of outcome Present Absent Non-smoking Drinking 10 10 50% Non-drinking 8 72 10% Smoking Drinking 25 25 50% Non-drinking 5 45 10%

John R Attia · Michael P Jones · Alexis Hure

16 00491

Subscribe to MJA email alerts

No spam, you can unsubscribe anytime you want.

By providing your information, you agree to our Terms of Use and our Privacy Policy.

Thanks for Subscribing! Tell us more

Your email updates will use your name.

Good one! Your updates are coming

Thank you for subscribing to the MJA email alerts. Receive the latest content in your inbox.