Volume 211 - Issue 7

Difficulties in knowing which critical care trial data warrant change in practice

Authors:  Benjamin Reddi, Mark Finnis and Sandra Peake

Med J Aust 2019; 211 (7): 306-307.e1. || doi: 10.5694/mja2.50331
Published online: 7 October 2019
Why is some strong evidence ignored while some weak evidence is rapidly acted upon?

Why is some strong evidence ignored while some weak evidence is rapidly acted upon?

Most clinicians aspire to practise evidence‐based medicine, no longer believing it acceptable to implement novel interventions simply because they “make sense” or remain untested. However, external influences, psychological factors, and misapplied statistical techniques may hinder rational decision making. Using examples from intensive care literature, we discuss why well supported therapies are not always readily adopted, while poorly supported interventions may be unduly welcomed into practice.

Nearly 20 years ago, data from a multicentre randomised controlled trial (RCT) indicated that ventilating patients with acute respiratory distress syndrome with lower tidal volumes was associated with improved survival.1 These findings were replicated in other health care systems around the world, with biological plausibility established from pre‐clinical and ex vivo investigation of trauma to the lung caused by mechanical ventilation. Yet despite these robust data, low tidal volume ventilation (LTVV) is inconsistently applied in intensive care units (ICUs) around the world.2

Similarly, several RCTs have identified a mortality benefit from selective decontamination of the digestive tract (SDD).3 SDD is a strategy of administering systemic antibiotics and non‐absorbable enteral and topical antibiotics to ICU patients in an attempt to eradicate carriage of potentially pathogenic organisms and minimise nosocomial infection. These data were generated in European health care systems, arguably comparable to Australasian systems, although with slightly lower prevalence of hospital‐acquired infections and multiresistant organisms. Nevertheless, SDD has never been widely implemented in Australasia.

Despite strong supportive evidence, why have clinicians failed to implement LTVV and SDD into their practice? Data indicate that one reason that LTVV is not consistently applied to patients with acute respiratory distress syndrome is the failure to recognise the condition2 — which requires radiological interpretation and, by some definitions, an estimate of left atrial pressure — and perhaps more importantly, a lack of awareness of the diagnosis. Clinicians may also find it implausible that simple calculations based on body weight are the best and safest way to select tidal volume, and may perhaps consider that basing ventilation around individual patient pulmonary mechanics might be superior. Another reason that clinicians may be reluctant to institute LTVV is that low tidal volumes and high respiratory rates are distinctly non‐physiological and often demand sedation (known to be harmful to ICU patients in general) and muscle paralysis. Moreover, LTVV is initially associated with worsening oxygenation and respiratory acidosis;1 these potential harms are seen at the bedside, while the longer term benefits are not immediately apparent.

Suspicion of unmeasured harm probably underlies reluctance to embrace SDD, as routine administration of broad‐spectrum antimicrobials is anathema to clinicians fearful of inducing antibiotic resistance. This circumspection seems prudent, with subsequent data indicating that short term mortality benefits may be negated by increased gram‐negative bacterial resistance rates in the respiratory tract and gut.4 To address these concerns, the Australian and New Zealand Intensive Care Society (ANZICS) Clinical Trials Group (CTG) is currently supporting a multicentre RCT evaluating SDD in Australasia.

Conversely, interventions with less compelling evidence have been adopted with undue enthusiasm. In 2004, the first Surviving Sepsis Campaign (SSC) guidelines were published, giving a clear and comprehensive list of recommendations for managing sepsis and septic shock.5 These industry‐sponsored guidelines were widely adopted by clinicians and incorporated into practice guidelines internationally, being endorsed by many professional societies. However, other societies, including the Infectious Diseases Society of America and ANZICS, pointed out that only five of 52 recommendations were based on level A evidence (supported by two or more high quality RCTs), with most based on low, very low or unclassified levels of evidence. For example, justification for the resuscitation strategy rested heavily on a single trial of early goal‐directed therapy (EGDT) — a bundle of interventions instituted to achieve certain physiological targets — undertaken in a single emergency department in the United States. The external validity of this study, which reported remarkably high control group mortality,6 was significantly overestimated. Furthermore, some bundle components lacked physiological plausibility7 and promoted arbitrary thresholds over clinical evaluation. Fixed resuscitation regimens and physiological targets were favoured over patient‐centred assessment and titrated therapies. Given these criticisms, why did this bundle of care achieve such widespread acceptance?

Intensive insulin therapy (IIT) — maintaining blood sugar levels below 6.1 mmol/L — was similarly embraced on weak data. A 2001 Belgian study identified that IIT in the ICU was associated with a 34% reduction in hospital mortality.8 Again, despite data being gathered from just one surgical ICU in which parenteral hyperalimentation was common (casting doubt on external validity) and with, perhaps, an implausibly high survival benefit, IIT was broadly adopted.9

ANZICS CTG was responsible for undertaking large, multicentre RCTs evaluating both EGDT for septic shock and IIT in the ICU. The Australasian Resuscitation In Sepsis Evaluation (ARISE) study joined a chorus of trials that identified that EGDT was not associated with improved mortality, but increased hospital costs.10 The Normoglycaemia in Intensive Care Evaluation–Survival Using Glucose Algorithm Regulation (NICE‐SUGAR) study11 failed to reproduce the benefits of IIT in a general ICU population, rather IIT was associated with increased mortality. It has been estimated that implementing IIT may have been responsible for 26 000 additional deaths annually in the US alone.9

Why were EGDT and IIT embraced by clinicians on such flimsy evidence? The diagnosis of sepsis can be as challenging as that of acute respiratory distress syndrome. However, the SSC promoted sepsis awareness remarkably successfully and, in fact, evidence would suggest that sepsis and septic shock are frequently overdiagnosed.12 The SSC received generous support from stakeholder industries, 11 of the 24 SSC guideline authors declared a conflict of interest with corporate entities that potentially stood to gain from adoption of the SSC guidelines. Furthermore, the publicity around SSC predisposed the incorporation of EGDT into government‐imposed quality indicators,13 adding financial incentives to further cement them into clinical practice. Likewise, IIT was a measured metric easily incorporated into quality assessment and reimbursement models.9

Clinicians are also disposed to strategies that appear intuitive and easily applied. The SSC included recommendations that seemed self‐evident, such as early antibiotic administration, and an easily followed framework that offered simple solutions to complex pathophysiology. This clear therapeutic recipe, promoted by the august SSC, appealed to busy clinicians making rapid decisions in acute medical environments.

A further, pivotal explanation for unwarranted adoption of subsequently refuted trials may lie in our interpretation of study results themselves. Most clinical studies report whether evidence refutes the null hypothesis by generating a P value,14 the probability that an observed difference would be equal to or greater than its observed value, providing study conduct and assumptions used to compute the P value (including the null hypothesis) are correct. Despite its prominence, the P value has shortcomings. First, it measures the probability of observed data given the hypothesis, not the probability a hypothesis is true given the data. The latter depends on the P value, but if the prior probability of the hypothesis and probability of the data under other hypotheses are not also considered, an unsophisticated observer could considerably overestimate effectiveness based on a low P value and erroneously adopt ineffective therapy. The fallacy persists that a significance threshold of P = 0.05 risks only a one in 20 chance that an ineffective intervention will be wrongly identified as effective; rather, it can be calculated that the false positive rate is usually far higher.14 In the case of EGDT and IIT, considering prior knowledge of expected mortality rates from septic shock, the validity of fixed physiological targets as resuscitation targets or associations between hypoglycaemia and mortality might have provided a valuable context to the interpretation of statistical significance tests in isolation. Conversely, misconceptions such as P > 0.05 proving the null may discourage further investigation of potentially effective therapies.15 While it is patently wrong to suggest that P = 0.049 proves an intervention is beneficial while P = 0.051 proves futility, the fault lies not with the statistical significance test itself. After all, a broadly accepted threshold value in a continuous parameter is one useful piece of information among others to consider when making what in clinical practice are frequently binary decisions.

How then to proceed? When good, reproducible evidence is not implemented, reasons must be identified and resolved. For example, in the case of LTVV, the proliferation of integrated electronic records and decision support tools may trigger consideration of the diagnosis and appropriate action; meanwhile, clinicians convinced that better or safer methods of tidal volume selection exist should subject their views to trial evaluation.

When small single centre trials generate positive results, we must remember that statistical analysis involves inherent uncertainty, there is no alchemy by which significance tests guarantee proof. Fisher never suggested that a single dataset generating a P value near 0.05 would reduce the false positive rate to an acceptable level, and emphasised the importance of reproducing the result.16 Reproducing studies should be encouraged and resourced and has proven particularly crucial when the external validity of single‐centre studies is questionable (such as EGDT or IIT), biological plausibility is in doubt or there is concern regarding safety (EGDT, IIT, SDD). A P value must be interpreted alongside evaluation of existing data, study methodology, informed appraisal of underlying plausibility, potential risks, and reproducibility. An RCT of a highly plausible, safe intervention supported by robust observational data, which has a positive but non‐significant result may be a more promising avenue than an implausible, potentially dangerous intervention that generates a statistically significant benefit in a poorly conducted single‐centre study. Lastly, clinicians should be vigilant that reimbursement models and industry concerns do not unduly flatter modest data.


Authors


Competing interests


References


Provenance: Commissioned; externally peer reviewed.