Deciding when quality and safety improvement interventions warrant widespread adoption
Authors: Ian A Scott and John B Wakefield
Published online: 20 May 2013
Evaluative criteria are needed to determine the likelihood of successful implementation and acceptable return on investment
Determining when a specific quality and safety improvement intervention (QSII) has sufficient evidence of effectiveness to warrant widespread implementation is highly controversial.1,2 Some large-scale QSIIs have been shown to be less effective than originally thought (Box).3-8 Reporting guidelines for QSII studies stipulate sufficient detail to allow users to gauge the feasibility and reproducibility of a specific QSII within local contexts.9 Some authors have focused on study designs and statistical methods used to evaluate QSIIs.10 An international expert group has distilled several key themes that researchers should consider and discuss when describing experiences with specific QSIIs.11
While considerable resources are being directed at QSIIs in Australia and elsewhere, recent literature reviews show significant shortcomings in research on QSII effects.12-14 Based on these reviews and our experience with various QSIIs, we propose a checklist of evaluative criteria that decisionmakers — clinicians, quality teams, policymakers and statutory bodies — can apply to existing literature relating to specific QSIIs to determine whether they are fit for purpose and whether widespread adoption is justified.
What is the problem; where, when and how often does it occur; who does it affect and by how much; what are the predisposing or mitigating factors; and what are the potential levers for remediation? Qualitative and quantitative data are necessary to elucidate the root cause of the problem, which should inform the design of a responsive QSII.
What individual or organisational behaviour is the QSII trying to change and how will it do this? Many QSIIs are complex, multifaceted, socially embedded, non-linear interventions which vary in their context (target population and setting), content and application (the QSII itself and how it will be delivered), and outcomes. The QSII should have a sound theoretical construct which explains and predicts how it will effect change in care and is fully cognisant of the beliefs and attitudes of target groups.15 Validated theories of behavioural and organisational change need to have been considered in developing a model of change which addresses the key issues listed in Appendix 1.16,17 A review of guideline implementation studies found that only 23% mentioned a theoretical framework, most referring to only one theory.18
Pilot testing of a QSII should have demonstrated its feasibility and potential benefit, while exposing any “weak links,” learning curves, unanticipated contextual barriers and undesirable consequences. Systematic literature reviews should have been conducted to identify prior experience with similar QSIIs, including field studies or modelling exercises that assess feasibility regarding up-front implementation costs.19
Demonstrating successful results from implementation of a single-site QSII does not guarantee generalisability of effect. If a QSII is to be replicated and tested in multiple settings, it must be standardised to some degree. However, strict standardisation may impede local adaptation required for successful implementation. During implementation of the World Health Organization’s Surgical Safety Checklist, which was associated with significant reductions in mortality and complications across eight sites in different countries, local refinement of each step according to perceived need was allowed.20 However, we propose that a QSII implemented in more than one setting should have — as a minimum level of standardisation — common objectives, theoretical framework, target populations and core components.
Measuring the success of a QSII is prone to bias if it relies on qualitative self-reports of individuals directly involved in its design and implementation,21 so externally verifiable outcome data are preferred. Also, outcome measures should be standardised and appropriate, data should be collected accurately and comprehensively, and study designs should minimise the risk of confounding.
Were outcome measures standardised and appropriate? Well defined and objective patient-important outcomes minimise ascertainment bias3,22 (Appendix 2). QSII studies which report surrogate or intermediate outcomes (eg, change in medication error rates, or compliance with surgical site identification policies) should indicate how strongly such measures correlate with hard clinical end points. For example, improvements in “safety culture” as measured by survey tools show tenuous associations with reductions in patient harm.23
Were data collected accurately and comprehensively? The extent of inaccurate or missing data in many QSII studies is significant, and cherry-picked data from sites that perform better than others are often presented as the generalisable result. In addition, the intervention itself can alter how data are collected. For example, greater pharmacist participation in clinical teams may not only prevent prescribing errors but also unearth previously undetected errors.24
Did study designs minimise the risk of confounding? Investigators should have used study designs which minimise bias (Appendix 2). Randomised studies minimise selection bias in attributing improved patient outcomes to QSII effects. Cluster randomised trials involving multiple sites avoid contamination of control groups within sites. If extensive rollout of a QSII is already occurring or about to occur, then stepped wedge designs which insert randomisation into the phasing of implementation are preferred. Where randomisation is impractical, non-randomised studies (controlled before–after trials, interrupted time series studies, statistical process control charts) can be used.
Were data that adequately tested the theory collected? Evaluations should have assessed whether what the theory predicted occurred in terms of behaviour change, and whether contingencies were accurately foreseen and responded to. Variables which may, theoretically, have an impact on QSII effectiveness (eg, participant characteristics, intervention intensity, or effect modifiers) should be measured quantitatively and qualitatively.
Were detailed process evaluations reported? Theory-driven process evaluations of QSIIs which describe actual implementation (intervention as performed) versus original intention (intervention as planned) enables users to differentiate between lack of effect due to potentially avoidable implementation failure and across-the-board ineffectiveness. This helps identify instances where no amount of intervention re-engineering is likely to render it sufficiently effective to be worth pursuing. However, process evaluations that suggest good execution of a QSII do not guarantee effectiveness. For example, in one study, educational visits for general practitioners aimed at influencing prescribing practice were well received and associated with high recall, but prescribing behaviour changed little and was constrained by patient preference and local hospital policy.25
The potential for some QSIIs to harm patients should have been considered. For example, it has been suggested that decreasing junior medical staff working hours to reduce fatigue-related errors might increase errors due to greater discontinuity of care and multiple handovers. However, this concern has been allayed.26
While formal economic analyses of QSIIs are rare, some attempt should have been made to quantify resource use and costs involved in implementation (personnel, equipment, training programs, consumables, etc) to compare investment required with achievable benefits. While cost savings may accrue by minimising expensive safety errors in patient care, QSIIs may incur considerable opportunity costs, as has been claimed for the 100,000 Lives Campaign.27
Studies of QSIIs that report large benefits over short periods are more likely to be true if:
prevalence of suboptimal or unsafe care, in the absence of the QSII, was quite high
the effects are plausibly explained by the theory underpinning the QSII and supported by process evaluations
plausible confounders that would have reduced the observed effect have been accounted for
similarly large effects have been observed across multiple studies
levels of uncertainty of effect estimates, as expressed by confidence intervals, are relatively small.
QSIIs should be favoured if there is evidence of sustainability of effects across multiple sites over 2 years or more.
Methodological limitations and possible sources of confounding, particularly for observational studies, should have been openly acknowledged, together with any conflicts of interest involving researchers who may benefit financially from providing QSII consultancies.
Before applying the checklist to a specific QSII, users must retrieve as much published evidence relating to the QSII as possible. By applying the checklist to this evidence, it is possible to build a profile of the QSII according to the evaluative criteria. Responses to the criteria may be dichotomous (yes or no) or, if the evidence is more subjective and uncertain, graded (using a 5-point Likert scale). Users may also wish to give different weightings to individual criteria depending on how critical they regard them to the overall utility of the QSII. We do not imply that all 12 criteria must attract favourable responses before proceeding with QSII implementation, although we feel that most QSIIs should satisfy Criteria 1–8. If they do not, we recommend that detailed longitudinal evaluations are undertaken as the QSII is implemented in pilot sites.
If a major quality problem is evident and requires urgent remediation, QSIIs that appear promising but have not been extensively evaluated may need to be considered. In such cases, we advise rigorous evaluation during implementation.
The strength of the checklist is that it encourages a structured appraisal of how QSIIs have been developed, implemented and evaluated. As an example, the checklist is applied to evidence around hospital rapid-response teams in Appendix 3. The results indicate that if this checklist had been available some years ago, it may have tempered early enthusiasm for rapid-response teams. The checklist also highlights the need for frequent and systematic evaluations of newly developed QSIIs. In an era of limited resources, the potential effectiveness and likely return on investment of specific QSIIs must be assessed. The checklist may contribute to greater discipline and transparency of investment decisions and help clarify which QSIIs require further refinement and testing before large-scale implementation.
Comparison of initial and later experience of two large-scale quality and safety improvement interventions
RRTs are multidisciplinary teams of medical, nursing and airway management staff charged with prompt bedside evaluation, triage and treatment of clinically deteriorating patients throughout all hospital wards outside intensive care units (ICUs). Their aim is to reduce preventable deaths, cardiac arrest, unplanned ICU admissions and postsurgical complications.
Initial experience: Early trials suggested a large potential benefit of RRTs in reducing unexpected cardiac arrests (by up to 50%), unplanned ICU admissions (by up to 44%), postoperative deaths (by up to 37%) and mean length of hospital stay (by up to 4 days).3,4 As a result of such observations and advocacy for RRTs from the Institute for Healthcare Improvement’s 100,000 Lives Campaign, hundreds of hospitals worldwide have implemented RRTs.
Later experience: The validity of earlier positive observations has been challenged and a meta-analysis of 18 high-quality trials confirmed no reduction in mortality, although cardiac arrest calls were reduced by a third.5
Pay-for-performance (P4P) schemes
P4P schemes involve defined changes in reimbursement to clinical providers (individual clinicians, group practices or hospitals) in direct response to a change in one or more performance measures as a result of one or more practice innovations. Their aim is to incentivise optimal provider performance and improve quality and safety of care.
Initial experience: In the United Kingdom, large-scale implementation of P4P contracts for family practitioners over 12 months was reported in 2006 to have resulted in practitioners achieving a median of 97% of their available points covering quality of clinical care, well in excess of the predicted 75%.6 However, no baseline was established for most indicators. The United States Institute of Medicine and high-profile quality experts recommended greater use of P4P programs to improve quality of care, and by 2009 more than 200 P4P programs covering over 50 million beneficiaries were implemented.
Later experience: A review of 17 studies (12 controlled trials) showed modest improvement (4%–8% absolute increases) in some or all process-of-care measures in five of six studies of clinician-level financial incentives and seven of nine studies of group practice-level incentives.7 Four studies showed unintended adverse effects (gaming, patient exclusion, and tick-box documentation of undelivered care). A 2009 review of P4P schemes in the UK showed that, within 2 years of commencement, there was no further improvement in quality-of-care indicators despite a more than £1 billion budget overrun and a decline in continuity of care.
Competing interests
References
- Auerbach AD, Landefeld CS, Shojania KG. The tension between needing to improve care and knowing how to do it. N Engl J Med 2007; 357: 608-613. 0_i1115712
- Berwick DM. The science of improvement. JAMA 2008; 299: 1182-1184. 0_i1115714
- Buist MD, Moore GE, Bernard SA, et al. Effects of a medical emergency team on reduction of incidence of and mortality from unexpected cardiac arrests in hospital: preliminary study. BMJ 2002; 324: 387-390. 0_i1115716
- Bellomo R, Goldsmith D, Uchino S, et al. Prospective controlled trial of effect of medical emergency team on postoperative morbidity and mortality rates. Crit Care Med 2004; 32: 916-921. 0_i1115718
- Chan PS, Jain R, Nallmothu BK, et al. Rapid response teams: a systematic review and meta-analysis. Arch Intern Med 2010; 170: 18-26. 0_i1115720
- Doran T, Fullwood C, Gravelle H, et al. Pay-for-performance programs in family practices in the United Kingdom. N Engl J Med 2006; 355: 375-384. 0_i1115722
- Petersen LA, Woodward LD, Urech T, et al. Does pay-for-performance improve the quality of health care? Ann Intern Med 2006; 145: 265-272. 0_i1115724
- Campbell SM, Reeves D, Kontopantelis E, et al. Effects of pay for performance on the quality of primary care in England. N Engl J Med 2009; 361: 368-378. 0_i1115726
- Davidoff F, Batalden P, Stevens D, et al; SQUIRE Development Group. Publication guidelines for improvement studies in health care: evolution of the SQUIRE Project. Ann Intern Med 2008; 149: 670-676. 0_i1115728
- Fan E, Laupacis A, Pronovost PJ, et al. How to use an article about quality improvement. JAMA 2010; 304: 2279-2287. 0_i1115730
- Shekelle PG, Pronovost PJ, Wachter RM, et al. Advancing the science of patient safety. Ann Intern Med 2011; 154: 693-696. 0_i1115732
- Alexander JA, Hearld LR. What can we learn from quality improvement research? A critical review of research methods. Med Care Res Rev 2009; 66: 235-271. 0_i1115734
- Brown C, Lilford R. Evaluating service delivery interventions to enhance patient safety. BMJ 2008; 337: a2764. 0_pgfId-1115736
- Kaplan HC, Brady PW, Dritz MC, et al. The influence of context on quality improvement success in health care: a systematic review of the literature. Milbank Q 2010; 88: 500-559. 0_i1115737
- Wakefield J, McLaws ML, Whitby M, Patton L. Patient safety culture: factors that influence clinician involvement in patient safety behaviours. Qual Saf Health Care 2010; 19: 585-591. 0_i1115739
- Francis JJ, Stockton C, Eccles MP, et al. Evidence-based selection of theories for designing behaviour change interventions: using methods based on theoretical construct domains to understand clinicians’ blood transfusion behaviour. Br J Health Psychol 2009; 14 (Pt 4): 625-646. 0_i1115741
- Dixon-Woods M, McNicol S, Martin G. Ten challenges in improving quality in healthcare: lessons from the Health Foundation’s programme evaluations and relevant literature. BMJ Qual Saf 2012; 21: 876-884. 0_i1115743
- Davies P, Walker AE, Grimshaw JM. A systematic review of the use of theory in the design of guideline dissemination and implementation strategies and interpretation of the results of rigorous evaluations. Implement Sci 2010; 5: 14-19. 0_i1115745
- Eldridge S, Spencer A, Cryer C, et al. Why modelling a complex intervention is an important precursor to trial design: lessons from studying an intervention to reduce falls-related injuries in older people. J Health Serv Res Policy 2005; 10: 133-142. 0_i1115747
- Haynes AB, Weiser TG, Berry WR, et al; Safe Surgery Saves Lives Study Group. A surgical safety checklist to reduce morbidity and mortality in a global population. N Engl J Med 2009; 360: 491-499. 0_i1115749
- Morganti KG, Lovejoy S, Haviland AM, et al. Measuring success for health care quality improvement interventions. Med Care 2012; 50: 1086-1092. 0_i1115751
- Cornish PL, Knowles SR, Marchesano R, et al. Unintended medication discrepancies at the time of hospital admission. Arch Intern Med 2005; 165: 424-429. 0_i1115753
- Nieva VF, Sorra J. Safety culture assessment: a tool for improving patient safety in healthcare organizations. Qual Saf Health Care 2003; 12 Suppl 2: ii17-23. 0_i1115755
- Leape LL, Cullen DJ, Clapp MD, et al. Pharmacist participation on physician rounds and adverse drug events in the intensive care unit. JAMA 1999; 282: 267-270. 0_i1115757
- Nazareth I, Freemantle N, Duggan C, et al. Evaluation of a complex intervention for changing professional behaviour: the Evidence Based Out Reach (EBOR) Trial. J Health Serv Res Policy 2002; 7: 230-238. 0_CHDJIIHB
- Goldman L, Fiebach N. Hippocrates affirmed? Limiting residents’ work hours does no harm to patients. Ann Intern Med 2007; 147: 143-144. 0_i1115761
- Wachter RM, Pronovost PJ. The 100,000 Lives Campaign: a scientific and policy review. Jt Comm J Qual Patient Saf 2006; 32: 621-627. 0_i1115763
Provenance: Not commissioned; externally peer reviewed.