Volume 198 - Issue 8

Deciding when quality and safety improvement interventions warrant widespread adoption

Authors:  Ian A Scott and John B Wakefield

Med J Aust 2013; 198 (8): 408-410. || doi: 10.5694/mja12.10858
Published online: 20 May 2013
One tool that decisionmakers could use to assess whether such interventions are fit for purpose is a comprehensive checklist of evaluative criteria, starting with questions about how well the problem to be addressed by the intervention has been defined.

Evaluative criteria are needed to determine the likelihood of successful implementation and acceptable return on investment

Determining when a specific quality and safety improvement intervention (QSII) has sufficient evidence of effectiveness to warrant widespread implementation is highly controversial.1,2 Some large-scale QSIIs have been shown to be less effective than originally thought (Box).3-8 Reporting guidelines for QSII studies stipulate sufficient detail to allow users to gauge the feasibility and reproducibility of a specific QSII within local contexts.9 Some authors have focused on study designs and statistical methods used to evaluate QSIIs.10 An international expert group has distilled several key themes that researchers should consider and discuss when describing experiences with specific QSIIs.11

While considerable resources are being directed at QSIIs in Australia and elsewhere, recent literature reviews show significant shortcomings in research on QSII effects.12-14 Based on these reviews and our experience with various QSIIs, we propose a checklist of evaluative criteria that decisionmakers — clinicians, quality teams, policymakers and statutory bodies — can apply to existing literature relating to specific QSIIs to determine whether they are fit for purpose and whether widespread adoption is justified.

Checklist of evaluative criteria
5. Have the effects of the QSII been evaluated with a sufficient level of rigour?

Measuring the success of a QSII is prone to bias if it relies on qualitative self-reports of individuals directly involved in its design and implementation,21 so externally verifiable outcome data are preferred. Also, outcome measures should be standardised and appropriate, data should be collected accurately and comprehensively, and study designs should minimise the risk of confounding.

Were outcome measures standardised and appropriate? Well defined and objective patient-important outcomes minimise ascertainment bias3,22 (Appendix 2). QSII studies which report surrogate or intermediate outcomes (eg, change in medication error rates, or compliance with surgical site identification policies) should indicate how strongly such measures correlate with hard clinical end points. For example, improvements in “safety culture” as measured by survey tools show tenuous associations with reductions in patient harm.23

Were data collected accurately and comprehensively? The extent of inaccurate or missing data in many QSII studies is significant, and cherry-picked data from sites that perform better than others are often presented as the generalisable result. In addition, the intervention itself can alter how data are collected. For example, greater pharmacist participation in clinical teams may not only prevent prescribing errors but also unearth previously undetected errors.24

Did study designs minimise the risk of confounding? Investigators should have used study designs which minimise bias (Appendix 2). Randomised studies minimise selection bias in attributing improved patient outcomes to QSII effects. Cluster randomised trials involving multiple sites avoid contamination of control groups within sites. If extensive rollout of a QSII is already occurring or about to occur, then stepped wedge designs which insert randomisation into the phasing of implementation are preferred. Where randomisation is impractical, non-randomised studies (controlled before–after trials, interrupted time series studies, statistical process control charts) can be used.

Applying the checklist to a specific QSII

Before applying the checklist to a specific QSII, users must retrieve as much published evidence relating to the QSII as possible. By applying the checklist to this evidence, it is possible to build a profile of the QSII according to the evaluative criteria. Responses to the criteria may be dichotomous (yes or no) or, if the evidence is more subjective and uncertain, graded (using a 5-point Likert scale). Users may also wish to give different weightings to individual criteria depending on how critical they regard them to the overall utility of the QSII. We do not imply that all 12 criteria must attract favourable responses before proceeding with QSII implementation, although we feel that most QSIIs should satisfy Criteria 1–8. If they do not, we recommend that detailed longitudinal evaluations are undertaken as the QSII is implemented in pilot sites.

If a major quality problem is evident and requires urgent remediation, QSIIs that appear promising but have not been extensively evaluated may need to be considered. In such cases, we advise rigorous evaluation during implementation.

The strength of the checklist is that it encourages a structured appraisal of how QSIIs have been developed, implemented and evaluated. As an example, the checklist is applied to evidence around hospital rapid-response teams in Appendix 3. The results indicate that if this checklist had been available some years ago, it may have tempered early enthusiasm for rapid-response teams. The checklist also highlights the need for frequent and systematic evaluations of newly developed QSIIs. In an era of limited resources, the potential effectiveness and likely return on investment of specific QSIIs must be assessed. The checklist may contribute to greater discipline and transparency of investment decisions and help clarify which QSIIs require further refinement and testing before large-scale implementation.

Comparison of initial and later experience of two large-scale quality and safety improvement interventions

Rapid-response teams (RRTs)

RRTs are multidisciplinary teams of medical, nursing and airway management staff charged with prompt bedside evaluation, triage and treatment of clinically deteriorating patients throughout all hospital wards outside intensive care units (ICUs). Their aim is to reduce preventable deaths, cardiac arrest, unplanned ICU admissions and postsurgical complications.

Initial experience: Early trials suggested a large potential benefit of RRTs in reducing unexpected cardiac arrests (by up to 50%), unplanned ICU admissions (by up to 44%), postoperative deaths (by up to 37%) and mean length of hospital stay (by up to 4 days).3,4 As a result of such observations and advocacy for RRTs from the Institute for Healthcare Improvement’s 100,000 Lives Campaign, hundreds of hospitals worldwide have implemented RRTs.

Later experience: The validity of earlier positive observations has been challenged and a meta-analysis of 18 high-quality trials confirmed no reduction in mortality, although cardiac arrest calls were reduced by a third.5

Pay-for-performance (P4P) schemes

P4P schemes involve defined changes in reimbursement to clinical providers (individual clinicians, group practices or hospitals) in direct response to a change in one or more performance measures as a result of one or more practice innovations. Their aim is to incentivise optimal provider performance and improve quality and safety of care.

Initial experience: In the United Kingdom, large-scale implementation of P4P contracts for family practitioners over 12 months was reported in 2006 to have resulted in practitioners achieving a median of 97% of their available points covering quality of clinical care, well in excess of the predicted 75%.6 However, no baseline was established for most indicators. The United States Institute of Medicine and high-profile quality experts recommended greater use of P4P programs to improve quality of care, and by 2009 more than 200 P4P programs covering over 50 million beneficiaries were implemented.

Later experience: A review of 17 studies (12 controlled trials) showed modest improvement (4%–8% absolute increases) in some or all process-of-care measures in five of six studies of clinician-level financial incentives and seven of nine studies of group practice-level incentives.7 Four studies showed unintended adverse effects (gaming, patient exclusion, and tick-box documentation of undelivered care). A 2009 review of P4P schemes in the UK showed that, within 2 years of commencement, there was no further improvement in quality-of-care indicators despite a more than £1 billion budget overrun and a decline in continuity of care.


Authors


Competing interests


References


Provenance: Not commissioned; externally peer reviewed.