Volume 197 - Issue 2

A dog walking on its hind legs? Implications of the CareTrack study

Authors:  Ian A Scott and Christopher B Del Mar

Med J Aust 2012; 197 (2): 67-68. || doi: 10.5694/mja12.10957
Published online: 16 July 2012
Despite its limitations, this important study highlights a genuine need for systematised performance monitoring[It] is like a dog’s walking on his hinder legs. It is not done well; but you are surprised to find it done at all. Samuel Johnson1 The results of the CareTrack study, reported by Runciman and colleagues in this issue of the Journal,2 and the accompanying commentary by the study authors3 represent an ...

Despite its limitations, this important study highlights a genuine need for systematised performance monitoring

The results of the CareTrack study, reported by Runciman and colleagues in this issue of the Journal,2 and the accompanying commentary by the study authors3 represent an important milestone in health services research. The headline message — that the quality of care delivered in Australia is so poor that, on average, only 57% of episodes of care assessed by trained surveyors met the minimum standard set by the study authors (surprisingly similar to the headline from the United States study on which it was based4) — will cause consternation among clinicians and finger-pointing from the lay press. What should we make of this?

Firstly, is the study an accurate reflection of Australian health care? The sampling was the innovation — more than 1000 individuals with at least one of 22 conditions were sampled using the online White Pages and assessed in terms of the actual care they received for these conditions in more than 35 000 episodes of care.2 But several factors make us uneasy. White Pages sampling is likely to bias selection of subjects away from mobile phone users and those without any phone, whose health care quality might be different from those with landline phones. Moreover, after exclusions, only 1154 respondents from 15 292 telephone numbers called (7.5%) had index conditions and health care providers who agreed to allow access to their medical records. The fact that many health care providers declined to have their records audited introduces another source of bias.

The small sample meant that there were few patients, and correspondingly few encounters, for each condition (eg, 30 participants with heart failure, 28 with chronic obstructive pulmonary disease, 21 with pneumonia, 19 with stroke, and 12 with alcohol dependence). This resulted in very wide confidence intervals around the point estimates of compliance with the chosen standards, meaning there is considerable uncertainty about the real values. In turn, this precludes sensitivity analyses to check the robustness of the findings against any number of different parameters, such as type of records audited (structured summaries v unstructured notes), level of evidence of the standards (randomised trials v expert opinion), clinical relevance of the standards (high v low), or level of agreement between surveyors and their trainer (high v low). Very few medical records (a mean of 1.3 per participant) were assessed by the surveyors — even though most participants had seen at least two health care providers — meaning we have only a small snapshot of care. Finally, most care episodes were delivered by general practitioners (which is to be expected, as this reflects the Australian health care system), meaning that there were few encounters in referred care to assess.

The next problem centres on the standards used in the study to define quality indicators. Some were selected from evidence-based guidelines — but not many. Most were established by self-selected experts and supported by poor evidence. Consensus-based recommendations supported 369 (71%) of the 522 indicators. This means that clinicians will challenge not only the validity of some indicators (“There’s insufficient evidence for that”) but also the appropriateness of their selection (“That’s not a clinically important criterion”). Nor do the indicators reflect individualised patient care — something that guidelines acknowledge either explicitly or implicitly, and one reason why using guidelines as audit standards is often cautioned against.

Some indicators look suspiciously like things that are easy to measure, such as documenting discussion with men about prostate screening, rather than discriminatory indicators of clinically important care. Others are redundant, such as avoiding use of warfarin in patients with haemophilia or active bleeding. The indicators focus unduly on underuse rather than overuse of care (certainly harder to measure by these techniques), which may drive health care further into overuse — one of the prime causes of iatrogenic harm and waste.5 The equal weight assigned to all indicators will also irritate clinicians. For example, failing to use a CURB-65 prediction rule for pneumonia or discussing use of local heat or cold therapy for osteoarthritis are far less important than not appropriately managing a patient with a blood pressure level ≥ 180/110 mmHg. Lastly, many indicators may actually be measuring deficiency of the reporting in the medical record, rather than deficiency in the care. For example, advising patients to use hot or cold therapy for osteoarthritis or considering a reduction in a maintenance dose of inhaled corticosteroids for patients with asthma are unlikely to be recorded. This will be even more of a problem for assessing care that should not be provided, where recording the circumstances under which care was prudently withheld is even less likely. That there were difficulties in unpicking these issues is evident in the disagreement between the surveyors and the trainer in scoring clinical indicators, which was checked in a subset of records and yielded κ scores as low as 0.50.

Nonetheless, despite all these issues, CareTrack is an important study for Australia. Even with its limitations, the authors are right to conclude that Australian health care is suboptimal and that ongoing, systematised performance monitoring is needed to stimulate and document improvement.

So what should happen next? The study authors suggest nationally consistent digital records would facilitate such audits in the future (and perhaps the universal Personally Controlled Electronic Health Record will assist with this) and that a greatly rationalised national ethics approval system would mean that the vast army of ethics committees (over 220 in this study!)3 need not be individually applied to for such purposes. These are grand implementation visions, but the devil is in the detail. No country has come close to achieving both these goals at the national level. In the US, Intermountain Healthcare in Utah (covering one relatively small geographical region)6 and the Veterans Affairs health care system (covering one defined patient population)7 are good examples. However, both have taken decades to evolve, were internally driven (not externally by government), required transformative culture change, and rely on sophisticated information technology and linked datasets.

Runciman and colleagues also suggest ways of validating and updating quality indicators, for which they plan to use a web-based “wiki” tool.3 This may well be the major bottleneck. Getting a relatively small group of clinical experts to agree on quality indicators, even in a narrow field with reasonable-quality evidence, is challenging. Trying to achieve consensus among almost every auditable clinician, plus interested members of the lay public, for indicators that are evidence-poor sounds virtually impossible.

Nevertheless, the CareTrack team should be applauded for attempting this first census of health care quality. How the Australian Commission on Safety and Quality in Health Care and the National Health Performance Authority will progress such efforts, and whether they will use a top-down or bottom-up approach, will be interesting to watch.


Authors


Competing interests


References


Provenance: Commissioned; externally peer reviewed.

More like this