Topics
Statistics
Web and telecounselling in Australia
Ron Borland,* Catherine J Segan† * Nigel Gray Distinguished Fellow, † Behavioural Scientist, The Cancer Council Victoria, 1 Rathdowne Street, Carlton, VIC 3053 ron.borlandATcancervic.org.au To the Editor: The editorial on web and telephone counselling in Australia1 has the capacity to seriously mislead readers. It asserts: Despite this extensive use, the review confirmed that no randomised controlled trials (RCTs) have been conducted of the efficacy of web or telecounselling either in Australia or internationally.1 The assertion was based on a review commissioned by the Commonwealth Department of Health and Ageing and from a review in the United Kingdom, but is simply not true. It may be true for services designed to deal with mental health problems, narrowly defined, but it is patently false if it is taken to include services to facilitate smoking cessation. We note that nicotine dependence is a recognised mental disorder,2 and thus, strictly speaking, even if the review asserted that it was restricted to mental health, it would still be wrong. We do not know about the accuracy of the statements in relation to other drug use problems, but for smoking cessation there are a number of randomised trials of telephone-based systems,3 and at least two web-based resources are translations to the Internet of tailored computer advice services shown to be effective in RCTs.4,5 Both include Australian examples of RCTs. Our group demonstrated that the Quitline callback service as operated by Quit Victoria (phone 131 848) enhances cessation outcomes.6 Another study showed that an interactive personalised computer advice program called the QuitCoach (www.theQuitCoach.org.au) is effective in facilitating cessation, particularly by reducing relapse.4 It is currently available through the Department of Health and Ageing’s website at www.quitnow.info.au. We wonder why this omission has happened. What makes a health issue as important as smoking so invisible? Are other drug and alcohol issues similarly invisible? Tobacco kills about 19 000 Australians each year, and disables many more. There is increasing evidence suggesting it plays an important aetiological role in the development of some mental disorders. Smoking rates among people with schizophrenia and depression are extraordinarily high.7-9 The Victorian Quitline has pioneered the integration of support for psychiatric conditions with smoking-cessation counselling,10 and, although this service has not yet been subject to outcome evaluation, it is apparent that it meets the proximal needs of both smokers with concurrent mental disorders and their carers. Telephone and web-based services hold tremendous potential both as stand-alone services and as integrable components of comprehensive, coordinated care. High quality evaluations are required, and they need to be seen as an integral part of service delivery. People in other healthcare areas could learn a lot from what has been achieved in smoking cessation.
Ron Borland · Catherine J Segan
Web and telecounselling in Australia
Helen Christensen,* Barbara Hocking,† Dawn Smith‡ * Deputy Director, Centre for Mental Health Research, Australian National University, Canberra, ACT 0200; † Executive Director, SANE Australia, Melbourne, VIC; ‡ Chief Executive Officer, Lifeline Australia, Canberra, ACT. helen.christensenATanu.edu.au In reply: Borland and Segan are correct in assuming that we did not include substance disorder randomised controlled trials (RCTs) in the assessment of the efficacy of web and telecounselling services in our editorial.1 Our definition of web and telecounselling was also strict in that we included only contact that involved a person (a counsellor) online or by telephone. We specifically excluded interactive personalised web programs such as www.theQuitCoach.org.au or others specifically in mental health (narrowly defined), which have been found effective when delivered by the Internet. (Such programs include Panic Online,2 MoodGYM and BluePages.3) The use of RCTs in evaluating the areas of substance use, anxiety, depression and other mental health problems is to be applauded. Borland and Segan’s letter is also instructive in reminding us of the importance of coexistent substance-use disorders and mental health problems. Organisations such as SANE are committed to reducing the health costs of smoking in people with mental health problems and have developed specific programs for this purpose. Importantly, we are in agreement with Borland and Segan that telephone and web-based services “hold tremendous potential both as stand-alone services and as integrable components of comprehensive, coordinated care”. However, our editorial reported that web and telecounselling (not integrated web-based management systems) have yet to be evaluated through RCTs. One point we contest is the view that smoking is invisible. Our systematic review of funding allocations to mental health research has found that the category of substance-use disorders, of which smoking was the third-largest component (below alcohol and opioids), received the most Australian research funding in 2000.4 The level of funding for substance use exceeded that for childhood disorders and dementia. Compared with substance-use research, depression research received less than half, and psychosis and anxiety less than a third, of funding. Affective disorders contribute the highest disease burden, and dementia has the highest health system costs. Although all our projects in these important areas require proportionately more funding, it is not helpful to claim that the omission of smoking outcome research is due to failure to recognise its importance.
Helen Christensen · Barbara Hocking · Dawn Smith
The dearth of new antibiotic development: why we should be worried and what we can do about it
The emergence and spread of multidrug-resistant pathogens has increased substantially over the past 20 years. Over the same period, the development of new antibiotics has decreased alarmingly, with many pharmaceutical companies pulling out of antibiotic research in favour of developing “lifestyle” drugs. Reasons given for withdrawing from antibiotic development include poor “net present value” status of antibiotics, changes in regulations requiring larger drug trials and prolonged post-marketing surveillance, clinical preference for narrow-spectrum rather than broad-spectrum agents, and high new-drug purchase costs. Major improvements in infection control in Australia are needed to prevent further spread of resistant clones, buying some time to develop urgently needed new antibiotic agents. Perpetuating a culture of “pharma bashing” will simply lead to more pharmaceutical companies withdrawing from the market. A change in the health and research culture is needed to improve cooperation between public, academic and private sectors.
Patrick G P Charles MB BS, FRACP · M Lindsay Grayson MD, FRACP, FAFPHM
Generalising the results of trials to clinical practice
Randomised controlled trials should be the basis for developing clinical guidelines and for decisions about individual patient management. They should also inform public health policy. However, their capacity to fulfil these roles will depend on how closely a trial’s participants reflect the general population of patients with the disorder that has been investigated. The extent to which a trial’s findings are relevant to the broader population of patients with the disorder is referred to as the trial’s generalisability, or external validity. The CONSORT statement refers to generalisability under Item 21 (Box 1).1 Well-written reports should discuss the various factors that influence the generalisability of the trial’s findings. Julian and Pocock have proposed a checklist of questions to assist with this assessment (see Box 2).2 To determine the generalisability of a trial’s findings, several aspects require scrutiny. Is the patient population representative of the broad target group?To make this assessment, it is firstly necessary to examine the inclusion and exclusion criteria for the trial. These criteria determine the characteristics of the potential participants. They are particularly important for trials assessing new drugs, because patients with any significant degree of renal or hepatic impairment, or any significant comorbidity, are often excluded. Excluding such participants may result in a trial population that represents only a subsection of the broader population with the disorder for which the drug may be indicated. Secondly, the baseline data should describe the population that participated in the trial. Demographic variables, age range, as well as clinical data such as blood pressure, staging of disease and any listed comorbidities, will help readers decide whether the trial population closely resembles the patient population (or individual patient) for which a decision about management is required. In some trials, the entry criteria are considerably broader than the population actually recruited. This discrepancy will only be evident if adequate baseline data are presented. For example, in a study of combination chemotherapy in malignant breast cancer, no radiotherapy was allowed in the protocol.3 Results were reported on the basis of the extent of lymph node involvement — 0–3 nodes or 4 or more — and readers might assume that the findings of the trial would apply to participants with many (more than 10) involved nodes. However, less than 8% of patients with more than 10 involved nodes were included in the study, as clinicians referred these higher-risk patients for radiotherapy rather than enrolling them in the trial.4 The original trial report did not tell readers that the group with more than 4 involved nodes actually comprised patients with primarily 4 to 10 involved nodes.5 Participant flow diagramThese diagrams are useful for assessing generalisability of trials. If properly completed, flow diagrams will indicate the number of participants: screened for participation; with the condition of interest; classed as ineligible (on the basis of exclusion criteria); and who did not elect to participate. If the population randomly allocated to groups within the trial represents only a small proportion of those with the condition of interest and assessed for eligibility, it is probable that the generalisability of the findings of the study will be limited. Large trials assessing warfarin therapy for atrial fibrillation have enrolled only about 15% of those who were potentially eligible, and this substantially limits the generalisability of their results.6,7 Where eligible patients who entered randomised trials have been compared with those who were eligible but did not participate, differences have emerged. In a study of therapy for temporomandibular disorders, 18 eligible patients did not consent and 60 were randomly allocated to trial arms.8 The 18 patients who did not consent reported more pain than those who participated, perhaps restricting the findings of the trial to those with milder pain. Differences were also evident between enrolled and unenrolled patients in the Thrombolysis in Myocardial Infarction (TIMI 9) trial.9 The TIMI 9 registry prospectively evaluated patients with ST-segment-elevation myocardial infarction. There were no exclusion criteria for the registry, but there were exclusion criteria for the randomised trial. Patients in the registry, but not enrolled in the trial, had higher baseline risk for adverse outcomes. Screening logsScreening logs list the numbers of individuals screened, eligible and enrolled, as well as reasons for not enrolling eligible patients. They thus allow readers to judge whether there are differences between patients who were and were not enrolled in the trial. If the two populations are similar, the generalisability of the trial is increased. A template of the typical information collected in the screening log is presented in Box 3. This information should be limited to the most important characteristics of the relevant population to minimise the burden on trial staff collecting the data. Screening logs also describe the range of participants with the disease being seen at each investigation site, and the patterns of care of these people. For those deemed ineligible, the criteria excluding them from the study are documented, providing further information as to the generalisability of the intervention to this cohort.10 The results of the study will apply more to subjects excluded because they were not available for follow-up or were just outside the age range than to those with concomitant disease. ComorbiditiesIn some trials, during random allocation, patients may be stratified by comorbidities regarded as potential confounders. In others, particularly in clinical trials of new drugs, comorbidities may be exclusion criteria. If these comorbidities are relatively common, the exclusion criteria will significantly limit the generalisability of the trial outcomes. This is an important issue in drug trials, as comorbidities are often exclusion criteria. When the drug is registered for use, the listed indication is often relatively broad, so it is necessary to scrutinise the clinical trials section of the product information to obtain a clearer picture of the patient population which was studied. If a patient for whom the drug is being considered has one of the comorbidities which was an exclusion criterion, whether the drug will be efficacious or safe is unknown. Such a dilemma exists with the “statin” lipid-lowering drugs. Over 160 000 patients have participated in trials of statins, but almost all of these trials have excluded patients with significant renal or hepatic disease.11-13 If some patients have characteristics which were not reported in the trial’s population, it is conceivable that the trial’s results are not relevant to these patients. Subgroup analysesAs adolescents and pregnant women are not usually included in trials, most clinical trials are of limited relevance to these groups. If there are a priori reasons to expect differences between subgroups, the trial may have stratified participants by these potential confounders. The findings of the trial will be most generalisable if the benefits are evident in each subgroup of the trial, as well as across the entire study population.14 In clinical trials with large numbers of patients or events, it is possible to have reliable subgroup analyses which may help prescribers to relate the trial’s findings more closely to patients for whom they are trying to select appropriate therapies. ConclusionsThe main purpose of conducting randomised clinical trials is to identify improvements in clinical care. Ideally, the findings of trials should apply to a wider population than those included in the trial. It is therefore vital that every effort is made to have a broad selection of patients from the population of interest to minimise selection bias. Numerous exclusion criteria will restrict the patient population and progressively diminish the generalisability of the findings of the intervention under evaluation. It is the responsibility of those reporting trials to include aspects of generalisability when discussing their findings. It is the task of those responsible for treating patients, producing clinical guidelines and formulating public health policy to carefully assess the generalisability of clinical trials before applying their findings.15 1 CONSORT checklist of items to report when reporting a randomised trial.1 Section and topic Item no. Descriptor Discussion Generalisability 21 Generalisability (external validity) of the trial findings. 2 Concepts covered in Julian and Pocock’s criteria for assessing generalisability2 Representativeness of patients for the condition in practice. Proportion of eligible patients participating. Conformity of the treatments and background care (doses, durations, follow-up period, etc) to standard practice patterns. Consistency of measured outcomes with conclusions drawn. Appropriate balance of surrogate and clinical outcomes. Reliability of evidence on efficacy and safety findings. Coverage of all relevant outcomes (adverse events and side-effects). Consideration of the study findings in the context of other available evidence. 3 Screening log template Site identification/name: ______________________________________ Date of visit: __________________________________________ Subject initials: ____________ Clinician initials: ____________ Age: ____________ Sex: ____________ Is the subject recruited into the study? yes / no If no: then: a) Main reason for exclusion: ______________________________ _____________________________________________________ _____________________________________________________ Reasons for exclusion: 1. Subject refusal 2. Clinician refusal (with possible reasons) 3. Ineligible (specify which inclusion criteria are not met and which exclusion criteria exist) 4. Language difficulties 5. Other b) Treatment actually given: _______________________________ _____________________________________________________ _____________________________________________________
J Paul Seale PhD, FRACP, FRCP · Val J Gebski MStat · Anthony C Keech FRACP, MClinEpi
Predicting death in young offenders: a retrospective cohort study
Objective: To examine predictors of death in young offenders who have received a custodial sentence using data routinely collected by juvenile justice services.Design: A retrospective cohort of 2849 (2625 male) 11–20-year-olds receiving their first custodial sentence between 1 January 1988 and 31 December 1999 was identified.Main outcome measures: Deaths, date and primary cause of death ascertained from study commencement to 1 March 2003 by data-matching with the National Death Index; measures comprising year of and age at admission, sex, offence profile, any drug offence, multiple admissions and ethnic and Indigenous status, obtained from departmental records.Results: The overall mortality rate was 7.2 deaths per 1000 person-years of observation. Younger admission age (hazard ratio [HR], 1.4; 95% CI, 1.0–1.9), repeat admissions (HR, 1.8; 95% CI, 1.1–2.9) and drug offences (HR, 1.5; 95% CI, 1.0–2.1) predicted early death. The role of ethnicity/Aboriginality could only be assessed in cohort entrants from 1996 to 1999. The Asian subcohort showed higher risk of death from drug-related causes (HR, 2.5; 95% CI, 1.1–5.5), more drug offences (relative risk ratio [RRR], 13; 95% CI, 8.5–20.0) and older admission age (oldest group v youngest: RRR, 9.3; 95% CI, 1.3–68.0) than non-Indigenous Australians. Although higher mortality was not identified in Indigenous Australians, this group was more likely to be admitted younger (oldest v youngest: RRR, 0.31; 95% CI, 0.15–0.63) and experience repeat admissions (RRR, 1.6; 95% CI, 1.0–2.4).Conclusions: Young offenders have a much higher death rate than other young Victorians. Early detention, multiple detentions and drug-related offences are indicators of high mortality risk. For these offenders, targeted healthcare while in custody and further mental healthcare and social support after release appear essential if we are to reduce the mortality rate in this group.
Carolyn Coffey BSc, GradDipEpi · Andrew W Lovett FRACP · Eileen Cini BSc(Hons) · George C Patton MD, FRANZCP · Rory Wolfe PhD · Paul Moran MD, MRCPsych
Research ethics committees: what is their contribution?
Perhaps a week of intensive training in critical thinking would be the best preparation for members of research ethics committees In a recent lecture at Monash University, the philosopher Raimond Gaita, Professor of Moral Philosophy at King’s College, University of London, and Professor of Philosophy at the Australian Catholic University, told the story of a woman facing a significant turning point in her life. It was a Friday, and a decision was needed by Monday, but she had unavoidable obligations over the weekend. She had a dear friend, a psychoanalyst and philosopher who had known her all her life. He knew her circumstances, her preferences and even her secret wishes. She contacted him and prevailed upon him to make the decision for her. In our private lives most of us would find it at least odd, and probably uncomfortable, to hand over responsibility for significant decisions to others. Yet, the prevailing paradigm for human research ethics committees has institutionalised this approach. Researchers themselves often do not consider the ethical implications of their work until it is time to fill out the various forms required by committees. Even then, the main concern is “getting through ethics” with minimal scarring of their proposal. The Nuremberg Code,1 the Helsinki Declaration,2 and even the National Health and Medical Research Council’s National statement on ethical conduct in research involving humans (the Statement),3 are not documents with which many researchers can claim significant familiarity. The reasons for their existence are faintly recalled, and current debates are only of interest if they impede research with which the researcher has a personal concern. Once an ethics committee has made its decision, there is no need to consider “ethics” again unless there is a significant adverse event. Although they undoubtedly provide a “safety net” to detect and prevent grossly unethical research, ethics committees must not and cannot be seen as the repositories for moral decision-making. Consider an imaginary (but highly plausible) ethics committee. It meets the Statement’s requirements for membership. Some members have attended the occasional seminar sponsored by the Australian Health Ethics Committee (AHEC), and some diligently read the AHEC Bulletin sent to registered committees. Some, though not all, of the members have actually read the Statement all the way through. One committee member doesn’t really agree with some of the content. The committee faces regular criticism by researchers for the amount of paperwork that must be submitted to it, and significant anger when it wishes to alter an aspect of a proposal for a multicentre trial. Although the committee’s deliberations are thorough, most of its recommendations consist of minor changes to the plain language and consent statement. It faces considerable (and understandable) pressure to reach rapid consensus. Rarely does a member ever register his or her dissent concerning a decision about which all other members of the committee feel comfortable. The committee is proud that it has never ultimately rejected any proposal. Some members of the committee are aware that they have acquiesced in decisions about which they had some misgivings. One or two of the most senior members know they can nearly always sway the committee to their point of view. In a recent editorial discussing clinical ethics committees, Margaret Somerville noted that: Committee decisions, as compared with individual ones, can spread the responsibility. A committee can make a decision that no one person — in particular, no committee member — acting alone would make.4 She uses the real-life example of decisions to shorten life by withholding treatment, or aborting a fetus, and the physicians doing this being morally reassured by the involvement of an Acute Clinical Ethics Service. She asks Might [this involvement] have allowed the caring team to implement decisions that their moral intuitions were indicating were unethical? While these decisions may have been ethical, we must always be aware that we ignore such intuitions at our ethical peril.4 Although clinical ethics committees perform a somewhat different function to human research ethics committees, there are significant similarities as far as the points made by Somerville are concerned. There is indeed comfort in allowing ourselves to be relieved of having to think about the implications of our actions, especially when the research dollar is concerned. It is problematic when the ethics review process is seen as a test of how much we are able to get away with. It cannot be persuasively argued that actions must be ethical because they have been approved by another person thought to be morally wiser. Nor is it helpful to allow oneself to be influenced to reach a decision in a group setting because of reasons of time, or because others have already reached consensus. If researchers have not thought through the ethical implications of their proposals, but instead leave that to the committee, and the committee makes decisions about which some of its members would be individually uncomfortable, this cannot be regarded as a satisfactory process. Perhaps the most essential preparation for members of research ethics committees is not studying the content of the Statement or the relevant law, but undertaking a week of intensive training in critical thinking. Perhaps we all must consider how best to deal with situations about which not all agree, and about which objections are morally relevant. Furthermore, there are many issues that are not well addressed by guidelines or law. What research should be done in the first place? How should communities from which participants are drawn be involved in the planning, implementation, monitoring and evaluation of research? What are the human rights implications of a study (particularly in populations significantly deprived of rights)? What responsibilities do researchers have to the larger community from which their subjects are drawn, and what do they owe to subjects after their research is completed? These questions and many more have important ethical dimensions. Many researchers are unaccustomed to thinking through the broader implications of their work. However, they are capable of doing what is necessary in order to fill out a form. In a seminal article in the New England Journal of Medicine, Henry Beecher stated The ethical approach to experimentation in man has several components; two are more important than the others, the first being informed consent . . . Secondly there is the more reliable safeguard provided by the presence of an intelligent, informed, conscientious, compassionate, responsible investigator.5 In concluding his lecture, Gaita rejected the notion of a “moral expert”, and called for us all to identify and rigorously analyse morally important issues without sentimentality. It is very difficult to improve on this.
Bebe Loff MA, LLB · Jim Black MB BS, PhD, FAFPHM
Multiple analyses in clinical trials: sound science or data dredging?
Clinical trials typically require the collection of many data to describe the participants and for measuring their response to an intervention. In addition to the primary analysis of treatment effect, investigators can use these data to perform multiple analyses, but there are important pitfalls with their use.1,2 Here, we discuss three common types of secondary analyses: analyses of multiple outcome variables; analyses of trial outcomes that account for prognostic factors (adjusted analyses); and using trial data to answer secondary research questions (see definitions in Box 1). The use of trial data for population subgroup analyses has been discussed earlier in this series.3,4 What are the problems?The two main problems introduced by multiple analyses are, firstly, the increased probability of detecting intervention effects where none exist (“false positives” owing to multiple comparisons — type I errors), and secondly, the limited capability (“power”) of trials to detect a true treatment effect in secondary outcomes if not enough participants are enrolled to show a statistically significant difference in these outcomes (“false negatives” — type II errors). One study compared trial protocols with their subsequent publications, and provided empirical evidence of the selective reporting of positive trial results.5 The use of multiple analyses is therefore of particular concern when these are conducted post-hoc as a “fishing expedition”, and undue emphasis is given to positive findings. Item 18 of the CONSORT checklist (Box 2) recommends that investigators report on all multiple analyses and declare which were prespecified and which were conducted as exploratory activities after investigators were unblinded to the treatment allocation of participants.6 Analyses of multiple outcomesAdvantagesInvestigators may choose multiple outcomes to measure the effect of treatment. This is an advantage when different parameters provide information about different aspects of the treatment response.1 Secondary analyses may also assist the interpretation of the primary analysis. For example, in a recently reported trial comparing chemotherapy regimens in the treatment of patients with metastatic breast cancer, investigators selected two primary outcomes as the most important measures of treatment effect: the overall tumour response rate (measured as complete and partial response) and time to treatment failure. Secondary outcomes, including overall survival, toxicity and quality of life, provided additional information about the treatment effect.7 Including a set of supplementary outcomes may also be a practical solution when different investigators value outcomes differently. Multiple-outcomes analysis is particularly useful when a statistically significant benefit of treatment on the primary outcome can be confirmed or strengthened by a consistent effect on other relevant outcomes. The Long-Term Intervention with Pravastatin in Ischaemic Disease (LIPID) trial evaluated the effectiveness of pravastatin for preventing cardiovascular events in patients with diabetes or impaired fasting glucose and a history of coronary heart disease.8 The finding of a statistically significant reduction in the risk of a major coronary event was supported by a similar reduction in the risk of a revascularisation procedure or stroke. Such findings may also advance the understanding of the relationships between outcomes. PitfallsThe type and number of analyses performed should be reported so that readers can assess the probability of detecting a treatment effect by chance alone. This is often poorly documented in trial reports.5 Additionally, major discrepancies have been observed between the primary outcome specified in the trial protocol and that reported in the published article.5 Chan et al reported that of 76 trials that prespecified a primary outcome in the trial protocol, 20 (26%) did not report on this outcome in the published article, and of 63 trials that specified a primary outcome in the published article, in 11 (17%) it was not mentioned in the trial protocol.5 An example of the latter was a study reporting on the percentage of patients with graft occlusion as the primary outcome, even though the study was not originally designed to measure a difference in this outcome.5 Caution is needed when unexpected results from multiple analyses are interpreted. Inconsistent results are more credible if the outcome variables are restricted to those that were prespecified in the trial protocol, are clinically relevant, and are based on plausible biological mechanisms. A variety of statistical corrections can be performed to take into account the increased probability of a chance finding with multiple testing. A statistically significant treatment effect for one outcome and not other clinically related outcomes may also indicate that the sample is too small — that is, the study lacks power. Interpreting the results of analyses that are underpowered is difficult. This is a common problem, for example, in the mandatory reporting of adverse events in drug trials. A trial may report on a large set of adverse events, but it will commonly not have been powered to detect a statistically significant difference in these outcomes between the trial’s study groups.1 Composite endpointsTo overcome the problem of insufficient power, investigators may combine data from clinically related outcomes to form one or more composite endpoints. This approach reduces the number of analyses required while retaining all the potentially valuable information. The TAXUS IV trial was a double-blind randomised controlled trial to determine the safety and effectiveness in coronary artery disease of paclitaxel-eluting stents compared with bare metal stents.9 The primary outcome of the trial was the incidence of revascularisation procedures due to reocclusion of the target vessel at 9 months. A composite endpoint, “major adverse cardiac events” (defined as death from cardiac causes, myocardial infarction, or revascularisation procedures), was a secondary outcome. At one year after the procedure, the rates of cardiac death and myocardial infarction were similar between the study groups, while the rate of target-vessel revascularisation was 62% lower (P < 0.0001) in the patients receiving paclitaxel-eluting stents than in those receiving bare metal stents.9 The treatment effect on revascularisation rates appeared to drive the results for the composite endpoint, resulting in a reported 49% reduction in major adverse cardiac events (P < 0.0001) at 12 months. Combining disparate events can lead to an overestimate of the clinical importance if a positive finding is largely driven by the less important events. Overall, as a result of the potential to overinterpret or misinterpret the results of analyses of multiple outcomes, readers should seek information in the methods section of the report about the primary purpose and outcomes that the trial was designed to address and interpret any additional findings in this context. Adjusted analysesClinical trials use a concealed randomisation process, with or without stratification by key prognostic factors, such as age and sex, to help to ensure the baseline similarity of the study groups.10 However, even well-conducted random allocation may still result in chance imbalances.11 If an imbalance in an important prognostic factor occurs, statistical methods can control for this imbalance by including the factor as a “covariate”. This is referred to as adjusted analysis, or a multivariate analysis if more than one covariate is included. While adjusted analyses can statistically accommodate imbalances between study groups in non-randomised studies, in randomised studies they should usually be considered supplementary to the unadjusted analysis of the primary outcome. If the adjusted effect estimate differs from the unadjusted estimate, interpretation may be a problem. For example, if some covariate data are missing and these participants are excluded from the adjusted analysis, it will not be clear whether observed differences result from controlling for this factor or another, unknown effect of these exclusions. Adjusted analysis may be indicated when a factor is known to strongly predict the outcome (for example, age and survival), even when the imbalance observed between study groups does not reach statistical significance.12 In general, adjusted analyses frequently improve the precision of the estimate of treatment effect, even when the correlation of the covariates with the study outcome is not strong.13 Another recent review of 50 consecutive published trials showed that the methods and reporting of adjusted analyses vary widely in clinical trials.14 Of the 36 trials with an adjusted analysis of the primary outcome, 42% did not report on the methods used to select the covariates.14 Using inappropriate methods for adjusted analyses may cause inaccurate and misleading results. Ideally, investigators should prespecify any prognostic factors that, if unbalanced, may affect the study outcomes and should plan for adjusted analyses accordingly. However, some strong predictors of the outcome may only become apparent at data analysis on formal testing (so-called exploratory analysis). In this situation, investigators should clearly describe when and how covariates were selected for the adjusted analysis. In any case, the primary emphasis should be on the unadjusted results, because investigators are able to conduct multiple adjusted analyses using different sets of covariates, which may lead to overinterpretation or selective reporting of significant findings. The findings of the primary unadjusted analysis are strengthened if the results of the adjusted analysis are consistent with them. Other ancillary analysesA clinical trial may seek to address ancillary questions unrelated to the primary question so as to optimise the use of resources required for a large clinical trial. Ancillary questions may relate to the treatment effect on other conditions of interest, such as the association between hormone replacement therapy and dementia in women recruited to a large trial investigating hormone replacement therapy and cardiovascular disease.15 Substudies may also use trial data to investigate epidemiological questions about the natural history of disease, the biological mechanisms of the disease16 or the treatment response.17 Ideally, these ancillary studies should be designed before the trial starts. However, important new information or scientific debate may arise during or after the trial to justify the use of trial data to investigate new hypotheses. Their results are more convincing if the decision to conduct the analysis has been made before unblinding. The same potential for overinterpretation and selective reporting of the results of multiple comparisons and reduced power apply to exploratory analysis, and any new findings should be regarded as new hypotheses for validation in future studies. The principles of planning, reporting, analysing and interpreting multiple analyses are shown in Box 3. These are not intended to discourage investigators from conducting potentially important exploratory analyses of plausible new hypotheses. Rather, they encourage the balanced reporting of all analyses to prevent unsound manipulation of data or undue emphasis on particular findings that may misdirect future research or compromise the interpretation of results for clinical practice. 1 Definitions Primary outcome: The health parameter measured in all study participants to detect a response to treatment. Conclusions about the effectiveness of treatment should focus on this measurement. Primary analysis: The statistical test performed to determine whether there is a difference in the primary outcome between participants allocated to receive the treatment and those allocated to the control arm. Secondary outcomes: Other parameters that are measured in all study participants to help describe the effect of treatment. Baseline variables: The characteristics of each participant measured at the time of random allocation. This information is documented to allow the trial results to be generalised to the appropriate population/s. Specific characteristics associated with the patient’s response to treatment (such as age and sex) are known as prognostic factors. Multiple analyses: Comparisons between the study groups for more than one outcome. They increase the likelihood of detecting a difference between the treatment and control group owing to chance alone (false positive). Common examples of multiplicity in trials include the use of: multiple outcomes, including surrogate endpoints; multiple treatment comparisons (in a multiarm trial); subgroup analyses to detect differences in the treatment effect in one or more subsets of trial participants; adjusted analyses to control for imbalances in prognostic factors between the study groups; repeated measures over time of the same outcome; and interim analyses of the treatment effect at different stages in the trial. Exploratory analyses: Analyses that were not specified before the trial or, for blinded studies, analyses planned after the investigators were unblinded to the treatment allocation of participants. These analyses may be driven by the results of the primary analysis. 2 CONSORT checklist of items to include when reporting a trial6 Selection and topic Item no. Descriptor Ancillary analyses 18 Address multiplicity by reporting any other analyses performed, including subgroup analyses and adjusted analyses, indicating those prespecified and those exploratory. 3 Checklist for multiple analyses Design and methods Were the primary and secondary outcomes for the detection of treatment response prespecified? Was the trial designed to have adequate power for the analyses of all outcomes? Were the covariates for the adjusted analyses and/or the method used to select these covariates prespecified? Were the substudies based on an existing trial or biological data? Were the substudies planned prior to unblinding of data? Analysis Have corrections for multiple-significance testing been performed? Was the combination of data into a composite outcome appropriate? Was the interpretation of the composite endpoint results appropriate? Reporting Are the total number of analyses performed reported? Was the power calculation reported for the primary outcome? Secondary outcomes? Are the rationale and methods of any adjusted analyses reported? Are the number and type of covariates in the adjusted analyses reported? Are the unadjusted and adjusted results reported? Are the prespecified analyses clearly distinguished from the exploratory analyses? Interpretation Is appropriate emphasis given to the primary outcome? Have the relationships between interrelated outcomes been explored with equal interest? Are the findings of the multiple analyses discussed in the context of current biological knowledge and current research?
Sarah J Lord MB BS, MScEpid · Val J Gebski BA, MStat · Anthony C Keech MScEpid, FRACP
Breast cancer in Western Australia: clinical practice and clinical guidelines
Objectives: To review changes in patterns of care for women with early invasive breast cancer in Western Australia from 1989 to 1999, and compare management with recommendations in the 1995 National Health and Medical Research Council guidelines.Design and setting: Population-based surveys of all cases listed in the Western Australian Cancer Registry and Western Australian Hospital Morbidity Data System.Main outcome measures: Congruence of care with guidelines.Results: Data were available for 1649 women with early invasive breast cancer (categories pT1or pT2; pN0 or pN1; and M0). In 1999, 96% had a preoperative diagnosis by fine-needle aspiration or core biopsy (compared with 66% in 1989), with a synoptic pathology report on 95%. Breast-conserving surgery was used for 66% of women with mammographically detected tumours (v 35% in 1989) and 46% of those with clinically detected tumours (v 28% in 1989), with radiotherapy to the conserved breast in 90% of these cases (83% in 1989). Adjuvant chemotherapy was given to 92% of premenopausal women with node-positive disease and 63% with poor-prognosis node-negative tumours (v 78% and 14%, respectively, in 1989). Among postmenopausal women with receptor-positive tumours, tamoxifen was prescribed for 91% of those with positive nodes (85% in 1989) and 79% of those with negative nodes (30% in 1989). Among postmenopausal women with receptor-negative tumours, chemotherapy was prescribed for 70% with positive nodes (v 33%) and 58% with negative nodes (v none).Conclusions: Patterns of management of women with early invasive breast cancer in Western Australia during the 1990s changed significantly in all respects toward those recommended in the 1995 guidelines.
Suzanne P McEvoy MAppEpid, FAFPHM · Claire Haworth RN · Jennett M Harvey FRCPA · Lin Fritschi PhD, FAFPHM · David M Ingram MS, FRACS · Michael J Byrne BMedSci, FRACP · Joanna Dewar FRACP · David J Joseph FRANZCR · James Trotter MD, FRACP · Chris Harper FRANZCR · Greg F Sterrett FRCPA, FIAC · Konrad Jamrozik DPhil, FAFHM
Complementary medicine research in Australia: a strategy for the future
Research funding for CAM is inadequate, resulting in too few good quality studies to support its use. Widespread use of CAM, as well as its media promotion, make this a vital public health issue, and the Australian government has a social and ethical obligation to respond by developing a research infrastructure (as has been done by the United Kingdom and United States governments). We propose a funding model that neither draws directly from the CAM industry nor from current health research budgets, yet would strengthen Australia’s international role in CAM research. Establishing and applying focused research methods in CAM is imperative for strengthening its evidence base and creating fresh options for safe and effective patient care.
Alan Bensoussan PhD, MSc · George T Lewith DM, FRCP
Complementary and alternative medicine: the convergence of public interest and science in the United States
CAM research is leading to changes in the vitamin cabinet and the clinic Many Americans use one or more health promotion, illness prevention or healing practices that are considered as complementary and alternative medicine (CAM).1 In recognition of this, the United States Congress legislated in 1991 to establish the Office of Alternative Medicine to “investigate and evaluate promising unconventional medical practices”. In 1998, Congress expanded this mandate by enacting legislation that created the National Center for Complementary and Alternative Medicine (NCCAM), endowing it with the resources and authority to fund research, train researchers, and disseminate information to the public and healthcare professionals. A number of factors contributed to the creation of NCCAM. First was the popularity of unproven medical practices, with users of one or more CAM modalities tending to be women, people with higher education, those with an interest in the role of the mind in health, or those with some chronic illness.2 Second was an increasing recognition of the importance of traditional healing practices among an ethnically diverse American population. Third, in 1994, the US Congress passed legislation that permitted wide access to dietary supplements without confirmation of their composition, safety or efficacy, and a concomitant loosening of legal restraints on alternative practices such as chiropractic medicine and acupuncture. Setting priorities — science firstGiven the diversity of CAM approaches and questions about their safety and efficacy, setting research priorities for NCCAM is a significant challenge. The US$117.7 million allocated to NCCAM in 2004, while generous by most standards, permits only a limited sampling of possible CAM approaches. To develop its approach, NCCAM sought input from diverse communities of stakeholders. The resulting first strategic plan stressed investment in basic and clinical research, training, dissemination of findings, and integration of safe and effective practices.3 NCCAM made a commitment to aspire to the same rigorous standards that characterise National Institutes of Health (NIH) research in general, while its research priorities would focus on the most promising scientific opportunities. As NCCAM celebrates its fifth anniversary, it is possible to list its not inconsiderable achievements to date (see Box 1), and to reflect on some of the lessons learned for current and future directions. Investing in research centresTo create a sustainable research infrastructure, NCCAM funded a first generation of research centres spanning a range of health disorders and disciplines. Based on formal reviews, a second generation of more focused centres is being developed. Some of these are involved in elucidating mechanisms of action of CAM therapies. Others promote collaborations between CAM and conventional institutions. Still others represent new initiatives to forge scientific partnerships between investigators at US and foreign institutions. Botanical trials — overcoming obstaclesWhen NCCAM was created, it was assumed that existing literature on herbal supplements would be sufficient to justify and design major studies of their safety and efficacy. It quickly became apparent that many botanical preparations are not standardised, and may be contaminated with heavy metals or drugs,4 precluding the conduct of meaningful and ethical studies. As a result, NCCAM is now working with academic and industrial partners to identify more optimal research-grade materials, and requires evidence of product quality for all its sponsored research. These approaches raise the quality of the studies, but consume time and resources. Phase III clinical trials — balancing pressure for progressThe results of NCCAM’s very first large clinical trial dictated a more deliberate and phased approach for future studies, even when high quality products are available. The three-arm randomised controlled trial failed to show that Hypericum perforatum (St John’s wort) ameliorates major depression.5 Advocates for the product faulted the study for having addressed too serious a form of depression. While the target population, the product, its dose, and endpoints had been thoroughly discussed, it was ultimately clear that more preliminary research and consensus development was needed to determine the optimal design of other large trials. Such efforts are being made now in trials of Ginkgo biloba for cognitive decline in the elderly, and glucosamine for osteoarthritis. Each of these, and other ongoing trials (Box 2), are being conducted with input from relevant communities of patients, practitioners and scientists, and cofunded by other NIH institutes. Brain–mind–body medicine — an emerging scienceThe capacity of the brain and mind to affect health is a CAM domain that is receiving more attention at NCCAM. Surveys indicate that about one in five adults use at least one mind–body therapy.6 Functional neuroimaging provides powerful new tools to identify changes in brain structures involved in generating emotional responses, interaction of distress and pain, and response to treatment. With this and other new laboratory and ambulatory methods, NCCAM is funding investigators to identify pathways of influence, and to test interventions, such as meditation in preventing illness, slowing disease progression and promoting well-being. Planning for the futureLessons learned at NCCAM forecast a future with a greater emphasis on preclinical and early-phase clinical studies that are designed to elucidate mechanisms, identify optimal dosing and schedules, and select appropriate target populations and control conditions before launching clinical trials. Clinical studies at all levels are increasingly being conducted as collaborative efforts between funded investigators and NCCAM staff, who provide technical guidance, from sophisticated design and statistical consultation, to advice on recruitment and retention. Research on natural products is being conducted within a framework in which NCCAM either provides well-characterised and standardised clinical trial materials for investigators to use, or tests products being used by investigators to assure characterisation and standardisation. Finally, NCCAM is taking full advantage of new technologies, from genomics to brain imaging, and applying them to new areas, such as the capacity of the mind to affect health, in its multidisciplinary research. While NCCAM first built a domestic scientific constituency, it seeks to engender and strengthen relationships with other countries that have both established research in conventional medicine, and a tradition that is rich in indigenous practices. We invite scientific leaders in these countries to join the global CAM research effort. 1 Activities of the National Center for Complementary and Alternative Medicine in its first 5 years It built a centre responsive to its mission and integrated into the other institutes at the United States National Institutes of Health It funded over 780 projects at 123 institutions, resulting in over 700 scientific publications It awarded more than 100 individual doctoral and postdoctoral training and career awards It enrolled nearly 40 000 participants in clinical protocols It received over 1.5 million visitors to the website <www.nccam.nih.gov> each year who search for information about CAM, clinical trials, and research opportunities It developed a database known as “CAM on PubMED” that lists nearly 400 000 articles on CAM-related subjects published in 45 languages from 70 countries It informed public policy, patient choice, and clinical practice through outreach activities, including public town meetings, public media, and scientific and professional conferences 2 Status of National Center for Complementary and Alternative Medicine Phase III Clinical Trials Complementary and alternative medicine modality Target disease Sample size Status National Institutes of Health Partner Acupuncture Osteoarthritis 570 Trial complete; analysis underway NIAMS Glucosamine/chondroitin Osteoarthritis 1 588 Enrolment complete; ongoing NIAMS Ginkgo biloba Dementia 3 073 Enrolment complete; ongoing NINDS, NIA, NIMH Shark cartilage Lung cancer 756 Patients enrolling; ongoing NCI Vitamin E Prostate cancer 32 400 Enrolment complete; ongoing NCI St John’s wort Minor depression 300 Patients enrolling; ongoing NIMH, ODS EDTA chelation therapy Coronary artery disease 2 372 Patients enrolling; ongoing NHLBI Saw palmetto Benign prostatic hyperplasia 2 860 Final protocol under development NIDDK, ODS NIAMS = National Institute of Arthritis and Musculoskeletal and Skin Diseases; NINDS = National Institute of Neurological Disorders and Stroke; NIA = National Institute on Aging; NIMH = National Institute of Mental Health; NCI = National Cancer Institute; ODS = Office of Dietary Supplements; NHLBI = National Heart, Lung, and Blood Institute; NIDDK = National Institute of Diabetes and Digestive and Kidney Diseases.
Margaret A Chesney PhD · Stephen E Straus MD
Balancing the outcomes: reporting adverse events
When decisions about a new intervention are being made, the “net clinical benefit” of the intervention needs to be assessed. This requires balancing all the reported benefits and side effects of the intervention. The adverse events experienced in a trial must be known in sufficient detail for their severity and relationship to treatment allocation to be judged. Reporting of such events is the subject of item 19 of the CONSORT statement (Box 1).1 What is an adverse event?The definition of adverse events (AEs) adopted by the International Conference on Harmonization (ICH) is shown in Box 2; it is designed to document all untoward events occurring in a clinical trial.2,3 AEs are thus both those events for which there is a known or plausible association with treatment and those for which there is none. Adverse drug reactions (ADRs) are those AEs that may reasonably be attributed to the medication (Box 2), distinguishing between medications that are used in accordance with their marketing approval (eg, in the approved dose, patient population and indication) or not.2,3 AEs are also classified as being serious or non-serious (Box 2 and Box 3). Standard schemes used to classify AEs, usually by body system, allow for easier comparison between different trial results and between different treatment options. Examples include the International classification of diseases,4 and the Medical dictionary for regulatory activities.5 Some classifications also grade severity of AEs, such as the Common terminology criteria for adverse events (CTCAE) system of the United States National Cancer Institute,6 which has five grades of severity, ranging from 1 (mild) to 5 (death). Regulatory requirements for reporting adverse eventsTo reliably report on AEs, procedures must be in place from the beginning of a trial for systematically recording and reporting them.2 All investigators participating in a clinical study, and their respective human research ethics committees (HRECs), must have been provided with an investigator’s brochure which includes all relevant information known about the safety, efficacy and pharmacodynamics of the investigational drug, and a description of the possible risks and adverse drug reactions associated with the drug and similar products.2 Good clinical practice guidelines require investigators to report immediately to the trial sponsor any serious AEs that occur during the conduct of a trial.2 In turn, many countries require the study sponsor to then report these to national regulatory authorities. An AE which is considered to be a serious unexpected adverse drug reaction7 must be notified by the sponsor to relevant regulatory authorities as an “expedited” SAE within 7 days of awareness for events that were fatal or life-threatening, and within 15 days for others.3 ICH guidelines for good clinical practice, as adopted internationally, also specify that all serious unexpected adverse drug reactions should be reported to the relevant HRECs.7 It is not possible to assess the significance of AEs from reports in which the treatment allocation remains blinded and the number of participants exposed to the trial medications is unknown. Hence, Data and Safety Monitoring Boards (with the ability to review events and their frequencies unblinded, if preferred) are an important (although insufficiently used8) mechanism for protecting the safety of trial participants9 and for ensuring that studies are stopped as soon as it becomes clear that the trial intervention is beneficial or harmful.10,11 Australian requirements for adverse event reportingIn Australia, the Therapeutic Goods Administration (TGA) only requires reports on serious unexpected adverse drug reactions that occur in Australia, and that it be informed of any significant safety concerns that arise from the sponsor’s monitoring of overseas safety reports and of any action undertaken by overseas regulatory agencies.3 In Australia, the section of the ICH good clinical practice guidelines on reporting to HRECs has been overridden by the National statement on ethical conduct in research involving humans.12 This mandates that investigators inform the TGA and HRECs of “all serious or unexpected AEs that occur during the trial and may affect the conduct of the trial or the safety of the participants or their willingness to continue participation in the trial”. One consequence of this directive is that HRECs in Australia are being inundated with large numbers of essentially uninformative AE reports.8 Presentation of adverse event reportsPatient selection can influence the rates of AEs. It is important that the trial population is described adequately so that clinicians can assess the risk of a particular treatment for an individual patient. It is also important that the mechanisms used to elicit reporting of AEs are documented. Volunteered reports of AEs can give incidences of AE markedly different from those ascertained through checklists or diaries.13,14 Counts of adverse eventsThe numbers of patients who had each type of AE should be clearly detailed in the study report. If some patients experience more than one type of AE, the numbers of each event type should also be documented.1 Both the number of patients experiencing at least one occurrence of the event of interest (for statistical analyses) and the total number of such events observed (for cost–benefit analysis) help interpretation (Box 4). Types of adverse eventsThe types of events chosen for reporting must be prespecified and may be selected on the basis of absolute numbers of events (ie, the most common events), biological relevance to the drug or study question, clinical relevance, or safety (ie, serious or severe events are reported). Adverse events by treatmentPresentation and comparisons of AEs are generally reported by allocated treatment (ie, the intention-to-treat [ITT] principle), but reporting by “treatment actually received” can also be useful in some settings. For example, where non-compliance rates with allocated treatment are substantial, ITT analyses will under-report treatment-related AEs. However, analyses by “treatment actually received” will provide only non-randomised comparisons and therefore contain a varying degree of selection bias.16 Therefore, ITT methods should be routinely reported, and data for treatment actually received should be added, with an explanation as to why, if there are high rates of non-compliance. Treatment withdrawal after adverse eventsAdverse events resulting in withdrawals from treatment should also be adequately described, as they reflect tolerability of treatment, and will be useful for both patients and clinicians to better assess the importance of particular reported AEs. Abstracts and keywordsFinally, when the study is published, the abstract and keywords should mention the term “adverse events”, even if none occurred in the study, to facilitate retrieval of AE data from databases such as MEDLINE.17 Current deficiencies in trial reportsA statement in the results section of a publication that “no adverse events were observed on the trial medication” without details in the methods section of the steps taken to ascertain AEs is difficult to interpret. What is less apparent is how difficult it really is to convey useful information on the types, severity and incidences of the AEs observed. Even large, multicentre studies published in first-rate journals can fail to report the number of patients who withdrew because of side effects of the trial medication. For example, a large study on the efficacy of irbesartan in preventing the development of nephropathy in patients with type 2 diabetes and microalbuminuria has the following description of the AEs observed: “Serious adverse events during treatment and up to two weeks after treatment were recorded in 22.8 percent of the patients in the placebo group and 15.4 percent of those in the combined irbesartan groups (P = 0.02). Nonfatal cardiovascular events were slighty more frequent in the placebo group (8.7 percent, vs. 4.5 percent in the 300 mg group; P = 0.11). The study medication was permanently discontinued in 18.9 percent of the patients in the placebo group, as compared with 14.9 percent of those in the combined irbesartan groups (P = 0.21).”18 In the report, the AEs that led to discontinuation were not described. Furthermore, it was unclear whether they were related to the condition being treated, to the trial medication, or to intercurrent illnesses. Guidelines have been developed to help researchers give useful information about AEs for reporting, both in general1 and for particular classes of trial, such as chemotherapy19 or postoperative analgesia.13 The need for such guidelines is evident from studies of adequacy of AE reporting. In one evaluation of reporting of safety data from clinical trials of HIV treatment, the severity of AEs was regarded as adequately defined in only a third of trials.20 In a subsequent study in other clinical areas, only 39% of trials adequately reported clinical adverse effects and only 29% adequately reported laboratory-determined toxicity.21 Other studies show similar rates of deficiencies in AE reporting in a variety of circumstances.22-25 These deficiencies are serious, as they can prevent clinicians from being able to provide patients with balanced information about the scale and scope of risks associated with different treatment strategies. Even when AEs have been well described, retrieving data about them can be problematic — of a sample of 37 trials indexed on MEDLINE or EMBASE known to present AE data, only 49% could be found by searching for text words such as “adverse event”, “side effect” or “h(a)emorrhage”, while adding indexing terms relevant to AEs only improved retrieval to 78%.17 ConclusionLarge randomised controlled trials and meta-analyses of randomised controlled trials are very effective in distinguishing AEs that are caused by the underlying condition from those that are related to the intervention. They can provide clinicians with unbiased information about the frequency and severity of adverse effects at a time when a drug or procedure is new. While other mechanisms for obtaining safety data are needed to detect AEs that are too rare to be detected by even the largest studies (Box 5), or that occur in groups of patients who would normally be excluded from trials (eg, because of illness severity, comorbid conditions or the need for potentially confounding therapies), these do not allow clinicians to quantify the risk of a treatment26 and can give markedly different impressions of the incidence of adverse drug reactions compared with those obtained from randomised clinical trials.27 Only proper reporting of AEs from randomised controlled trials allows adequate assessment of the potential net clinical benefits of interventions. 1 CONSORT checklist of items to report when reporting a randomised trial1 Section and topic Item no. Descriptor Results Adverse events 19 All important adverse events or side effects in each intervention group. 2 Definitions adopted by the International Conference on Harmonization, and adopted by Australia’s Therapeutic Goods Administration2 Adverse event An adverse event is any untoward medical occurrence in a patient or clinical investigation subject administered a pharmaceutical product and which does not necessarily have a causal relationship with this treatment. An adverse event can therefore be any unfavourable and unintended sign (including an abnormal laboratory finding), symptom, or disease temporally associated with the use of a medicinal (investigational) product, whether or not related to the medical (investigational) product. Adverse drug reaction Before marketing approval: all noxious and unintended responses to a medicinal product related to any dose should be considered adverse drug reactions. The phrase “responses to a medicinal product” means that a causal relationship between a medicinal product and an adverse event is at least a reasonable possibility. After marketing approval: a response to a drug which is noxious and unintended, and which occurs at doses normally used in man for prophylaxis, diagnosis, or therapy of diseases or for modification of physiological function. Unexpected drug reaction An adverse drug reaction, the nature or severity of which is not consistent with the applicable product information (eg, investigators’ brochure for an unapproved investigational medication). Serious adverse event or reaction Any untoward medical occurrence that at any dose: Results in death Is life-threatening Requires inpatient hospitalisation or prolongation of existing hospitalisation Results in persistent or significant disability or incapacity Causes a congenital anomaly or birth defect 3 Severity and causality of adverse events (AEs) * Or more severe than previously known. NSAE = non-serious adverse event; SAE = serious adverse event; SADR = serious adverse drug reaction. 4 Checklist for presenting adverse event (AE) reports Describe methods used to ascertain AEs (eg, reported by physician or patient) Show events which are unexpected in the context of the treatment given Categorise the seriousness of events where relevant Report AEs by the number of patients affected and by the number of events Report AEs by intention-to-treat methods (may additionally be shown by treatment actually received) Highlight AEs whose severity causes withdrawal or modification of treatment Highlight substantial differences in the risk of an AE for different subgroups of participants Use time-to-event (Kaplan–Meier survival) methods to avoid inflated estimates, where discontinuation rates have been high in long-term trials15 Mention key AE findings in the abstract and keywords of the report 5 Numbers of patients that need to be exposed to a medication to ensure that an adverse drug reaction has a 95% probability of being observed at least once Frequency of adverse drug reaction Minimum no. of patients required* Very common (≥ 10%) 29 Common (1%– < 10%) 299 Uncommon (0.1%– < 1%) 2 994 Rare (0.01%–< 0.1%) 29 956 * Number is based on the lower boundary of each category of frequency.
Anthony C Keech FRACP, MClinEpi · Susan M Wonders BDS · Val J Gebski BA, MStat · David I Cook FAA, FRACP
How people with chronic illnesses view their care in general practice: a qualitative study
Objectives: To explore the perceptions of patients with chronic conditions about the nature and quality of their care in general practice.Design: Qualitative study using focus group methods conducted 1 June to 30 November 2002.Participants and setting: 76 consumers in 12 focus groups in New South Wales and South Australia.Main outcome measures: Recurring issues and themes on care received in general practice.Results: Three groups of priorities emerged. One centred on the quality of doctors, including technical competence, interpersonal skills, time for the patient in the consultation and continuity of care. A second concerned the role of patients and consumer organisations, with patients wanting (i) recognition of their knowledge about their condition and self-management, and (ii) for GPs to develop closer links with consumer organisations and inform patients about them. The third focused on the practice team and the importance of practice nurses and receptionists.Conclusion: GPs should consider the amount of time they spend with chronically ill patients, and their interpersonal skills and understanding of patients’ needs. They need to be better informed about the benefits of patient self-management and consumer organisations, and to incorporate them into their care. They also need to review how their practice nurses and receptionists can maximise the care of patients.
Fernando A Infante MPH · Judith G Proudfoot PhD, MA, BEd(Hons) · Gawaine Powell Davies MHP · Mark F Harris DRACOG, FRACGP, MD · Tanya K Bubner BSocSc · Chris H Holton GDPH, GDAcc, BA(Acc) · Justin J Beilby MD, MPH
Ageing and healthcare costs in Australia: a case of policy-based evidence?
There have been dire predictions that population ageing will result in skyrocketing health costs. However, numerous studies have shown that the effect of population ageing on health expenditure is likely to be small and manageable. Pessimism about population ageing is popular in policy debates because it fits with ideological positions that favour growth in the private sector and seek to contain health expenditure in the public sector. It might also distract attention from the need to evaluate the appropriateness and effectiveness of current patterns of care. Pessimistic scenarios have stifled debate and limited the number of policy options considered. Policy making in Australia would be improved if we took a more realistic view of the effect of population ageing on health expenditure.
Michael D Coory MB BS, PhD, FAFPHM
My pseudoscientific nightmare
Everyone knows medical research is about improving healthcare Recently I had some trouble with people one might call “pseudoscientists”. These individuals are often technically quite competent; they seem to know their craft and they produce seemingly good work. But there is something amiss. It took me a long time to find out what that might be. Now I think I have identified it — the pseudoscientist has entered the field of science for the wrong reason: to advance not medicine, but himself. This theme must have been on my mind the other night, when I had the most vivid dream. The dream took me to the “Annual Festival of Pseudoscience”, where a panel of the most distinguished pseudoscientists reached consensus on how to become a fellow pseudoscientist. Here are the 10 commandments of pseudoscience that they dictated to their audience. Ensure that passionate belief rather than reason is the force that drives you. In science, one tests (more accurately, falsifies) hypotheses. In pseudoscience you want to “prove” what you already “know”. Only the biased researcher can mislead the world effectively. Avoid scientific training. Pseudoscience needs enthusiastic amateurs who have picked up the rules of science while busy doing other things. The worst that could happen to pseudoscience is for properly trained career scientists to join its arena. Maintain your bias. Bias usually originates from interests that create financial, personal or emotional conflicts. Nurture those interests and never disclose these conflicts to anyone, particularly not when publishing. Use publicity to obtain funding. Research funds are becoming scarcer by the minute. Lack of funds can seriously delay your endeavours. If you find it difficult to compete, the time-tested approach is to make more noise than anyone else. Hire a PR firm, for instance. Once the daily papers regularly sing your praises, your pseudoscience will thrive. Do not lose sight of what you intend to prove. Some people say that good research can never be “negative” — even showing that therapy X is not effective would yield the positive result of enabling patients to choose something that does work. Make sure your goals are not obscured by such old-fashioned nonsense — your aim as a pseudoscientist is to assist your friends, the manufacturers or promoters of therapy X. Let your goals drive your data analysis. Even with safeguards in place, you might one day generate a result that does not fit your preconceived ideas (or those of your sponsors). Subanalyse and subanalyse until you have what you were fishing for — a significant result showing what you want. Suppress unwelcome results. If things should go disastrously wrong and even extensive data dredging does not yield the desired outcome, the professional pseudoscientist must resort to the last, desperate, but usually effective, measure. Make the unwelcome finding disappear — don’t ever publish anything that does not confirm your beliefs or that might upset your friends. Overinterpret. More often than not, you will create data that you and your sponsors like. Now you must ruthlessly overinterpret these findings to ensure that everyone knows about your work. Publish your results as often as you possibly can. Journal editors don’t like duplicate publications, so it would be foolish to tell them. Attack opposing scientists. There is always a danger that scientists will publish papers that upset pseudoscientists. In such cases, initiate a campaign of defamation against your opponent — this will decrease their credibility and increase yours, and all will be fine again. At this point, I woke up feeling sick and anxious. Where does the dream end and reality begin? Did I have a nightmare or a vision? And, horror of horrors, did I not recognise some of the faces of the panellists? But it must have been a nightmare! These 10 commandments are just a guide on “how not to conduct medical research”. Surely, medical research is not abused as a career springboard? This would result in chaos and lead us badly astray without a compass for orientation. Surely, any responsible researcher knows that medical research is about improving healthcare? If not, how could we continue with the progress medicine has made so far? Surely, medical research has not been taken over by pseudoscientists. Or has it?
Edzard Ernst MD, PhD, FRCP
What explains falling asthma mortality?
Elizabeth J Comino Senior Research Fellow, School of Public Health and Community Medicine, University of New South Wales, Liverpool Hospital, Liverpool, NSW. E. CominoATunsw.edu.au To the Editor: The Australian Bureau of Statistics recently released details of asthma mortality for 2002. These figures indicate that asthma mortality has continued to decline in 2002 and that deaths in young people aged 5–34 years are at their lowest level since the early 1950s (Box). In this age group, the number of deaths fell from 43 in 2001 to 33 in 2002 (a 23.3% drop), and for all ages the number of deaths fell from 422 in 2001 to 397 in 2002 (a 5.9% drop). This suggests that the various asthma awareness activities, spearheaded by the National Asthma Council and other interest groups, have been successful in raising awareness of asthma and its management. Or does it? Emerging evidence suggests other changes in the epidemiology of asthma in Australia. Robertson recently reported a 26% decrease in the prevalence of asthma in Melbourne school children between 1993 and 2002, but increased reporting of rhinitis and eczema over the same period.1 The significance of these findings is difficult to interpret without measures of airway function. A second study supported these results and also observed a small decline in the prevalence of parent-reported asthma, but found little change in atopy or airway hyperresponsiveness.2 Age-adjusted hospital separation rates for asthma decreased by 31.1% in young people and 31.4% in all ages between 1989–90 and 1999–00.3 There is little evidence of improved management of asthma in the general practice setting. Data published from the BEACH (Bettering the Evaluation and Care of Health) survey of general practice activity indicates a significant reduction in rates of presentation for asthma among children but not adults, with no changes in indicators of severity over time.4 Our recent research in south-western Sydney, examining uptake of the “asthma 3+ visit plan”5,6 by general practitioners and their patients, is disappointing. It suggests reluctance on the part of both GPs and patients to participate in the plan. Clearly, there remains much that we do not understand about the natural history of asthma.7 We need to continue to monitor asthma through regular surveys and routine data collection in order to understand more about fluctuations in asthma prevalence, the relationship to changing child-rearing practices (such as use of childcare facilities) and the impact of management practices. Asthma mortality in Australians aged 5–34 years, 1920–2002* * Points on the graph represent 3-year “moving” averages — for example, the 2001 value is the average of 2000, 2001 and 2002 data; the 2000 value is the average of 1999, 2000 and 2001 data, etc. This technique is used to smooth annual fluctuations that occur in data of this kind.
Elizabeth J Comino
Using AUDIT to classify patients into Australian Alcohol Guideline categories
Julia E Fawcett,* Anthony P Shakeshaft,† Mark F Harris,‡ Alex Wodak,§ Richard P Mattick,¶ Robyn L Richmond** * PhD Candidate, † NHMRC Research Fellow, ¶ Director, National Drug and Alcohol Research Centre, University of New South Wales, Sydney, NSW 2052; ‡,** Professors, School of Public Health and Community Medicine, University of New South Wales, Sydney, NSW; § Director, Alcohol and Drug Service, St Vincent's Hospital, Sydney, NSW. A. ShakeshaftATunsw.edu.au To the Editor: Revised Australian Alcohol Guidelines1 were released in 2001. Although general practitioners (GPs) can be influential in initiating and supporting behaviour change to reduce levels of alcohol misuse among their patients,2,3 the extent to which their advice remains relevant and effective depends largely on the extent to which screening tools can be modified to take account of revised versions of such guidelines. The Alcohol Use Disorders Identification Test (AUDIT)4 is a clinical instrument used widely to screen patients for problematic alcohol use. The aims of our study were to examine the ability of AUDIT to classify general practice patients’ alcohol consumption into the categories specified in the revised Australian guidelines, and to identify any additional information needed for such classification. Patients aged at least 16 years attending a general practice surgery in western Sydney were asked by receptionists to complete a health-related survey by means of a hand-held computer while waiting for their consultation. Items covered a number of domains, including demographics, the AUDIT, and two additional questions about consumption of specified quantities of alcohol. The use of computers ensured patients were only asked questions relevant to them. Risk of harm in the long term: Respondents’ average number of standard drinks per week was calculated from the first two AUDIT questions, using a previously devised method.5 Risk of harm in the short term: AUDIT question 3 is not specific enough to distinguish short-term risk of harm, so additional, sex-specific questions on how many occasions in the previous 30 days the patient had consumed “7–10” and “11 or more” (men) or “5–6” and “7 or more” (women) standard drinks were asked. Of the 115 patients who completed the survey, 62% were female; their mean age was 42 years; 10% were unemployed; 34% had had tertiary education; 65% were married or in a de facto relationship; and 80% were born in Australia. Their alcohol consumption patterns are shown in the Box. AUDIT is a reliable and valid instrument, and is widely used as a clinical tool. However, as national guidelines are updated, clinical tools such as AUDIT need to remain consistent with them. Ideally, revisions would build on the benefits of existing tools rather than rendering them obsolete. For example, a major advantage of AUDIT is that it measures a number of drinking dimensions within the one, brief, validated instrument. This multidimensionality could be preserved while promoting AUDIT’s consistency with new guidelines by adding two items, with high face validity, to more accurately assess risk of harm in the short term. Incorp-orating the two additional consumption items we used in this study with AUDIT allows drinkers to be classified according to the guidelines as “low-risk”, “risky” or “high-risk” both in the long term and short term, with minimal additional response time. Alcohol consumption patterns in one general practice in western Sydney, as defined by the recently revised Australian Alcohol Guidelines1 Characteristic Males (%) Females (%) Total (%) Abstinent 18.2 29.6 25.2 Long-term harm Low-risk 68.2 67.6 67.8 Risky 11.4 1.4 5.2 High-risk 2.3 1.4 1.7 Short-term harm Low-risk 61.4 52.1 55.7 Risky 9.1 8.5 8.7 High-risk 11.4 9.9 10.4 Bold text represents categories that cannot be distinguished using AUDIT alone.
Julia E Fawcett · Anthony P Shakeshaft · Mark F Harris · Alex Wodak · Richard P Mattick · Robyn L Richmond
Estimating disease likelihood: a case of rubbery figures
In diagnosis and prognosis, we should avoid intuitive “guesstimates” and seek a validated numerical aid One of the axioms of clinical practice is that, in medicine, there are few, if any, certainties. When assessing the likelihood of a specific disease in a particular patient, or the chance of a future adverse event in a patient with known disease, clinicians are estimating probabilities or risk. These estimates derive from a clinical gestalt — the process of interpreting findings from history, examination and simple investigations (diagnosis), or of disease-specific correlates of complications or death (prognosis). Clinicians use these estimates of probability or risk to decide whether they should intervene immediately, particularly if effective treatments are available. Alternatively, if the disease likelihood is low, or treatments toxic or only marginally effective, these estimates are used to decide whether to defer treatment and either observe expectantly or conduct more sophisticated tests whose results may substantially alter pre-test likelihood estimates. If the estimate is too high, patients may incur unnecessary treatments or confirmatory investigations, or, if the estimate is too low, they may suffer the consequences of delayed intervention. Thus, a fair bit is riding on how accurately we can judge the likelihood of current or future disease. Available research suggests that, for various reasons, we are not that good at it.2-4 Common pitfalls include: framing a clinical problem in a way that may exaggerate risk; overweighting or underweighting certain clinical features; erroneously extrapolating past, vividly recalled cases to current patients; or manipulating risk subliminally to better fit with a preferred course of action (or inaction). Overall, most of us, not surprisingly, are risk averse and will commit to action to avoid personal regret at witnessing an unfavourable but possibly preventable event, even if our perception of risk of such an occurrence seems low.5 In this issue of the Journal, Attia and colleagues (page 449) evaluate the extent to which clinicians’ estimates of probability or risk for commonly encountered case scenarios vary from the “correct” estimate, and which clinician-related factors may influence such variation.6 They distributed three hypothetical case scenarios to groups of general practitioners and physicians in Australia and the United Kingdom, and compared respondents’ estimated probabilities of angina (in a patient with chest pain), deep vein thrombosis (DVT) (in a patient with a swollen leg), and future stroke (in a patient with chronic atrial fibrillation) with the “correct” estimates derived from statistically validated clinical-decision rules. Two cautions come to mind: were the clinicians given sufficient information on which to base a reasoned judgement (keeping in mind that they could not examine the patients); and how accurate was the rule-based estimate as the reference standard? One could argue that, in the chest-pain scenario, few experienced clinicians would be comfortable estimating the likelihood of angina simply on being told of a 65-year-old man presenting with exertional chest pain, without more detail about the character of the pain, the existence of coronary risk factors, and any signs of vascular disease seen on physical examination. The decision rule applied to the same case is also suspect, as it includes, for example, rapid relief with nitrogylcerine as being positively predictive of angina, which recent evidence would challenge.7 In the other two scenarios, the clinical details provided were more complete, and the decision rules more robust. Another concern is that the “correct” estimate was stated as a single percentage, which clinicians were expected, perhaps unfairly, to closely approximate. This ignored the fact that, in developing the rule, the “correct” estimate is actually a mean within a range of observed frequencies, all of which would probably lead to the same clinical action. On the positive side, the strengths of the study were its large, representative samples of clinicians, use of three different scenarios, and use of logistic regression to identify clinician-specific predictors of accuracy. Setting aside methodological limitations, how did the respondents fare? Only slightly more than half of the whole group were within 20 percentage points of the “correct” probability estimate for the angina and stroke scenarios, and less than one in 10 achieved a similar result with the DVT scenario. In keeping with my earlier comments, most respondents overestimated rather than underestimated the risk, with estimates spread over a huge range, from 10% to 100% at least, for all cases. There was a notable lack of association between accuracy and experience as measured by age, years of practice, or field of specialty, with GPs performing as well as physicians. Unfortunately, the study by Attia et al did not have the power to determine whether graduating from a medical course that used problem-based learning — with emphasis on evidence appraisal — predisposed to better performance. The implications of this study and others are several. First, all clinicians, irrespective of experience, appear to have problems quantifying probability or risk of disease, and, while there may be exceptions, this difficulty is independent of the clinical circumstances. Consequently, we should avoid intuitive “guesstimates” and seek instead a validated decision-rule, scoring scheme or other numerical aid that gets us closer to the mark. Fortunately, an increasing number of such tools are becoming available8 and in a form compatible with hand-held computers. Second, if we are to choose the best rules and use them appropriately, we need to understand how such rules should be constructed and tested.9 Third, we may need to “unlearn” some of our cherished clinical “rules of thumb” if evidence arises that questions their validity.10 Finally, we should advocate for more research into decision aids that will help us to more accurately estimate and communicate likelihood of disease in individual patients. The results of such efforts should facilitate a more rational use of investigations and treatments and lead to better patient outcomes.
Ian A Scott FRACP, MHA, MEd
Generating pre-test probabilities: a neglected area in clinical decision making
Objective: To assess the accuracy and variability of clinicians’ estimates of pre-test probability for three common clinical scenarios.Design: Postal questionnaire survey conducted between April and October 2001 eliciting pre-test probability estimates from scenarios for risk of ischaemic heart disease (IHD), deep vein thrombosis (DVT), and stroke.Participants and setting: Physicians and general practitioners randomly drawn from College membership lists for New South Wales and north-west England.Main outcome measures: Agreement with the “correct” estimate (being within 10, 20, 30, or > 30 percentage points of the “correct” estimate derived from validated clinical-decision rules); variability in estimates (median and interquartile ranges of estimates); and association of demographic, practice, or educational factors with accuracy (using linear regression analysis).Results: 819 doctors participated: 310 GPs and 288 physicians in Australia, and 106 GPs and 115 physicians in the UK. Accuracy varied from about 55% of respondents being within 20% of the “correct” risk estimate for the IHD and stroke scenarios to 6.7% for the DVT scenario. Although median estimates varied between the UK and Australian participants, both were similar in accuracy and showed a similarly wide spread of estimates. No demographic, practice, or educational variables substantially predicted accuracy.Conclusions: Experienced clinicians, in response to the same clinical scenarios, gave a wide range of estimates for pre-test probability. The development and dissemination of clinical decision rules is needed to support decision making by practising clinicians.
John R Attia MD, PhD, FRCPC · David W Sibbritt PhD · Ben D Ewald BMed, MMedSci · Balakrishnan R Nair FRCP, FRACP · Neil S Paget MA, DipEd · Rod F Wellard MEd, PhD · Lesley Patterson · Richard F Heller MD, FRCP
Subgroup analysis: application to individual patient decisions
Clinical trials provide evidence of effectiveness of treatments as an average for a group of patients, yet, in clinical medicine, we usually wish to apply these results to individuals. Can we simply apply the overall trial result for each patient, or can the result be tailored to individual patients in some way? Consider a hypothetical example: a randomised trial comparing treatments A and B shows that treatment A is more effective than B among men (P < 0.001), but not among women (not signficant). Does this mean men should receive the new treatment, but women should not? Box 1 illustrates these results from three different studies. In Study 1, the estimated treatment effect in men and women is the same — a 25% reduction in mortality associated with treatment A — but the much smaller number of women in the study gives rise to wider confidence intervals for this subgroup. In this case, there is no basis to consider that the treatment is any less effective in women (no heterogeneity; ie, non-significant test for interaction). Treatment could be considered effective for any patient regardless of sex. In Study 2, the treatment effect in women is less than that in men (8% v 25% relative reduction, or 0.92 v 0.75 relative risk, respectively), but the effects in both are still consistent with the overall result of a 20% relative reduction, and there is no evidence of significant heterogeneity between groups (test for interaction, P = 0.20). Here, the different results between men and women could be simply due to chance, and it would still be appropriate to apply the overall estimate to both men and women (unless there was additional evidence).1 In Study 3, the observed effects for men and women are sufficiently different to suggest that this difference is unlikely to be due to chance (test for interaction, P = 0.01), and it is reasonable to conclude that the treatment effect differs between men and women. In this instance, the trial evidence should be considered separately for these subgroups. However, even here, a test of interaction can still give a low P value simply on the play of chance if many subgroups have been evaluated.2 A practical approachConsider applying the overall trial treatment effect to each subgroupHow should we decide in practice whether to consider the treatment effects for these subgroups separately? A practical approach is shown in Box 2. A controlled trial is usually designed with a sample size large enough to show an overall treatment effect, but not necessarily adequate to show significant effects in each subgroup separately. The overall treatment effect is considered the best estimate for each subgroup of patients in the trial (this is sometimes referred to as the effect domination principle1). Hence, the treatment-effect results should only be applied differently for different subgroups if there is evidence of heterogeneity (a significant difference between subgroups, sometimes called interaction or treatment-effect modification).3-7 In considering whether there is evidence of heterogeneity, it is also worth reviewing the other questions outlined in the checklist for subgroup analyses in the previous article in this series.1 If there is clear and reliable evidence of heterogeneity, using the treatment effects for each subgroup may be appropriate. However, as these may be unreliable (based on smaller numbers of patients) they may still be considered exploratory and motivate further trials rather than necessarily leading to different treatment guidelines. Further, evidence of heterogeneity may also lead to a search for underlying factors which may be linked to the particular subgroup and provide a more plausible biological explanation for such variation. For example, an apparent difference in treatment effect between men and women may truly relate to differences between these groups in age or smoking status (so-called confounding). Seek confirmatory evidenceIf there is still uncertainty whether differences between the subgroups in treatment effect are real, the following steps should be taken: seek confirmation from the results of an independent trial, a meta-analysis, or both; determine whether the effect is also present for a composite (expanded) endpoint, or surrogate endpoints; and establish whether independent evidence exists of a-priori biological plausibility of differences in the treatment effect. Without a good a-priori rationale for subgroup differences, the overall treatment effect provides a reasonable estimate for each subgroup, unless confirmatory evidence of treatment differences becomes available. Estimate treatment effect according to baseline riskIf the same (or similar) relative treatment effect applies to different subgroups of patients, then those with a greater baseline risk of an event will derive a larger treatment effect. The absolute risk reduction associated with treatment is simply the absolute baseline risk multiplied by the relative risk reduction (Box 2).8 For example, for a patient group with a 20% baseline risk, a treatment with a relative risk reduction of 25% would translate into a 20 × 0.25 = 5% reduction in absolute risk. For a patient group with a 10% baseline risk, this would translate into a 10 × 0.25 = 2.5% absolute risk reduction. The number of patients needed to treat (NNT) to avoid one event can be calculated as one divided by the absolute risk reduction (Box 2).8 This corresponds to 1 ÷ 0.05 = 20 NNT for a patient with a baseline risk of 20%, and 1 ÷ 0.025 = 40 NNT for a patient with a baseline risk of 10%. Smaller numbers needed to treat will result for patients at higher baseline risk. A practical exampleA 75-year-old woman who has had a previous myocardial infarction (MI) presents within 4 hours of symptom onset with suspected acute MI and ST elevation on electrocardiogram (ECG); she is being considered for aspirin and reperfusion therapy. Data from randomised trials of aspirin, thrombolytic therapy and immediate coronary angioplasty are considered. The patient has no known contraindications to these treatments. Based on evidence from the ISIS-2 trial,9 there is strong evidence that aspirin reduces the risk of short-term mortality, by about 23%, both overall and within most subgroups examined, including women and patients aged over 70 years. However, in this trial, among patients with a prior MI, no significant treatment benefit was observed. While the treatment effect in this subgroup apparently differed from patients without a prior MI, the interaction may have been a chance finding owing to the many subgroups examined. (For example, the chance of at least one significant result at the 5% level among 20 independent tests is over 50%.1) Consequently, the trial evidence still strongly supports the use of aspirin therapy in this patient. Randomised trials of thrombolytic therapy in the FTT overview10 have also demonstrated clear evidence of a reduction in mortality from such treatment for patients with acute ST elevation presenting within 12 hours of symptom onset. This overview also suggested diminished effectiveness of such treatment in the elderly (test for trend with older age, P = 0.01). However, in this case, much of the heterogeneity could be explained by the fact that older patients more often presented later (after 12 hours) and without the specific diagnosis of ST elevation on ECG. Lack of ST elevation and late presentation to hospital relate directly to the underlying biology and are linked to diminished effects of treatment. Once these confounding factors have been taken into account, there is much less rationale for considering different treatment or withholding thrombolytic therapy simply on the basis of the age of our patient.11 Next, the role of immediate coronary angioplasty in such a patient could be considered. Randomised trials, particularly in specialised centres, have suggested an additional treatment benefit for immediate coronary intervention compared with thrombolysis. An individual patient data overview of earlier randomised trials suggests a relative reduction in death or reinfarction of about 50%, with similar relative effects in each of the subgroups examined.12 However, the absolute benefits of treatment (absolute risk reductions) were estimated to be much greater in the patients at high baseline risk, particularly those aged over 70 years (see Box 3). Consequently, if treatment with immediate angioplasty is considered appropriate in the particular hospital setting, it would be likely to have greater absolute benefit for this older patient than the average patient. Finally, the role of long-term treatment in this patient could be considered. Should statin therapy be considered on the basis of the evidence from such trials as the LIPID and CARE studies?13,14 Both of these had insufficient evidence to show reductions in mortality with treatment for women separately. In the LIPID trial, older and younger patients had similar relative reductions in events (Box 3), but older patients at higher baseline risk had greater absolute benefit. Fewer women than men were studied in these trials, and yet the results for women were not inconsistent with those for men (Box 3). The effect of statin therapy for women with prior CHD is illustrated further by the results of the 4S, CARE and LIPID trials.14 The combined results of these three trials show an overall significant reduction in coronary events; the estimates from the separate trials vary but are still consistent with the overall result. Evidence of a similar relative treatment effect from statin therapy in both women and men has recently been confirmed by the results of the Heart Protection Study.15 Finally, for some of these decisions, different recommendations for treatment may still apply even when a similar relative treatment effect seems valid and patients are at the same baseline risk. Circumstances in which different recommendations will be appropriate include: Where the importance of different outcomes (of benefit and harm) varies for different patients Where patient preference varies for other reasons Where there are limitations in applying the trial results in a particular setting, related to such factors as the skill or experience of practitioners and access to technologies. Principles for applying subgroup analysis to decisions about individual patients are summarised in Box 4. Cautious interpretation of the results of subgroup analyses is generally advisable. 1: Treatment effects in subgroups of men and women in three hypothetical trials Overall relative risk: 0.75 for Study 1; 0.80 for Study 2; 0.87 for Study 3; represented by the vertical dashed line in each case. 2: Interpreting treatment effects in different subgroups within a controlled clinical trial* * The decision pathway assumes there was a reasonable basis to consider the subgroups to have the same underlying condition. 3: Absolute risk reduction (ARR) and numbers needed to treat (NNT) in age and sex subgroups In none of these cases is there evidence of treatment effect modification (all P values for interaction are non-significant). ARRs and NNTs are derived from the overall relative treatment effect. 4: Principles for using subgroup evidence for making decisions about individual patients Use the subgroup-specific result only when there is (unconfounded) evidence of interaction and, ideally, confirmatory evidence. Use the estimated overall treatment effect if there is no evidence of heterogeneity (no interaction or treatment-effect modification). Adjust the size of the treatment benefit (and harm) according to the patient’s baseline risk. Consider patient preferences regarding each outcome when there are significant trade-offs in benefit and harm.
R John Simes MD, SM, FRACP · Val J Gebski BA, MStat · Anthony C Keech MB BS MScEpid FRACP
Parallel infusion of hydrocortisone ± chlorpheniramine bolus injection to prevent acute adverse reactions to antivenom for snakebites
Simon G A Brown Emergency Physician, Fremantle Hospital & Health Service, and Clinical Senior Lecturer in Emergency Medicine, University of Western Australia, Department of Emergency Medicine, Fremantle Hospital, Alma Street, Fremantle, WA 6160. simon.brownAThealth.wa.gov.au To the Editor: Gawarammana et al present an interesting study of premedication to prevent adverse reactions to snake antivenom,1 but I have concerns with the data presentation and analysis. The 95% confidence intervals presented in Box 3 of their article are impossibly narrow. For example, recalculation of the first cell (hypotension rate for treatment A) by the binomial method gives a 95% confidence interval of 12%–62%, not 30%–36% as presented. Also, my analysis of the aggregate endpoint differs. Using the Fisher exact test to compare reaction rates in the hydrocortisone + antihistamine group (11/21) with placebo (13/16), for the difference in proportions of 0.29 I obtain a 95% confidence interval of 2 0.04 to 0.58. This has a P value of 0.14 using a 2-tailed test, according to the program Analyse-it.2 This comes as no surprise given that steroids take hours to work, and that histamine is just one of many mediators released during anaphylaxis, rising early and only transiently during severe and protracted anaphylactic reactions.3 Although antihistamines reverse the effects of histamine infusions which mimic anaphylaxis, animal studies indicate that they are ineffective for treating anaphylaxis mediated by mast-cell degranulation.4,5 Human studies have shown that H1 blockade is useful in preventing mild reactions to immunotherapy that are confined to the skin, but do not appear to prevent severe reactions.6,7 One human study has compared H1 + H2 blockade with H1-only blockade for the management of mild allergic reactions, finding a small benefit of combined H1 + H2 blockade.8 However, a confounder was that adrenaline was administered more frequently in the combined H1 + H2 antihistamine treatment group. A study of reactions to immunotherapy has failed to show any benefit from H1 + H2 blockade.7 I suspect that, instead of proceeding to study combination prophylaxis with antihistamines, steroids and adrenaline, it might be better to examine the preparation for and management of allergic reactions to antivenom. Many doctors fear intravenous adrenaline infusions, but in our experience this approach to managing anaphylaxis is safe, well tolerated and immediately effective, without the unpredictable, unpleasant, and (in the setting of venom-induced coagulopathy) potentially dangerous side effects sometimes seen with intramuscular, subcutaneous and intravenous bolus injections.9 Perhaps it is time for a trial of premedication with subcutaneous adrenaline versus an “as required” approach using a carefully titrated intravenous infusion.
Simon G A Brown
Parallel infusion of hydrocortisone ± chlorpheniramine bolus injection to prevent acute adverse reactions to antivenom for snakebites
S Abeysingha M Kularatne,* Indika B Gawarammana† * Senior Lecturer, † Lecturer, Department of Medicine, Peradeniya University, Sri Lanka. samkulATsitnet.lk In reply: Brown has pointed out some interesting observations about antihistamines and mediators such as histamine, released during anaphylactic reactions. First, we must address the statistical issues. We are grateful to Brown for detecting an important error in Box 3 of our article.1 When we were asked to calculate confidence intervals for the percentages, we used a formula recommended by Spiegel.2 As this formula was not available in our software packages, we programmed it ourselves. We made an error in our programming and we apologise for this. (See Correction, page 428.) For the difference in reaction rates for the treatment regimens, we used the following calculation. The difference P1 – P2 = − 0.2887 To calculate the confidence interval, we calculated the standard error (SE) of the difference by the following formula. This interval does not include 0, so Brown’s P value of 0.14 is not reasonable. We do accept that the statistical significance is marginal. We agree that antihistamine is ineffective as a prophylactic agent in anaphylaxis. This was observed in studies in Sri Lanka and Brazil, where chlorpheniramine and promethazine, respectively, were used to prevent reactions to antivenom.3,4 However, we defended the observed reduction of mild to moderate acute reactions to antivenom in our study by highlighting the counter-effect of antihistamine on released histamine after mast-cell degranulation.1 Our study was to test the usefulness of an established practice in Sri Lanka, where steroid infusion is used to counter reactions to antivenom. The crux of the problem is the highly antigenic antivenom preparations used in Sri Lanka, despite ever-increasing reaction rates, because of the lack of facilities for developing purified antivenom. The proposal of using intravenous adrenaline infusion to reduce allergic reactions is novel and exciting. This should be tested by randomised controlled trials.
Indika B Gawarammana
Parallel infusion of hydrocortisone ± chlorpheniramine bolus injection to prevent acute adverse reactions to antivenom for snakebites
Val J Gebski Principal Research Fellow, NHMRC Clinical Trials Centre, Level 5, Building MO5, Mallett Street Campus, University of Sydney, NSW 2006. valATctc.usyd.edu.au Comment: Kularatne and Gawarammana compared the proportion of side-effects given in their table above using the well-known (and easily understood) test of the difference between two proportions using the normal distribution. However, as each of these proportions follows a binomial distribution (13 “successes” out of 16 trials and 11 “successes” out of 21 trials), the normal approximation is only useful if the number of trials is greater than 30. Because of the small sample sizes, “exact” methods better reflect the true difference (in terms of P values and confidence intervals) between the two proportions. The more complicated methods (different formulations of the exact test) and the resulting confidence intervals provide a clearer indication of whether the proportions are indeed different (statistically). In this instance, the evidence is not sufficient to declare the two proportions statistically different.
Val J Gebski
Subgroup analysis in clinical trials
Clinical trials represent a major investment by investigators, sponsors and participants, and it is reasonable to attempt to gain the maximum information from them. Practitioners and regulatory agencies are keen to know whether there are subgroups of trial participants who are more (or less) likely to be helped (or harmed) by the intervention under investigation, and a recent survey of trials published over 3 months in four leading journals found that 70% included subgroup analyses.1,2 Furthermore, regulatory guidance documents (such as the Committee for Proprietary Medicinal Products September 2002 document Points to consider on multiplicity issues in clinical trials3) strongly encourage appropriate subgroup analyses. The results of subgroup analyses can also drive changes in practice guidelines. For example, the United States National Institutes of Health issued a clinical alert following the unexpected finding in the BARI (Bypass Angioplasty Revascularisation Investigation) trial that mortality after angioplasty in patients with diabetes was nearly double that after bypass-graft surgery (P = 0.003).4 Meaningful information from subgroup analyses within a randomised trial is restricted by multiplicity of testing and low statistical power. There is therefore a tension between our wish to identify heterogeneity in the responses of trial participants to trial interventions and our technical capacity for doing so. Surveys on the adequacy of the reporting of clinical trials consistently find the reporting of subgroup analysis to be characterised by poor practice.2,5-7 Item 18 of the CONSORT checklist (Box 1) deals with the multiplicity issues that arise in subgroup analysis.8 Problems in subgroup analysisThe problem of multiple testingStatistical investigation of large numbers of subgroups inevitably shows significant interactions with the effectiveness of the trial intervention. By definition, testing at the 5% level of significance will erroneously report a statistically significant difference between subgroup categories in about 5% of the tests performed (so-called false-positive results). Trials with multiple comparisons to assess the comparability of randomised groups at baseline confirm this prediction.1,9 In subgroup analysis, where a plethora of factors (eg, sex, age, race, centre, smoking status, stage of disease, and coexistent disorders) may influence outcome, the risk of false-positive results is high.10 Overly enthusiastic analysis of subgroups can reveal statistically significant differences in outcome between subgroups even where neither arm of the study receives any intervention.11 In some cases, such as in the ISIS-2 study, which found a slight adverse impact of aspirin therapy on patients born under the star signs Gemini and Libra, and that aspirin helped after the first, but not subsequent, infarctions,12 the results of the subgroup analysis may be dismissed as contrary to current understanding of biological mechanisms. In other cases, such as the BARI trial,4 whether the finding was valid could only be established by additional studies.13,14 The problem of statistical powerMost studies enrol just enough participants to ensure that the primary hypothesis can be adequately tested. Therefore, statistical tests on subgroups will have only power to detect substantially larger effects on the same endpoint. Loss of compliance, together with adjustments for multiple testing, will exacerbate this lack of power.6 In consequence, when tested separately, many of the subgroups will fail to show the statistically significant treatment effect that was shown in the main population; at the same time, genuine differences in response to treatment (so-called heterogeneity) between study subpopulations may also go undetected. Can the problems be overcome?Despite subgroup analyses generally lacking statistical power, when used repeatedly to look for differences across many factors (eg, sex, age, smoking status, blood pressure) they have a proclivity to detect spurious effects. We are thus forced to reconcile our wish to find genuine differences between subgroups with the need to minimise the risk of accepting and publishing false positives.2,6 One solution to this dilemma is to accept that the results of subgroup analysis are hypotheses. Guidelines such as those given in Box 2 are intended to help readers identify which hypotheses are strong and which are weak. However, even among experts, opinions range from only accepting pre-specified subgroup analyses supported by a very strong a priori biological rationale15 to a more liberal view in which subgroup analyses, if properly carried out and carefully interpreted, are permitted to play a role in assisting doctors and their patients to choose between treatment options.16 Trial designAre the subgroups appropriately defined?Subgroups based on characteristics measured after randomisation, such as compliance, should be avoided, as allocation to the subgroup may be influenced by the intervention. Similarly, it is preferable to use the intention-to-treat population, as reasons for withdrawal may not be balanced between treatment arms. For example, adverse drug events may be a more important reason for withdrawals from an active treatment arm, whereas lack of efficacy may be more important in a placebo-controlled arm.17 Were the subgroup analyses planned before commencement of the study?In general, subgroup analyses should be defined a priori and purposely on the basis of known biological mechanisms or in response to findings in previous studies. Ideally, the choice of the subgroups and the expected direction of the subgroup difference should be justified in the trial protocol. Where a particular subgroup analysis is of great interest, adequate power to show the results can be designed into the trial, for example by using an expanded endpoint for the subgroup analysis. At the other extreme, subgroup analyses that are decided on once the dataset has been examined should be treated with scepticism. Intermediate between these two extremes are cases, such as occurred in the BARI trial, in which the subgroup analysis, although not originally planned, was decided on during the course of the trial in response to findings in other studies (with the investigators remaining blinded to the interim results of BARI).4 ReportingThe study report should include all the information required to assess the validity of subgroup analyses reported. In particular, the number of subgroup analyses should be declared, as this will enable readers to assess whether the issue of multiple testing is being dealt with. Analyses planned a priori, and the rationale for choosing them, should be clearly stated. Summary data, including event numbers and denominators for all the subgroup analyses, even the uninteresting ones, should be included, as this will facilitate future meta-analyses of the data and help prevent publication bias.18 The impact of multiple tests on the chance of declaring as statistically significant at least one false-positive result is shown in Box 3. Statistical analysisSome investigators avoid the issue of multiplicity of testing by tabulating the observed outcomes for the subgroups of interest without undertaking any formal statistical analysis. The data become available for meta-analysis,18 but there is the disadvantage that the investigator may fail to detect and draw attention to an important heterogeneity in the population. The statistical methods used should be appropriate for the hypothesis being tested. The common practice of performing subgroup-specific tests of treatment effect is flawed in that it is testing the wrong hypothesis.19 The hypothesis that should be tested is whether the treatment effect in a subgroup is significantly different from that in the overall population.19 Testing for a statistically significant treatment effect in a subgroup is hindered by a small sample size. The appropriate tests to use when analysing heterogeneity of responses among subgroups are interaction tests,2,10 for which worked examples are available.19,20 One study found that these were used in only 43% of 35 trials which reported subgroup analyses in their sample.2 Finally, the article should state whether the statistical tests used included adjustments for multiplicity. InterpretationBecause subgroup analyses have less power to detect a therapeutic effect than the main study, the trial report, especially in the Abstract or Conclusions, should emphasise the overall result. Given the risks of false-positive findings when multiple subgroup analyses are performed, it is not surprising if a subgroup-specific test shows a significant (P < 0.05) or suggestive (P = 0.05 to P = 0.10) effect of treatment, even when the trial failed to do so overall.2,7 Investigators are often tempted to highlight a particular subgroup analysis.2,7 For example, in one trial the suggestion that a psychosocial nursing intervention following myocardial infarction was harmful for women (P = 0.064), but not men (P = 0.94), was highlighted, even though the intervention did not affect survival in the overall population21 (and a test for interaction was not significant2). A number of arguments may be used to support the validity of a claimed subgroup effect (see, for example, the BARI trial4 and Rathore et al22): replication in another independent study; the presence of a dose–response relationship; reproducibilty of the observation in independent samples within the study, such as within individual sites; and the availability of a biological explanation. Of these, the first is the strongest evidence. For example, even though the BARI study found no difference in survival following bypass surgery or angioplasty in the overall population, the validity of the subgroup findings was supported by other studies.4 On the other hand, the report by Rathore et al that digoxin use is associated with a significantly increased risk of death among women (P < 0.014)22 is weakened by the fact that it was a post-hoc analysis which was motivated by “biological suspicion” rather than by suggestive findings in earlier trials. Biological justifications for the findings of a posteriori (exploratory) analyses, on the other hand, carry little weight6,23 — the reports that diabetes is more common in boys born in October,24 and that lung cancer is more common in people born in March,25 included (in)credible biological explanations after the findings had been revealed. The strategies for overcoming some of these difficulties in interpreting subgroup analyses will be explored in a forthcoming article in this series. 1: CONSORT checklist of items to include when reporting a trial8 Selection and topic Item no. Descriptor Ancillary analyses 18 Address multiplicity by reporting any other analyses performed, including subgroup analyses and adjusted analyses, indicating those pre-specified and those exploratory. 2: Checklist for subgroup analyses Design Are the subgroups based on pre-randomisation characteristics? What is the impact of patient misallocation on the subgroup analysis? Is the intention-to-treat population being used in the subgroup analysis? Were the subgroups planned a priori? Were they planned in response to existing trial or biological data? Was the expected direction of the subgroup effect stated a priori? Was the trial designed to have adequate power for the proposed subgroup analysis? Reporting Is the total number of subgroup analyses undertaken declared? Are relevant summary data, including event numbers and denominators, tabulated? Are analyses decided on a priori clearly distinguished from those decided on a posteriori? Statistical analysis Are the statistical tests appropriate for the underlying hypotheses? Are tests for heterogeneity (ie, interaction) statistically significant? Are there appropriate adjustments for multiple testing? Interpretation Is appropriate emphasis being placed on the primary outcome of the study? Is the validity of the findings of the subgroup analysis discussed in the light of current biological knowledge and the findings from similar trials? 3: Probability of at least one significant result at the 5% significance level given no true differences Number of tests Probability 1 0.05 2 0.10 3 0.14 5 0.23 10 0.40 20 0.64
David I Cook MD, FRACP · Val J Gebski BA, MStat · Anthony C Keech MScEpid, FRACP
Public funding of large-scale clinical trials in Australia
Alan Rodger Medical Director, and Professor of Radiology Oncology, Beatson Oncology Centre, Western Infirmary, Dumbarton Road, Glasgow, G11 6NT, UK. alan.rodgerATnorthglasgow.scot.nhs.uk To the Editor: I strongly support the editorial comments and recommendations of McNeil et al1 on public funding of clinical trials. Having worked in the Australian healthcare system for 11 years, and having returned to a changed National Health Service in Scotland a few months ago, I can vouch for the benefits that accrue from adequate funding for clinical trials. While McNeil and colleagues focus on large-scale trials, their comments apply equally to smaller trials. Certainly, in oncology, several trials organisations in Australia have struggled for years to continue conducting trials in spite of inadequate government funding. The ANZ Breast Cancer Trials Group and the Trans Tasman Radiation Oncology Group are but two organisations with which I am familiar. In addition to precarious funding, I believe the consequences in the past 2 years of upheaval in the insurance industry have placed all such groups on an uncertain and untenable footing. Clinical trials must be ethical, scientific and well managed. Clinicians entering patients into trials need support from essential data managers and clinical nurse specialists. Governments encourage and, in fact, demand evidence-based medicine. The only effective way to produce the evidence is to conduct clinical trials. That costs money. In Victoria, cancer trial management was supported by about $800 000 per annum in grants from the Cancer Council Victoria. Those funds provided start-up assistance to institutions new to clinical trials and supported the others that could not rely on pharmaceutical company largesse. However, only part of that money was state government funded and then only for rural and regional centres or for breast cancer. The diagnosis-related-group-based casemix funding of Victorian hospitals included a notional element for research. That sop was lost in budget deficits. In Scotland, where health matters are totally devolved to the Scottish Parliament, the latter’s Scottish Executive Health Department has very recently enhanced funding for cancer care. This includes £500 000 (A$1.25 million) annually to support clinical research in cancer. Each of the three cancer networks has a guaranteed share of that sum to resource clinical trials in all the associated health boards and cancer units. This is in addition to the excellent clinical trials units in the main cancer centres, often funded by the charity Cancer Research UK and industry. There is also a national system of considering and approving clinical research in cancer. Such approval places obligations on health boards to support such trials. Lastly, accreditation of cancer centres can depend on clinical trial participation. The evidence that this (still imperfect) system has an effect is seen in trial entry at our oncology centre, where 11% of patients are already entered into trials. The new funding should see that increase. Government needs to put its money where its mouth is: evidence needs resourcing. The alternative is to rely on charity or on industry (whose eye is more often on marketing than science).
Alan Rodger
Making sense of trial results: outcomes and estimation
The format in which the results of randomised controlled trials (RCTs) are presented can have a major impact on how they are interpreted, and the extent to which they will be adopted into clinical practice. A key element in the reporting of RCTs is the measurement scale on which outcomes are assessed. Scales which are presented as large whole numbers tend to attract the interest of clinicians and patients, independent of the reliability of the estimates.1,2 Enough information needs to be presented to allow clinicians to convert the size of the reported benefit into a format which allows easy comparison with other relevant trial results, including the range of certainty of the benefit (Box 1).3 As outcomes may be measured and collected in a variety of ways, it is essential that there is prior agreement on how any benefit or detriment of the intervention will be reported. Measurement of outcome efficacyFour critical measures of outcomes contribute to the interpretation of the benefits (or otherwise) of specific interventions: Observed effect of the intervention, which should reflect both the magnitude and direction of the effect, and be indicated as differences (between means or medians, proportions), ratios of quantities measuring association (odds, risk, hazards) or ratios of quantities measuring effect (variances or correlations). Other specialised measures (such as the “location” effect) also occur in certain statistical analyses such as the Wilcoxon rank sum test.4 Precision of the estimate of effect is usually called the standard error (SE), and provides a measure of how accurately this estimate measures the true intervention effect. In general the SE is proportional to the sample size — the larger the sample size the smaller the SE. The calculated SE of the effect takes into account the inherent variability (as measured by the standard deviation, SD) in the measured outcome in each of the comparison groups. Confidence interval for the true effect,5 which provides a range of feasible values within which the true effect may lie. The popularity of the 95% CI relates to the use of the 5% level of significance for testing whether the effect is likely to have occurred by chance alone. Confidence intervals are commonly two-sided, reflecting a two-sided hypothesis test (ie, compared with the control, the intervention can have either a benefit or detriment). A generic expression for a two-sided 95% CI is 95% CI = (effect – 1.96 × SE) to (effect + 1.96 × SE), where 1.96 is obtained from the standard normal distribution and relates to the 95% level chosen. P value, a statistical measure of the “strength” of the observed effects. Small P values suggest strong evidence of a real effect, while large ones suggest weak evidence. The conventional “cut point”, 0.05, sometimes referred to as the level of significance, is that value below which it is commonly deemed there is sufficient evidence to declare that the effect of the intervention is truly beneficial or detrimental. Thus a P value of 0.05 reflects a likelihood of 5% that the observed effect might have occurred by chance alone. However, statistical significance does not necessarily imply clinically meaningful differences, making it critical that the magnitude of the effect be accompanied by confidence intervals. Presentation of outcomesDifferent formats for presenting results can make comparisons with other studies confusing (see Box 2). Whichever format is chosen, sufficient information must be provided to allow readers to convert from one format to another. The most popular format is to present results as relative risk reductions (Box 2, Trial B). When the four trials in Box 2 were presented to clinicians, more than 70% considered the active treatments in Trials B and D worth using in clinical practice, while less than 20% considered the treatments in Trials A and C worthwhile. In fact, the “trials” were the same study and treatment.6 Even though the reliability of the statements in Box 2 cannot be determined without either a confidence interval or a P value, confidence intervals are rarely requested. Essential information to enable calculation of the results in all of these formats should be provided to readers. For studies of clinical events, this would include the numbers experiencing the event (numerators) in each group and the numbers at risk (denominators/group size) in each group (Box 3). From this, the proportion of participants in each group experiencing an event (risk) can be calculated, as well as the difference in proportions (absolute risk difference). A confidence interval around this difference and a P value can then be calculated. The relative risk reduction is then simply the risk difference divided by the risk in the control arm. The reciprocal of the absolute risk difference gives the number needed to treat (NNT),7 which is the expected number of patients who need to receive the intervention to see clinical benefit in one patient. A confidence interval for the NNT can be calculated by simply using the reciprocal of the confidence interval of the absolute risk difference. Interpretation of resultsGraphical presentation of results can give a clearer indication of effect sizes and clinical and statistical significance than presenting the results in a table, particularly when reporting treatment effect on multiple outcomes or in subgroups.8 Box 4 shows various scenarios demonstrating effect size and direction, confidence intervals (reliability) and the strength of the evidence (P value). The smallest useful clinical benefit underpins the interpretation of the treatment effect, and this is usually determined by expert clinical discussion before the study is undertaken. Cost and known side effects, as well as comparison with alternative treatment options, need to be considered when determining smallest useful clinical benefit. Box 4(a) shows a statistically significant benefit with a narrow confidence interval (confidence interval boundary does not cross the “no effect” [one] line) and a small P value. However, this effect is not large enough to achieve a clinically meaningful result and would not be considered important enough to change clinical practice, as the lower limit of plausible-effect magnitude falls above the smallest clinically useful benefit. Box 4(b) shows an effect which is both clinically and statistically significant (small P value). The magnitude surpasses the limit for a clinically useful benefit and, even though the confidence interval is wide, the minimum plausible effect is just beyond (more extreme than) the smallest useful benefit. Box 4(c) illustrates a null effect. This result is associated with a large P value and confidence intervals which cross the no-effect line. This is a reliably null result, with a zero-effect estimate, a narrow confidence interval and large P value. Box 4(d) is statistically non-significant and shows an inconclusive clinical effect. The true effect may well be beneficial, but the wide confidence interval is consistent with a broad range of possible effect sizes, suggesting the sample size for the study may have been too small (study lacks statistical power).9 Box 4(e) also shows an inconclusive result — both clinically and statistically. The very wide confidence interval indicates that the estimate of effect is unreliable. DiscussionSelecting the scale of measurement for the outcome is essential in study design, as this will be a critical factor in deciding the appropriate sample size.9 Some outcomes have a natural scale (eg, binary, ordinal), while others may lend themselves to reclassification. Thus, variables like blood pressure or cholesterol level may be easier to interpret when classified as high or low rather than being compared on their natural (continuous) measurement scale. The SE of the measured effect on outcome provides an estimate of the precision of the observed effect, while confidence intervals give a range of the plausible values in which the true effect may lie. Confidence intervals can also be used to aid in clinical decision making and to create clinical-significance curves and risk–benefit contours.5 Estimates of NNT provide a simple translation of the study results which can be directly applied to clinical practice. If these issues are considered, carefully planned and prospectively declared, the generalisability and validity of the final results will be enhanced. A checklist for good reporting of results is given in Box 5. 1: CONSORT checklist of items to include when reporting a trial3 Selection and topic Item no. Descriptor Outcomes and estimation 17 For each primary and secondary outcome, a summary of results for each group, and the estimated effect size and its precision (eg, 95% CI). 2: Challenges in comparing trial results Four treatments were tested against placebo in clinical trials for about 5 years. In no trial were there major side effects of the treatments. The results were reported as follows: Trial A 91.8% in the group allocated to the active treatment survived, compared with 88.5% in the placebo group. Trial B Patients allocated to the active treatment had a 30% reduction in the risk of death. Trial C Mortality was reduced by 3.4% in the group allocated to the active treatment. Trial D One death was avoided for every 30 patients treated. On the basis of these reports, and assuming all treatment costs are modest, which treatments would seem reasonable to introduce into your clinical practice? 3: The calculation of different effect measures in a placebo-controlled trial Treatment group (N = 1000) Control group (N = 1000) Number of events 60 100 Group risk (proportion with event) 60/1000 = 6% 100/1000 = 10% Absolute risk difference 10% - 6% = 4% Relative risk reduction 4%/10% = 40% Relative risk/risk ratio (treatment v control) 6%/10% = 0.6 95% CI for risk difference and P value 1.6%–6.4% reduction, 2P < 0.001 Number needed to treat and 95% CI 1/4% = 25; 1/0.064; 1/0.016 = 15.6–62.5 The “odds” of an event (and odds ratio) are commonly used instead of risk and risk ratio in reporting clinical trial results; odds are easily calculated as the number with an event divided by the number without an event. In this example, the odds of having an event are 6.4% for the treatment group and 11.1% for the control group, giving an odds ratio of 0.57 (95% CI, 0.41–0.80; 2P = 0.001). 4: Graphical representations of benefit from treatment 5: Checklist: ideal reporting of trial results Number of events expected in the control population, and the effect size assumed for the sample size calculation Numbers of events observed and numbers at risk in each comparator group separately The absolute risk reduction/difference for each event type Relative risk or odds ratio for treatment effect 95% confidence interval for either absolute risk reduction or relative risk (or odds ratio) 2-sided P value for determining statistical significance of either absolute risk reduction or relative risk (or odds ratio) Number needed to treat (NNT) and 95% CI and/or number needed to harm (NNH) and 95% CI The minimum clinically worthwhile benefit of the intervention
Rachel L O'Connell BMath, MMedStat · Val J Gebski BA, MStat · Anthony C Keech MScEpid, FRACP