Topics

Statistics

Ethics Letters 2 February 2004 Free

Multicentre research: negotiating the ethics approval obstacle course

Lynne M Roberts,* Lucy Bowyer,† Caroline S Homer,‡ Mark A Brown§ * Research Midwife [corresponding author], † Senior Lecturer in Obstetrics (University of New South Wales), ‡ Midwifery Consultant, Department of Women’s and Children’s Health, St George Hospital, Research Building, St George Hospital, Kensington Street, Kogarah, Sydney, NSW 2217; § Professor of Medicine (University of New South Wales), Department of Renal Medicine, St George Hospital. RobertslyATsesahs.nsw.gov.au To the Editor: The obstacles presented by Human Research Ethics Committees (HRECs) have caused a significant delay in commencing a valuable research project. We are currently conducting a multicentre study investigating the outcomes of hypertensive pregnancies in a cohort of 1620 women. It is a retrospective review of medical records and does not entail any participation of the women. Ethics approval was sought and gained from the New South Wales Health Department and one other NSW area health service (AHS) involved in the study. The bulk of the medical records (85%) are held by this AHS and a smaller proportion by eight other AHSs in NSW. Despite these prior approvals, the process of gaining ethics approval from the eight AHSs was fraught with obstacles at every stage. After 8 months’ work, we have received approval from the HREC of each of the AHSs. Our experience has revealed many inconsistencies in the requirements of the HRECs in the different AHSs, as summarised in the Box 1. These inconsistencies highlight discordances with the guidelines to support researchers and HRECs drawn up by the National Health and Medical Research Council (NHMRC). The NHMRC’s National Statement on Ethical Conduct in Research Involving Humans1 clearly outlines that, once approval has been gained from one HREC, other sites should accept that approval. Unfortunately, it seems that the HRECs involved in giving approval for our study did not follow the guidelines relating to multicentre projects. Other researchers have reported similar problems. 2-4 Breen and Hacker2 suggest that HRECs are slow to adopt a simplified review process because this interferes with traditional practices of each committee making its own assessment. It is indisputable that ethics considerations are a vital component when undertaking human research. It is also crucial to have a reliable and trustworthy process that evaluates research proposals in order to protect participants from physical and psychological harm. However, it has taken the research midwife (who is on a 1-year non-renewable grant) 8 months to secure ethical approval at all sites. This process is cumbersome and counterintuitive to the principles and guidelines for multicentre research in this country. 1: Summary, by area health service (AHS, coded S to Z), of different requirements for gaining ethics approval for a multicentre study Area health service S T U V W X Y Z No. of pages of application form 19 20 19 20 12 23 2 11 No. of copies of form required 1 1 17 15 20 16 1 14 No. of hospitals in AHS covered by approval 2 2 4 1 3 5 3 5 Approval covered private hospitals in AHS also na na Yes No No na No na No. of contacts made (phone/letter/email) to gain approval 20 15 20 15 20 30 10 20 Time taken to gain approval 3 months 5 months 4.5 months 8 weeks 6 weeks 6 weeks 1 week 3 weeks Special requests during approval process A F A, B, C, E A A, D A G, H Approved after first submission Yes No No Yes Yes Yes Yes Yes na = not applicable. A = Asked for local researcher to be a contact person for the study. B = Charged a $33 fee to submit application. C = Requested scientific protocol with references. D = Requested budget form. E = Reviewed by scientific advisory committee before human research ethics committee (HREC). F = Requested consent and subject information forms. G = University HREC’s approval as well as approval of area health service HREC required. H = Final approval required from chief executive officer of major hospital in that AHS.

Lynne M Roberts · Lucy Bowyer · Caroline S Homer · Mark A Brown

Statistics Letters 19 January 2004 Free

Overweight and obesity in Australia: an underestimate of the true prevalence?

Terry J Coyne,* Michael G Findlay,† Torukiri I Ibiebele,‡ David W Firman§ * Senior Lecturer, School of Population Health, University of Queensland, Public Health Building, Medical School, Herston Road, Herston, QLD 4029; † Acting Senior Analyst, ‡ Assistant Analyst, Epidemiology Services Unit, Queensland Health, Brisbane; § Team Leader, Surveys and Social Statistics, Office of Economic and Statistical Research, The Treasury, Queensland Government, Brisbane.t.coyneATsph.edu.au To the Editor: While the rates of overweight and obesity among Australian adults, as determined by the Australian Diabetes, Obesity and Lifestyle Study (AusDiab),1 may be alarming to some, they may in fact be underestimates of the true prevalence of overweight and obesity. The AusDiab study design2 and its low response rates indicate that the results will need to be interpreted with caution. Firstly, the AusDiab study design excluded rural and predominantly Indigenous census collection districts (CDs). In Queensland, all CDs selected were capital city or other major urban centres (Rural and Remote Areas Classification, categories 1 and 2);3 thus, people living in major rural centres (such as Rockhampton or Bundaberg) or major remote centres (such as Mt Isa) were excluded (ie, in Queensland, about 20% of the population were excluded). Secondly, another potential bias may have been introduced by the Socio-Economic Indexes for Areas (SEIFA) scores of the CDs sampled in the AusDiab study. For example, the overall SEIFA score for the CDs included in Queensland was 1035 (73rd percentile), well above the state average. Finally, the response rates in the AusDiab study were low: only 29% of those estimated to be eligible, and only 52% of those invited, actually completed the study. Our analysis of risk factors of Queensland-AusDiab participants suggests that these participants may have been more health conscious than the general Queensland population. Rates of smoking reported for men and women were considerably lower in the Qld-AusDiab cohort compared with those in the Queensland phase of the 2001 National Health Survey4 (17.3% and 14.5% v 28.4% and 19.8%, respectively). Compared with results of a Queensland Omnibus telephone survey5 conducted at about the same time, higher proportions of Qld-AusDiab participants reported greater intakes of vegetables (≥ 4 serves/day: 27.4% Qld-AusDiab v 16.4% Omnibus) and fruit (≥ 2 serves/day: 28.9% v 24.3%), and less frequent consumption of fast foods (> 1 day/week: 37.3% v 49.5%). Given the low response rate and possible selection bias in the AusDiab study, we suggest that the overweight and obesity data should be interpreted with caution. Several indicators suggest that these data could be underestimates of the true prevalence of overweight and obesity, and that the AusDiab population may have been of higher socioeconomic status, more health conscious (lower rates of smoking, better dietary intake), and more willing to participate in a lengthy examination than the general Australian population. These factors may all be associated with lower rates of overweight and obesity, and therefore future national surveys will need to take these factors into consideration to obtain more accurate estimates of important determinants of health.

Terry J Coyne · Michael G Findlay · Torukiri I Ibiebele · David W Firman

Statistics Letters 19 January 2004 Free

Overweight and obesity in Australia: an underestimate of the true prevalence?

Adrian J Cameron,* Paul Z Zimmet,† David W Dunstan,‡ Jonathan E Shaw§ * Epidemiologist, † Director, ‡ Research Fellow, § Physician in Diabetes, and Director, Clinical Research; Epidemiology Department, International Diabetes Institute, 250 Kooyong Road, Caulfield, VIC 3162. acameronATidi.org.au In reply: Coyne suggests that, based on comparisons within Queensland, the national prevalence of obesity in our article1 is an underestimate. It should be noted that the Australian Diabetes, Obesity and Lifestyle Study (AusDiab) was designed primarily to produce national, not state-specific, data. Forty-two census collection districts (CDs) were selected Australia-wide, with only six CDs selected within each state. The primary objective of this sample selection was to obtain a nationally representative population, not necessarily one representative of each state. Coyne states that none of the Queensland CDs were in major provincial centres. Of the six Queensland CDs, four were outside Brisbane. From the national sample, 17 of 42 CDs (40.5%) were outside capital cities. As a comparison, 36% of the Australian population lives outside capital cities.2 Regarding selection of CDs, we excluded only those in Statistical Local Areas defined as 100% rural, and those where the Indigenous population made up 10% or more of the overall population.3 This excluded only 5.8% of the total eligible population. If the prevalence of obesity among this group was double the overall prevalence, this would not significantly alter the national rate. While the smoking rates in AusDiab were lower than reported elsewhere, the prevalence of obesity, hypercholesterolaemia and hypertension were in line with trends in a series of surveys over the past 20 years.4 In an extensive analysis of food consumption between AusDiab and the 1995 National Nutrition Survey, the rates of fruit and vegetable consumption were within 4% between the surveys for those most commonly eaten. Since our conclusion was that obesity has increased, the possibility of an underestimate only reinforces our message.

Adrian J Cameron · Paul Z Zimmet · David W Dunstan · Jonathan E Shaw

Statistics Editorials 17 November 2003 Free

Public funding of large-scale clinical trials in Australia

Failure to provide public funding for clinical trials may come at a high cost to the community in the long term Large-scale morbidity–mortality trials have become fundamental to the evaluation of most new drugs intended for long-term administration. Such trials have the unique ability to determine the net balance of positive and negative outcomes from the long-term use of drugs and allow consumers to feel confident that long-term therapy is safe in otherwise healthy individuals. A randomised controlled trial is the only study design that can provide reliable and unbiased estimates of the moderate treatment effects of interventions for most chronic diseases. Smaller studies measure surrogate outcomes or are underpowered to answer important clinical questions. If underpowered, they may be unethical and also squander the community altruism that underpins trial participation. Large-scale trials are major logistical exercises. They involve several thousand people allocated randomly to different treatment groups and monitored for 4–6 years. Depending on the recruitment strategy and the mode of follow-up, the cost is typically $20–$50 million.1 Past attempts to interest Australian research organisations in funding such studies have floundered because of this cost. In spite of this, there are a number of groups in Australia with an excellent track record in initiating and running high-quality large-scale clinical trials (many of the “public good” variety), indicating local capacity to conduct this type of research. However, these trials have largely been the province of the pharmaceutical industry. Although industry-funded studies have yielded firm scientific foundations in many areas of clinical practice, they are almost all directed towards testing superiority or equivalence of specific products, or other aspects such as greater convenience of new agents or technologies over the existing ones. Indeed, it is naïve to believe that the interests of industry will align with broader societal interests in securing effective and affordable care.2 The failure of non-industry concerns, including governments, to fund large-scale clinical trials leaves some conspicuous gaps in evidence where the consequences may be forgone savings for the public purse. In other cases, the result may be prolonged community exposure to older agents whose long-term risks have not been adequately assessed. Some recent examples highlight the importance of this problem. Antihypertensive drugs make up a large component of the Australian pharmaceutical budget — $516 million for the Pharmaceutical Benefits Scheme (PBS) for the newer agents in the financial year 2002–03 alone.3 For many years, there has been a trend towards these newer, more expensive agents replacing older, cheaper drugs for first-line management of mild hypertension.4 The justification was provided by small trials involving surrogate endpoints, such as effects on blood pressure control, vascular changes, and other risk factors. However, clinical trials with surrogate endpoints do not provide an appropriate basis to underpin long-term drug therapy: they can not provide reassurance of the drug’s long-term safety or determine the balance of desirable and undesirable effects of new agents. When the necessary studies of antihypertensive drugs were finally undertaken, they demonstrated that the advantage of newer agents over diuretics was marginal, at best.5,6 In this instance, a lack of appropriate trial data on management of hypertension probably led to years of unnecessary expense to the PBS that greatly outweighed the cost of a large-scale trial. It is clearly in the public interest to ensure that PBS funds are not being spent on expensive therapies when much cheaper agents are just as effective. We recently estimated that the failure to provide funding for trials probably cost Australian taxpayers between $45 million and $108 million in 1998 alone.4 Another example of the false economy of failing to fund clinical trials is the recently reported Women’s Health Initiative study.7 Before the results of this trial were published, a generation of women was prescribed hormone replacement therapy (HRT), despite the lack of rigorous long-term safety data that could only have been obtained from a large-scale trial. There was little commercial imperative to fund such a long-term trial when large markets of regular users existed. Eventually the US National Institutes of Health (NIH) recognised the importance of funding such a study, as indeed they have funded a number of other “public good” studies. Release of the study results has led to a sharp drop in the use of combined HRT, except for short-term use to relieve significant perimenopausal symptoms. The “dividend” for the Australian government was $16 million less expenditure on HRT in the financial year 2002–03.3 A failure to learn from these experiences may cost the community in the future. For example, low-dose aspirin is an effective antiplatelet agent whose use has recently been advocated in the United States for people with a 10-year risk of coronary events and stroke of 10% or more.8 This recommendation may lead to widespread use of aspirin for primary cardiovascular prevention in the elderly, despite a lack of data to indicate that its benefits in this age group outweigh the risk of haemorrhage.9,10 There is little likelihood that commercial interests will supply the funding to overcome this lack of data. It is more likely that industry would fund a study using a newer, more expensive antithrombotic agent in the hope of establishing it as standard therapy. The means must be found to identify and target strategically important research questions that require public funding. A budget (in the order of $100 million) for national research funding of these large “public good” trials should be established and administered by the National Health and Medical Research Council (NHMRC). This sum, representing 12% of the NHMRC budget and 0.2% of the recurrent health expenditure of $60 billion, is commensurate with the importance of such trials to clinical medicine and public health. Using the NIH as a model, trials would involve a mix of requested and investigator-initiated research. Research groups, either alone or (more likely) collaboratively, would apply for competitive funding. Although this would be administered by the NHMRC, a number of other stakeholders would benefit, including federal and state governments and their agencies, departments of health, the Health Insurance Commission, and the PBS. New funds should be made available from these sources. States should contribute to this initiative as large-scale trials are usually multicentred, allowing research capacity building and employment in both metropolitan and rural areas throughout Australia. Failure to develop a policy that supports such strategic research may well lead to waste of public funds and a delayed recognition of unfavourable risk–benefit ratios.

John J McNeil PhD, FRACP · Mark R Nelson PhD, FRACGP, FAFPHM · Andrew M Tonkin MB BS, MD, FRACP

A comparison of buprenorphine treatment in clinic and primary care settings: a randomised trial

John R M Caplehorn Senior Lecturer, Clinical Epidemiology, School of Public Health, University of Sydney, Sydney, NSW 2006. johncAThealth.usyd.edu.au To the Editor: The trial of buprenorphine-assisted heroin detoxification in primary care and a specialist clinic by Gibson et al1 was intended to compare the effectiveness and cost-effectiveness of buprenorphine-assisted withdrawal in a specialist clinic with treatment by general practitioners. However, of the average $191 for primary care staff costs, $69 was incurred at the clinic. As at least a third of interactions between patients and staff actually took place in the clinic, the primary care arm of the trial was really a combination of specialist clinic and primary care. Another design problem was the study’s lack of statistical power. A study would need 550 participants to have an 80% chance of identifying (at the 0.05 level of statistical significance) a difference of 50% in self-reported abstinence during the 8-day detoxification (ie, improving the percentage reporting abstinence from 22% to 33%). The trial by Gibson and colleagues had only 115 participants. As expected, the trial produced statistically non-significant results. Yet, the authors highlight the finding that 23% of primary care patients reported being abstinent during the 8-day detoxification, compared with 22% of the clinic patients, (95% CI risk difference, –14.1% to 16.5%; P = 0.9 [χ2]). Moreover, the clinic group performed better on an objective and more reliable measure of abstinence: 20% of clinic patients versus 14% of primary care patients gave morphine-free 8th day urine specimens, (95% CI risk difference, –7.7% to 19.8%; P = 0.4 [χ2]). As the confidence intervals for these risk differences include zero, the confidence interval for any estimate of incremental cost-effectiveness includes infinity. It is quite misleading for Gibson and colleagues to claim that “it costs $20 to achieve a 1% improvement in outcome in primary care”, as this ignores both the conflict and the variability in their clinical outcomes.1 Moreover, the statement ignores the variability in the estimated costs of treatment (eg, mean cost per clinic patient, $332; SD, $70). Surprisingly, Gibson and colleagues did not collect any information on continuing abstinence at the 13-week follow-up. Rather, they collected information on patients’ current treatment. While patients in whom detoxification therapy fails should be offered other treatment, post-withdrawal engagement in maintenance treatment is not a meaningful measure of the effectiveness of detoxification. If anything, it is a measure of failure. The trial needed sufficient statistical power to identify clinically meaningful differences in abstinence at the end of the 8-day detoxification and at 13 weeks. Staff working in the specialist clinic should not have been extensively involved in the delivery of primary care. Gibson and colleagues should have summarised their findings using appropriate estimates of clinical effect and cost-effectiveness with 95% confidence intervals.2

John R M Caplehorn

A comparison of buprenorphine treatment in clinic and primary care settings: a randomised trial

Amy E Gibson Senior Research Officer, The National Drug and Alcohol Research Centre, University of New South Wales, Sydney, NSW 2052. amy.gibsonATunsw.edu.au In reply: The primary focus of our study1 was retention in treatment, and not differences in abstinence. Caplehorn has previously argued compellingly that an orientation to abstinence can have an adverse impact on treatment outcomes in opioid dependence.2 We were using buprenorphine to redefine detoxification, not as a treatment producing lasting abstinence but as a way of promoting engagement in ongoing treatment. The power of our study was calculated on the basis of the proportion of subjects entering post-detoxification treatment, not on their self-reported abstinence levels. During the detoxification stage in the primary care setting, we used a shared-care dosing arrangement. This was primarily because of the need to give an initial research assessment to all participants before they were randomly allocated to treatment arms — something that would only occur in the context of a research study, and noted in the discussion. Further details of the health economic analysis are soon to be published.3 Ours was a study of the setting for buprenorphine treatment. Its critical finding was that patients were equally as likely to be engaged in maintenance treatment with practitioners in primary care as in specialist clinics.

Amy E Gibson

Statistics EBM: Trials on trial 20 October 2003 Free

Inclusion of patients in clinical trial analysis: the intention-to-treat principle

Determining the sample of participants to be analysed is a crucial step in reporting clinical trials. For such analyses, the gold standard is the “intention-to-treat” principle. The question of which participants are included in the analysis appears as Item 16 of the CONSORT statement (Box 1).1 Intention-to-treat (ITT)Analysis by ITT is a strategy that compares the study groups in terms of the treatment to which they were randomly allocated, irrespective of the treatment they actually received or other trial outcomes. Regardless of protocol deviations and participant compliance or withdrawal, analysis is performed according to the assigned treatment group.2,3 Random allocation aims to ensure that trial participants’ risk factors that may affect the outcome under investigation are balanced between the allocated treatments. This is to ensure that any differences in outcomes observed between groups are actually a result of the trial interventions. Importantly, there can be no guarantee that participants from each group who do not comply with the allocated treatment have the same risk-factor profile. Any analysis other than an ITT analysis (eg, one that excludes non-compliant participants) will potentially compromise the balance of these factors and introduce bias into the treatment comparisons. Thus, the ITT strategy generally gives a conservative estimate of the treatment effect compared with what would be expected if there was full compliance. By accepting that non-compliance and protocol deviations are likely to occur in actual clinical practice,3,4 ITT essentially tests a treatment policy or strategy, and avoids overoptimistic estimates of the efficacy of an intervention resulting from the removal of non-compliers. Ensuring ITT produces meaningful answersThe reality of conducting clinical trials means that the ITT principle is not usually fully met, especially when outcome data are missing for some participants. However, clinical trial researchers should consider this principle an ideal, and steps to achieve it should be considered in both the design and conduct of a trial. Firstly, eligibility errors can be avoided by careful scrutiny before random allocation. Indeed, allocation of ineligible patients should be the exception, unless eligibility cannot be assessed quickly. Secondly, all efforts should be pursued to ensure minimal dropouts from treatment, crossover of participants between groups and losses to follow-up. An active run-in phase may be feasible to identify patients who are likely to drop out. A thorough consent process for participants and education of investigators will also minimise the number of dropouts. During the trial, adequate warning of the potential side effects of treatment, together with ongoing clinical support and reassurance, should be available to all participants. When a proportion of participants are expected to receive a treatment different from the assigned one, a dilution effect generally results. The subsequent potential loss of study power can be accounted for by increasing the planned sample size.5 Box 2 details the advantages and limitations of ITT analyses. Alternatives to ITT analysisPer-protocol (PP) analysisThere is a view that only patients who sufficiently complied with the trial’s protocol should be considered in the analysis.6 Compliance covers exposure to treatment, availability of measurements, and absence of major protocol violations. Such an analysis is often referred to as a “per-protocol” or “on treatment” analysis. The main issue arising from this approach is that it might introduce bias related to excluding participants from analysis. Therefore, the ITT analysis should always be considered as the ideal primary analysis, possibly supplemented by a secondary analysis using the PP approach. However, if investigators decide differently, their choice must be justified and should be subject to strict rules.7-9 Treatment-received (TR) analysisAnother approach is to analyse all participants according to the treatment they actually received, regardless of what treatment they were originally allocated. While this may have some initial appeal, once again the effect of random allocation is compromised, making the interpretation of the results difficult. The impact of various approaches is illustrated in Box 3. When ITT requirements are not fully metA number of strategies can be adopted if the assumptions underpinning ITT are not satisfied. If the crossover/non-compliance rates are small, then an ITT analysis should be the principal method of analysis. There is still some debate about whether ineligible subjects can legitimately be omitted from the final analysis.2 For instance, in a study involving a potentially life-threatening condition, such as severe acute respiratory syndrome, treatment may be routinely commenced before laboratory confirmation of the diagnosis. If the patients subsequently are not diagnosed with the condition, there may be a case for excluding them from the ITT population. In these instances, a “modified” or “quasi” ITT population may be defined, allowing for such exclusions. The following principles should be followed to allow participants to be excluded from such an analysis: the criteria for exclusion from the analysis should be pre-specified in the protocol, be objective and clearly defined;7,8 and, to remain unbiased, decisions to exclude participants need to be made (i) by researchers blinded to treatment allocation, and (ii) on the basis of information not related to either the allocated treatment or to events or outcomes that occur after random allocation. In all circumstances, all patients randomly allocated to a study arm should be followed up, as exposure to study treatment may still influence their safety and place them at risk of serious adverse events. All efforts must be made to ensure maximum compliance and that patients continue to take their allocated treatments, and that all patients are accounted for in the trial report.9 The modified or quasi ITT population may also be useful when outcomes are not assessed in all participants. For example, outcomes requiring colonoscopic follow-up can result in no information for patients who, for any reason, did not undergo colonoscopy during the study, requiring an analysis based on a subset of the patient population.10 In such a case, modifying the ITT population allows some clinical interpretation of the results. A more extreme example is a study evaluating hip protectors, in which only around 50% of those in the intervention arm were wearing a hip protector at the time of their fracture.11 In this situation, neither an ITT or per-protocol analysis would necessarily provide reliable information about the value of hip protectors when actually worn. There has been debate about the appropriateness of imputing missing values.4 If missing data are imputed, it is recommended that some sensitivity analysis be performed to ensure that study conclusions are not misleading.4,12 ConclusionITT analysis gives unbiased and consistent estimates of a treatment policy, and should, wherever possible, be the analysis of choice. Deviations from this principle compromise the balance between groups that is achieved by random allocation, and are rarely justifiable as a principal analysis. 1: CONSORT checklist of items to include when reporting a trial Selection and topic Item no. Descriptor Numbers analysed 16 Number of participants (denominator) in each group included in each analysis, and whether the analysis was by “intention to treat”. State results in absolute numbers (eg, 10/20, not 50%). 2: Advantages and limitations of an intention-to-treat (ITT) analysis Advantages Retains balance in prognostic factors arising from the original random treatment allocation Gives an unbiased estimate of treatment effect Admits non-compliance and protocol deviations, thus reflecting a real clinical situation Limitations Estimate of treatment effect is generally conservative because of dilution due to non-compliance In equivalence trials (attempting to prove that two treatments do not differ by more than a certain amount), this analysis will favour equality of treatments Interpretation becomes difficult if a large proportion of participants cross over to opposite treatment arms Requirements for an ideal ITT analysis Full compliance with randomised treatment No missing responses Follow-up on all participants ITT analysis is highly desirable unless: there is overwhelming justification for a different analysis policy (eg, an unacceptably high proportion of ineligible participants — those without the disease under study, for whom there is no potential benefit from the intervention. In these circumstances a “quasi” ITT approach (in which ineligible patients are excluded) is more appropriate. 3: Example illustrating the impact of intention-to-treat, per-protocol and treatment-received analyses in a placebo-controlled trial* Treatment group (n = 1000) Control group (n = 1000) Compliers Non-compliers (drop-outs) Compliers Non-compliers (drop-ins)‡ Compliance 80%†‡ 800† 200† 800‡ 200‡ Untreated baseline risk 10% 10% 7.5% 20% Number of events without any treatment 80 20 60 40 Overall event rate 100/1000 = 10% 100/1000 = 10% Expected number of events Expected benefit (relative risk reduction) Full compliance 80 100 20% benefit (1 – [80/100]) Intention-to-treat analysis 64 20 60 32 9% benefit (1 – [84/92]) Per-protocol analysis 64 — 60 — 7% detriment (1 – [64/60]) Treatment-received analysis 80 20 60 32 40% detriment (1 – [112/80]§) Trial assumptions * The average risk of each group is 10% over the long term trial duration, and active treatment, when taken, reduces the risk by 20%. † 20% of those allocated to receive the active drug do not take it because of early side-effects unrelated to the study outcome. ‡ 20% of those allocated to receive the matching placebo medication are prescribed the active therapy because of early clinical deterioration of their condition directly related to their risk of study outcome (these participants are a high-risk subset and have double the average risk [ie, 20%]). § This comprises expected events in those taking the active drug (treatment group compliers and control group non-compliers) divided by those not taking the active drug (control group compliers and treatment group non-compliers). A simple adjustment factor to obtain a better estimate of what might happen with full compliance (100%) compared with observed compliance (80% for each group) can be applied to the ITT benefit (ie, 9% x 100/80 x 100/80 = 13% benefit).

Stephane R Heritier PhD · Val J Gebski BA, MStat · Anthony C Keech MScEpid FRACP

Women's health Letters 6 October 2003 Free

Hormone replacement therapy: to use or not to use?

Michael D Coory Medical Epidemiologist, Queensland Health, GPO Box 48, Brisbane, QLD 4001. michael_cooryAThealth.qld.gov.au To the Editor: The randomised controlled trial associated with the Women’s Health Initiative (WHI) found that long-term hormone replacement therapy (HRT) with combined oestrogen–progestin causes net harm.1 Both the article by Baber and colleagues on HRT2 and a previous editorial by Patel and colleagues3 imply that the method used to calculate the confidence intervals in the WHI report is questionable. Baber et al suggest that “a trial such as this, with multiple endpoints, should use adjusted rather than nominal confidence intervals to test individual endpoints for significance”.2 It is important that this issue is clarified. In the WHI report in JAMA, Table 2 shows both nominal and adjusted confidence intervals for the primary and secondary outcomes.1 Nominal confidence intervals are appropriate for the preselected primary outcomes of the trial — breast cancer, coronary heart disease and the composite global-index score.4 Confidence intervals adjusted for multiple comparisons are possibly appropriate for the multiple secondary endpoints in the study, but are not advocated by all statisticians.5 In any case, the decision of Baber and colleagues to concentrate on adjusted confidence intervals for the preselected primary outcomes is not valid.4 The purpose of confidence intervals is to assess the effects of random variation or chance. It is not sensible to suggest that the extra harm that occurred in the combined HRT arm of the WHI study could be due to chance. Moreover, 42% of women in the HRT group stopped taking the drug, and 11% of women in the placebo group started taking it.1 Therefore, the reported findings of the intention-to-treat analysis underestimated the true harm to individual women taking long-term HRT. Also, if duration of treatment is important (as appears the case with breast cancer risk), and because compliance decreased over time, 5-year results underestimated longer-term treatment harm.4 The aim of the WHI trial was to assess whether long-term HRT is a useful preventive intervention for postmenopausal women. It did not assess the short-term use of HRT to relieve severe hot flushes. As Sackett points out, curative and preventive medicine are absolutely and fundamentally different in their obligations and implied promises to the individuals whose lives they hope to modify.6 As a long-term preventive intervention, HRT causes more harm than good. Although the absolute risks were small, millions of women were prescribed this treatment worldwide, causing harm to thousands. Billions of dollars were spent on an ineffective preventive intervention.6 The thousands of Australian women who stopped taking HRT on learning the results of the WHI trial made a sensible decision.

Michael D Coory

Women's health Letters 6 October 2003 Free

Hormone replacement therapy: to use or not to use?

Rodney J Baber,* Justine L O’Hara,† Frances M Boyle‡ * Clinical Senior Lecturer, Department of Obstetrics and Gynaecology, University of Sydney, NSW, 2006; † Medical Student, ‡ Oncologist, Royal North Shore Hospital, Sydney, NSW. rbaberATmail.usyd.edu.au In reply: We acknowledge that not all statisticians agree on the place of adjusted confidence intervals. However, we and others1,2 believe they represent a conservative choice for secondary endpoints in a study with multiple endpoints, such as the WHI trial. Results of recent randomised controlled trials of hormone replacement therapy (HRT) and cardiovascular disease certainly support the notion that HRT confers no protection. However, any real harm of HRT must be questionable in light of the rapid review by Beral and colleagues, which, also using nominal confidence intervals, showed no change in relative risk for HRT users.3 We are surprised that, having emphasised the importance of nominal confidence intervals for primary endpoints, Coory did not mention that the breast cancer risk in the WHI report was not statistically significant using either nominal or adjusted CIs, or that the global index used was a non-validated instrument designed for and used only in the WHI study.4 Intention-to-treat analysis is used to avoid overestimates of both harm and benefit. While drop-in and drop-out rates (equal in both arms) may have led to underestimates of harm from HRT, they may also have led to underestimates of benefit, with no net change to risk–benefit assessment. The aim of the WHI trial was to assess the benefit or otherwise of long-term HRT on disease processes in otherwise healthy women. There seems little doubt that in the group of older, overweight, somewhat hypertensive, women enrolled in this trial the use of HRT was not beneficial. The aim of our article was to assess the case for and against HRT use.5 In reaching our conclusions, we drew on a broad range of published data, including, but not confined to, the WHI data. Our conclusions make it clear that we believe the use of HRT is primarily for short-term relief of symptoms during the menopause transition. However, we sought to defend the right of a small number of women to choose to continue HRT for long-term improvement of quality of life and symptom relief after appropriate, balanced, individualised counselling about the risks and benefits of such a decision. We do not agree with Coory’s final comment. The thousands of Australian women who stopped taking HRT on learning the results of the WHI trial did so in fear and ignorance in an environment where their physicians were unable to offer balanced counsel — hardly a formula for good medicine.

Rodney J Baber · Justine L O’Hara · Frances M Boyle

Women's health Letters 18 August 2003 Free

Gestational diabetes mellitus: accuracy of Midwives Data Collection

Robert G Moses,* Alison J Webb,† Christine D Comber‡ * Clinical Director, † Nurse, Diabetes Service; ‡ Nurse, Department of Obstetrics and Gynaecology, Illawarra Area Health Service, PO Box W58, Wollongong West, NSW, 2500. mosesrATiahs.nsw.gov.au To the Editor: Gestational diabetes mellitus (GDM) is glucose intolerance of variable severity with onset or first recognition during the current pregnancy.1 GDM is one of the conditions requiring an entry on the New South Wales Midwives Data Form. Effective healthcare planning is dependent on accurate data collection. To our knowledge, the verity of the midwives data with respect to GDM, or indeed other entities, has not been checked for many years. A previous article has demonstrated that the accuracy of GDM data collection is poor, with the incidence of GDM being under-reported.2 Recently, an article from Victoria also showed a recorded rate of GDM about half that of the acknowledged incidence.3 We have recently completed a review of compliance with GDM testing in our area and, knowing the true incidence of GDM, this has allowed us to revisit the accuracy of the data being recorded on the Midwives Data Collection Form. In the city of Wollongong, NSW, with a population of around 280 000 and about 3000 births each year, all deliveries take place at two public hospitals (Wollongong and Shellharbour) and a private hospital (Illawarra Private Hospital). It is the policy of both the Obstetric Department and the Division of General Practice that all pregnant women should be tested for GDM in accord with the ADIPS guidelines.4 All women who delivered at the three hospitals over the 6-month period from January 2002 to June 2002 were identified from the Labour Ward records. A hospital-based delivery is used by 99.3% of women in the area.5 The results of testing for GDM were determined for all of these women. There were 1655 deliveries at the three hospitals over the 6-month period. Seven women with known type 1 or type 2 diabetes were excluded, leaving 1648 women whose data could be examined. Women were considered to have been tested for GDM (n = 1518) if they had had either a glucose tolerance test (n = 1502) or a glucose challenge test (n = 16). There were 101 women diagnosed with GDM, giving an overall incidence rate of 6.6% (prenatal clinic, 7.1%; shared-care, 6.6%; private patients, 6.3%). The most recent midwives data indicate an incidence of 5.7% at the public hospitals and 3.1% at the private hospital. It is thus apparent that the official statistics still underestimate the incidence of GDM. A similar degree of error may also be found for other entries, and hence data should be extrapolated with caution. A redesign of the collection form may help remove some of the errors and omissions. For the question regarding GDM, we feel accuracy could be enhanced if there were separate “Yes” and “No” boxes, rather than a single check box. This might encourage further consideration of the problem. Accuracy could be further enhanced by allowing space for the glucose tolerance test results at 0 and 2 hours — these would also be useful data in their own right.

Robert G Moses · Alison J Webb · Christine D Comber

Women's health Letters 18 August 2003 Free

Gestational diabetes mellitus: accuracy of Midwives Data Collection

Lee K Taylor Manager, Surveillance Methods, Centre for Epidemiology and Research, NSW Department of Health, Locked Bag 961, North Sydney, NSW 2059. ltaylATdoh.health.nsw.gov.au In reply: Moses et al are correct in noting that gestational diabetes mellitus (GDM) is under-reported to the New South Wales Midwives Data Collection (MDC). The most recent validation study of the MDC was carried out in 1998. We reviewed a random sample of 1680 medical records from NSW public and private hospitals, representing 1.9% of births reported in 1998. The sensitivity and specificity of reporting of GDM to the MDC were 86.7% and 99.6%, respectively.1 In this sample, the incidence rate of GDM was 3.5% according to the MDC, and 4.0% according to the medical record review. These population rates are lower than the rates reported by Moses et al among women attending hospitals in Wollongong. In addition to incomplete recording of diagnosed GDM on the MDC, the low rate of recording of GDM in medical records in our sample suggests that GDM was also under-ascertained at a population level. This is probably due to variations in the implementation of pregnancy screening for GDM between clinicians and across NSW hospitals. In February 2003, the Royal Australian and New Zealand College of Obstetricians and Gynaecologists endorsed the Australian Diabetes in Pregnancy Society GDM Management Guidelines.2 The guidelines recommend universal screening for GDM, noting that selective screening may be appropriate because of limited resources or known low GDM incidence. The suggestions for trying to improve reporting of GDM by redesigning the MDC form are welcome, and we will certainly consider them at the next review. We are also considering using the hospital Inpatient Statistics Collection (ISC), in which discharge diagnoses are classified according to the International Classification of Diseases, as an alternative source of information on maternal morbidity. We are currently reviewing a random sample of 500 medical records of mothers who gave birth in hospitals throughout NSW. The information obtained will be compared with matched ISC records provided to the NSW Department of Health to determine whether the ISC is a more reliable source of information on maternal morbidity than the MDC. In the longer term, I anticipate that the integration of the MDC with computerised medical records in hospitals will also contribute to improved reporting. Under-reporting of maternal morbidity, including GDM, is an issue for all state and territory perinatal data collections in Australia. The information is used for planning and evaluation of healthcare services, so it is important that we get it right. I would like to thank Moses et al for raising this issue.

Lee K Taylor

Information science EBM: Trials on trial 21 July 2003 Free

Baseline data in clinical trials

Although reporting baseline data seems simple, it is crucial information for readers in judging the validity of a trial. Knowing the baseline characteristics of the trial participants allows readers to assess how closely these match patients seen in their own clinical practice, and therefore how generalisable the results of the trial will be (so-called external validity). Baseline characteristics also allow the success of randomisation to be assessed. In studies where important baseline factors appear well balanced, it is likely that any differences in outcome between the intervention and control groups are a real effect of treatment (one component of internal validity). For these reasons, the reporting of baseline demographic and clinical characteristics of each group is a requirement of the CONSORT statement.1 The item and its descriptor as they appear in the CONSORT checklist are shown in Box 1, and a checklist for baseline data is provided in Box 2. ContentBaseline data should adequately describe the population in the trial. This means including demographic variables, known factors that influence the outcome (including medications being taken by participants), factors that are likely to modify any benefit of treatment, and those that may predict adverse reactions. These factors are called potential "confounders", because, if they are imbalanced between the treatment groups at baseline, they may result in an apparent treatment effect when none exists, or mask an effect that does exist. Baseline data should also include any factors (especially known potential confounders) that have been used as strata for randomisation. Stratified randomisation, described in detail earlier,2 is used when a baseline characteristic, such as tumour stage, is known to affect outcome risk; the characteristic is therefore included in the randomisation algorithm to minimise imbalances between treatment groups. This is particularly useful in small studies. If the study population contains subgroups of particular interest, the characteristics defining these subgroups, and numbers or proportion in each group, should be stated. For example, in a long-term trial of a new medication for preventing heart attack, diabetes mellitus would be a potential confounder (as people with diabetes have a much higher risk of heart attack than similar people without diabetes). Those with diabetes in this study would also be an interesting subgroup in whom the effects of the intervention might be different. Similarly, concurrent therapy with aspirin (which would substantially reduce the risk of heart attack) could confound the trial results if there was an imbalance between trial groups in the proportions of patients taking aspirin; aspirin therapy might also influence the likelihood of adverse reactions to study therapy. Baseline factors can be determined from interviews, physical examination, laboratory measures or imaging studies. MeasurementBaseline data are measured as close as possible to the time that participants are randomly allocated to study groups, and in all cases, should be measured before the allocated treatment commences (information collected after the commencement of trial treatment may have been altered by the treatment itself, and is generally not regarded as baseline data). Ideally, baseline data should be collected on all patients screened for eligibility, as this would provide further information about the generalisability of the trial population. However, this is not always practicable or affordable, so some variables (eg, tissue biopsy, measurement of genetic markers, expensive imaging tests) are measured only in actual participants randomly allocated to a trial group. For factors that are not constant, the conditions under which the baseline data are collected should be stated in the methods section of the study report. For example, it should be clear whether blood pressure recordings were measured sitting or supine, or after a specified rest period; also whether a single reading, the average of several readings, the highest of two, or the first of two or more, was used. Baseline data as entry criteriaIn some circumstances, threshold levels of one or more baseline variables will form part of the entry criteria for the study. In this case, if an extreme value of a baseline factor, such as high blood pressure, is required to qualify a person for entry into a study, potential participants whose value on the day of screening is more extreme (higher) than their usual level will be more likely to qualify for entry. A second baseline reading of the average blood pressure for this group will be lower and more accurately reflect their usual blood pressure; this is known as regression towards the mean.3 For this reason, remeasuring factors required for entry is desirable, to establish a more realistic group average value of the characteristic at baseline. PresentationThe baseline characteristics are usually presented in the first table in a report. Care should be taken to include the necessary descriptive information without overwhelming readers with unnecessary details. For example, in the recent AFFIRM trial comparing rate control with rhythm control of atrial fibrillation, the published first table has 16 baseline characteristics, each with a mean and percentage value for the overall group, and for both treatment groups separately, together with P values.4 The resulting table of 107 values and four footnotes may make it difficult for some readers to extract the key information.5 A simpler presentation appears in the FRISC II study of invasive compared with non-invasive treatments for unstable coronary artery disease.6 This presents more baseline characteristics (20), but by minimising detail (omitting overall group and P values), allows a more rapid comparison of the characteristics between groups. Comparability between groupsIf randomisation has been performed correctly, the groups should be similar in baseline characteristics, except for the play of chance. Stratification in the randomisation process further restricts the extent of chance imbalances.2 For continuous variables (such as blood pressure, age, cholesterol level), the similarity of the treatment groups should be assessed by comparing relevant summary measures (mean and standard deviation, or median and range). For categorical factors (such as sex, disease stage), the numbers and proportions in each category level should be shown for each treatment group. The more similar the treatment groups, the more credible are the trial results as reflecting a true result of treatment, especially if unadjusted analyses are presented.5,7 Use of P values to assess randomisationUse of statistical tests to compare the balance and/or values of baseline characteristics between the study groups and the presentation of P values are not uncommon. However, many authors assert that this is inappropriate.3,5,8-10 If randomisation has been performed correctly, chance is the only explanation for any observed difference between groups at the outset of the study, in which case statistical tests become superfluous. Consequently, only if it is suspected that the randomisation process has failed or was flawed, can performing significance tests on the baseline data be readily justified.8 It is worth remembering that, if 20 baseline characteristics are presented from a trial using simple randomisation, it is more likely than not that at least one characteristic will show a significant imbalance between groups at two-sided P < 0.05 by chance alone (actual likelihood, 64%). In any case, providing P values is not a substitute for carefully describing, in the results section, any imbalances between study groups that may be clinically important. For example, in a trial of a thrombolytic drug, a 1% baseline difference in history of previous intracranial haemorrhage may not be statistically significant, but could still affect haemorrhagic stroke rates after treatment (an outcome of the study), and hence could be regarded as potentially clinically significant. If there are imbalances that are considered important to the final study results, they should be accounted for by an adjusted analysis of the data, not simply noted with a P value in the first table.7 Other uses of baseline dataA longer-term benefit of collecting comprehensive baseline data is that, after outcome data become available, it allows the estimation of risk of the outcome in the control group, related to various baseline characteristics. This effectively uses the control group as an epidemiological cohort study, providing contemporary information about predictors of disease outcomes. In summary, careful planning and collection of baseline data enables performance of a high-quality trial and allows readers to clearly see the internal and external validity of the study. 1: CONSORT checklist of items to report when reporting a trial 1 Section and topic Item no. Descriptor Baseline data 15 Baseline demographic and clinical characteristics of each group 2: Checklist for baseline data Measurement Consider all important baseline variables to be measured and how they are to be measured before treatment starts: Demographic characteristics (age, sex, height, weight, etc) Known factors that predict the outcome (potential confounders) Factors that predict or alter the risk of adverse reactions Stratification factors Pre-specified subgroups. Reporting Tabulate relevant summary measures (eg, mean and standard deviation). Include all important baseline characteristics while keeping the table readable. Wherever possible, avoid displaying P values. Analysis In the results section, discuss the similarity of the two groups, highlighting any clinically important differences that may influence the outcome. Discussion Discuss the effect of the baseline data balance on the internal validity of the study and the comparability of the study population to patients seen in wider clinical practice.

David C Burgess BMed · Val J Gebski BA, MStat · Anthony C Keech MSc(Epid), FRACP

Statistics Letters 21 July 2003 Free

Statistical methods in clinical trials

Peter J Goadsby Professor of Clinical Neurology, Institute of Neurology, National Hospital for Neurology and Neurosurgery, Queen's Square, London, WC 1N 3BG, United Kingdom To the Editor: Gebski and Keech describe with clarity and accuracy the important basic concepts of statistical analysis for physicians.1 I would like to draw attention to two issues. Firstly, the authors refer to common measurement scales that are used in medicine. It is crucial to understand the limits of a measurement to begin to appreciate results from any study. They describe the continuous scale and offer blood pressure and temperature measurements as examples. This scale refers to data determined such that the distance between any two points is known and measureable. Siegel used the term "ratio scale" if there was a true zero point to the measurement.2 This contrasts to an ordinal categorical scale, in which the intervals are not constant. The scale referred to can be transformed, and is anchored with respect to the measurements to some reproducible point. The term "ratio" for this scale seems preferable, as the world is, in essence, discrete when measured, in the quantum sense. Certainly, the measured world is not continuous, at least as far as we can determine it. Secondly, the authors do not mention resampling methods.3 These can be very powerful and are attractive in biomedical research when the distribution may not be defined. While I realise these methods are relatively new, they do seem unreasonably ignored in undergraduate medical education.

Peter J Goadsby

Statistics Letters 21 July 2003 Free

Statistical methods in clinical trials

Val J Gebski,* Anthony C Keech† * Principal Research Fellow, † Deputy Director, NHMRC Clinical Trials Centre, University of Sydney, Locked Bag 77, Camperdown, NSW 1450. valATctc.usyd.edu.au In reply: While one can view the world as being "discrete", the assumptions underpinning most common statistical methods in analysis of clinical studies are "continuous" distributions. In fact, statisticians go to enormous lengths to approximate discrete systems as continuous ones (lifetime analysis, normal approximations, etc). The measurement scale by which study outcomes are assessed needs careful consideration (to ensure consistent precision and units of measurement). However, both practical and statistical considerations allow for the more common definitions of continuous and discrete measurements to be just as effective for statistical comparisons. Indeed, there is frequently little loss of statistical efficiency when "continuous" variables are appropriately categorised into ordinal groups.1 Resampling methods randomly sample the data repeatedly to estimate the underlying population distribution parameters (eg, mean, standard deviation, etc). They can be very useful in solving specific problems in which the underlying properties of the data used to make treatment comparisons are unknown and using other statistical approximations is deemed to be inappropriate. However, these are specialised computer-intensive techniques for use by trained biostatisticians, rather than commonly used analysis methods. Problems arise with resampling techniques (eg, obtaining confidence intervals), which require specialised statistical expertise.

Val J Gebski · Anthony C Keech

Epitaph for the EBM in action series

This issue of the Journal (page 575) features the last article of the EBM in action series,1 conceived to show how clinicians can effectively look for the best available evidence to answer clinical questions. In the current medical climate, clinicians clearly need systems to obtain the best available evidence, and the responsibility for creating these systems falls on both individual clinicians and the organisations for ...

Christopher B Del Mar MD FRACGP · Jeremy N Anderson MD FRANZCP

Statistics EBM: Trials on trial 2 June 2003 Free

Recruitment to randomised studies

Maintaining adequate and consistent accrual of patients is important for small clinical trials and crucial to large multicentre, multinational clinical trials, in which participants often number in the thousands. Trial recruitment that is slower than expected can result in prolonged trial durations, increased costs, uneven workload, and morale problems, both for trial participants and trialists.1 Furthermore, when recruitment is so poor that it needs to ...

Wendy E Hague MBA, MB BS · Val J Gebski BA, MStat · Anthony C Keech MscEpid, FRACP

Statistics 7 April 2003 Free

Clinical trials research in the new millennium: the International Clinical Trials Symposium, Sydney, 21–23 October 2002

The International Clinical Trials Symposium, held in Sydney in October 2002, brought together over 700 people to explore the issues and future directions for clinical trials. The support from government as well as industry and the diversity of people attending the symposium indicated the importance of clinical trials research in Australia. Established principles and new challengesAfter about 50 years of developing and refining the design and conduct of clinical trials, researchers have established the fundamentals. Randomisation, blinding, informed consent, adequate sample size and a prospectively stated study design are well accepted as means of producing unbiased evidence of the effectiveness of therapies. However, the symposium brought to light various issues that continue to challenge clinical trials researchers. Prominent among these were: how to make the conduct of trials more relevant to clinical practice (and, conversely, how to draw more of "real-world" care into the context of trials); the importance of consumers participating (not just as patients, but by voicing research questions); how to ensure that trials research concentrates on the clinical questions most needing answers (rather than on, for example, comparable products jostling for market share); and how success in these areas can practicably be achieved (not least by using technology to run trials efficiently). Some of these themes have been aired in the Journal.1 Bridging the gap between researchers and practitioners and patientsStephen Blamey (Chairman, Medical Services Advisory Committee, Commonwealth Department of Health and Ageing, Canberra) set the agenda in his opening address by stressing the importance of evidence-based medicine in allocating government funding. On behalf of the users of trial evidence, David Henry (Professor of Pharmacology, University of Newcastle) recommended that purchasers of services be represented when trials are being designed (and possibly share the cost of trials). Sue Lockwood (Chair, Breast Cancer Action Group, Melbourne), using the example of the Australian Sentinel Node Biopsy Trial (Royal Australasian College of Surgeons),2 showed that involving consumers ensures that recruitment is fast and that the outcome is evidence that matters to patients. Bob Temple (Director, Office of Medical Policy, US Food and Drug Administration, Rockville, Md, USA) raised the issue of the need for regulatory bodies to have a say in trial design to ensure that trials meet their objectives for rigour and new indications. Researchers can reach out to consumers by publicising the results of trials. Although journals have traditionally been the providers of this information, Richard Horton (Editor, Lancet) thought that promoting the results of research in other ways, such as communications aimed directly at consumers, is what really changes practice. Evidence may become known only from discussion in other journals and the mass media. As noted by Sally Redman (Chief Executive Officer, Institute for Health Research, University of Sydney), results published in specialised journals are often translated into practice by way of guidelines. "I carry a card that says, 'Invite me to participate in all randomized controlled trials for which I am potentially eligible'," said Iain Chalmers (Founder of the UK Cochrane Centre). He added that participation in controlled trials should become a widely available treatment option within routine healthcare. It was pointed out early in the symposium by Paul Glasziou (Professor of Evidence-based Practice, School of Population Health, University of Queensland), and others, that patients participating in clinical trials appear to do better than those receiving standard care outside of trials.3 Refinements in trial design, such as recruitment of clusters of patients rather than individuals (Judy Simpson, Associate Professor, Department of Public Health and Community Medicine, University of Sydney), and more control of confounders to widen entry criteria, can mean that more patients will be able to take advantage of participating in a trial. When commercial confidentiality is not an issue, trials can be advertised on web sites or public trials registers so that patients themselves can take the initiative in enlisting. Which diseases and treatments need evidence?Various speakers identified the many areas where more trials research is needed. Richard Horton championed an international perspective and the need for trials research in developing countries, focusing on diseases affecting people in these regions. In terms of the numbers of trials and the numbers of patients recruited, cardiovascular disease and cancer are way ahead. This is partly because of the large number of people with these diseases and the number of new treatments being developed. A challenge for researchers is to diversify to other areas. An innovative afternoon session comprised 12 concurrent forums in different clinical specialties. The objective was to identify current research priorities for particular clinical areas, and to develop new trial proposals or address concerns about trial methodology, such as measurement of outcomes, methods of randomisation and recruitment of patients. Besides cardiovascular disease and cancer, the forums embraced complementary medicine, diabetes, general practice, HIV medicine, perinatal medicine, reproductive and gynaecological medicine, rheumatology and surgery, as well as health technology assessment and information technology in research. The symposium also included speakers whose work focused on laboratory rather than human research. Some urged that the need for increased support of clinical studies should not compromise basic science research. A vision of the possibility of selecting the right drug for an individual patient is becoming a driving force in drug development (Peter Shaw, Director, Human Genetics and Pharmacogenomics, Bristol-Myers Squibb, Pennington, NJ, USA), particularly in oncology, where some drugs are especially effective for patients with certain genetic profiles. In HIV research, characterisation of viral gene sequences is affecting all aspects of trial design and analysis, adding to their complexity (Victor DeGruttola, Professor of Biostatistics, Harvard School of Public Health, Boston, Mass, USA). Systems for conducting trials more efficientlyMike Conlon (Chief Information Officer, University of Florida Health Science Center, Gainesville, Fla, USA) related that internet-based systems can reduce overall operational costs by a factor of three and significantly reduce workloads. Several groups in Australia are also working toward eliminating paper in the day-to-day running of trials. In view of the need for more and better trials in more areas of healthcare, these efficiencies will be essential. The necessity for an Australian trials registryPerforming trials is an expensive business. Researchers need to know about every trial in their field so as not to duplicate what is already being done (John Simes, Director, NHMRC Clinical Trials Centre, University of Sydney). Most of the trials in Australia are still unregistered (Paul Glasziou). Patients and their doctors want to know which trials are available and suitable. Ensuring that everyone can find out which trials are under way requires central coordination. The way forward would be an Australian register of clinical trials. A national register is essential for planning, for maximising patient participation, and for identifying relevant randomised trials for studies compiling the available evidence. Tony Keech (Deputy Director, NHMRC Clinical Trials Centre, University of Sydney) stressed that such a register would also facilitate meta-analyses, which ideally should be planned prospectively. Panel sessions — debates and hypotheticalsWhether (subject to consent and safety) every patient should be in a clinical trial was savagely contested by Richard Horton and Harvey White (Director, Coronary Care and Cardiovascular Research, Green Lane Hospital, Auckland) and David Celermajer (Professor of Cardiology, University of Sydney) and Martin Tattersall (Professor of Cancer Medicine, University of Sydney). Horton and White asserted that the key to better patient recruitment lay with doctors — they must offer more trial opportunities to their patients. And simpler inclusion criteria would allow typical patients with multiple morbidities to participate. Celermajer and Tattersall insisted that proof-of-concept trials do not require typical patients and that entry criteria do not require broadening. A hypothetical, moderated by David Celermajer, took a make-believe hospital ethics committee through an evolving scenario. The committee members agreed on the principles of justice, equity and non-maleficence, and that the trial research must be of adequate quality and not harmful to patients. However, when trials researchers want to extend the protocol to reuse data and analyse blood samples, a green light from the ethics committee may come only after strong debate. If blood analysis reveals a risk of disease, should the risk be disclosed to individual patients? If so, should it also be disclosed to their relatives? Some of the ethical dilemmas showed how closely the conduct of trials reflects usual medical practice, in that many ethical issues are the same. The ultimate panel session was moderated by Norman Swan (Producer and Presenter of the Health Report, Australian Broadcasting Corporation, Sydney), who examined the main themes emerging from the symposium with a broad-based panel. The panel recommended the following areas of action. More trials, especially prospectively designed trials in resource-poor countries; AIDS and malaria should have priority. Registration of trials. Trials in everyday practice. Development of funding mechanisms for trials. Guidelines as a means of translating trial evidence into practice. Trials of treatments (other than drugs) and technologies, to build on the work done by the pharmaceutical industry. Advancing alliances of triallists and government, particularly the NHMRC. ConclusionClinical trials have come of age. Researchers are now critically evaluating their methods to improve the precision of trial results and reduce bias. At the same time, a widening of inclusion criteria for patients entering trials is allowing more patients to take part. Therefore, the trial environment is becoming more like standard clinical practice; indeed, the next step may take us toward integrating trials with medical care so that research is part of the healthcare system.

Rhana Pike MA · Anthony C Keech FRACP, MScEpid · R John Simes FRACP, SM

Statistics 7 April 2003 Free

Flow of participants in randomised studies

In judging the results of randomised trials it is important to know from where and how participants were recruited, to what extent they received the intended interventions, whether they were followed up as planned, and whether their data were analysed as stated. These details are to ensure that readers of trial reports can appreciate both how closely participants reflect those more generally suffering from the condition under investigation, and how reliably the trial's results test its hypothesis. Participant flow diagramEnrolmentItem 13 of the CONSORT statement recommends a flow diagram to aid in the reporting of participant flow (see Box 1).1 Box 2 provides a checklist for tracking subject participation throughout the trial. In this scheme Part A refers to the number of participants with the condition of interest screened for eligibility criteria as specified in the trial protocol. For a recent example see the Second Australian National Blood Pressure trial.2 This study clearly details the process resulting in the final 6083 participants recruited. With 54 288 people screened to participate, 31 255 had the condition of interest (hypertension). Of these, 8273 were found to be ineligible and 16 899 refused to participate. The remaining 6083 were randomly allocated, corresponding to Part C of the flowchart in Box 2. The ratio of participants randomly allocated to those initially assessed helps determine how generalisable the results of the trial will be, and consequently may also affect the extent to which the results of the trial might influence health policy. Part B of Box 2 indicates assessed participants who do not subsequently participate, with reasons for non-participation given. Enough information should be given to identify separately the numbers who were deemed ineligible, refused to participate and those not randomly allocated to an intervention for other reasons. The ratio of the number of participants to the number of people initially assessed for eligibility may also provide an insight into the acceptability and practicability of the intervention. For example, if 2000 people were assessed and only 300 recruited, such a low ratio might be the result of highly restrictive eligibility criteria, participant requirements that are too complicated or impractical, or an intervention too intrusive for participants to readily accept over standard care. AllocationOf the total number of participants randomly allocated, the number assigned to each of the study arms should be separately presented (Box 2, Part D). A breakdown of the number of participants who actually received the allocated treatment and those who did not (with reasons given) should be included. This information, together with details of participant follow-up (see below), helps determine how well the intentions of the protocol were met. Follow-up detailsDuring the trial, participants' status in terms of outcomes (both efficacy and safety) is usually ascertained at predefined time intervals. However, some participants may withdraw from the study before completion; these participants are classified as "lost to follow-up" from that point onwards. Details of the number of participants who are lost to follow-up in each of the study arms are essential (Box 2, Part E), as this provides information on the reliability of the study's conclusions. When participants cannot be accounted for at the end of the study, their outcome status cannot be determined. If substantial numbers of participants are lost to follow-up, concerns may arise about the integrity of any observed effect of an intervention. A common approach to evaluating the potential influence of losses to follow-up is to use a sensitivity analysis, where worst-case and best-case scenarios relating to these losses are examined. For example, at one extreme, it could be assumed that all the control-arm and none of the intervention-arm participants who were lost to follow-up suffered the outcome of interest. At the other extreme, the opposite (ie, all intervention-arm participants lost to follow-up and none of the control participants lost to follow-up have suffered the outcome of interest) could be assumed. This provides bounds for the maximum possible influence of such losses on the observed treatment effect. Less extreme (more plausible) scenarios can also be examined. If this results in the effect of treatment disappearing or reversing direction, the robustness of the results should be questioned. Box 3 provides an example. Compliance lossesA further issue in interpreting study results is the degree to which participants adhered to their allocated treatment during the study period. Participants who stop or never take their allocated treatment are usually called "drop-outs", and those who begin active treatment when they have been allocated to the control group are called "drop-ins"; all are considered compliance losses. Compliance losses reduce the study power and dilute the observed effects of treatment, and should be documented in the participant flow (Box 2, Part E). Where such non-compliance can be predicted in advance, its potential effect on the power of the study may be reduced by increasing the study sample size.3 Differential compliance rates may offer clues as to the real side-effects of an intervention, or about the success of methods used to blind participants or clinicians to treatments.4 AnalysisPart F of Box 2 relates to the number of participants who were included in the statistical analyses. If any other than those lost to follow-up have been excluded from analysis, both the number in each treatment arm and the reasons should be detailed. As excluding patients from analysis can potentially undermine the effectiveness of the randomisation4 and produce comparisons which may no longer conform to the intention-to-treat principle,5 the nature of, and justification for, any such exclusions should also be provided. ConclusionsDifferent study types may slightly alter the participant flow diagram. For example, a cluster-randomised study will enumerate the clusters of randomised subjects (ie, the units of randomisation, such as whole communities, schools, hospital departments) rather than the number of individuals. Participant flow diagrams are an effective method of summarising all stages of key trial processes. They suggest how these processes should be reported in the trial, specifying separately trial enrolment, treatment allocation, subject follow-up and statistical analysis. In reporting the results of randomised studies, all individuals originally considered for participation in the study should be accounted for in the participant flow diagram. The diagram also provides an overview of many aspects of study quality that can have a major influence on the generalisability and reliability of the conclusions. 1: CONSORT checklist of items to report when reporting a trial Section and topic Item no. Descriptor Participant flow 13 Flow of participants through each stage of a clinical study (a diagram is strongly recommended). Specifically, for each group, report the numbers of participants randomly assigned, receiving intended treatment, completing the study protocol, and with data analysed for the primary outcome. Describe deviations from the planned study protocol, together with reasons. 2: Checklist: flow diagram of the process through key stages of a randomised trial1 3: Sensitivity analysis for a hypothetical study of 1000 participants with 300 events observed, but 84 participants lost to follow-up Patient no. Basis of analysis Intervention Control Odds 95% CI P Observed events only (lost subjects excluded) 130/440 170/476 0.76 0.57–0.996 0.047 All lost participants in placebo group, but none in intervention group assumed to have suffered an event 130/500 194/500 0.55 0.42–0.73 < 0.001 All lost participants in intervention group, but none in placebo group assumed to have suffered an event 190/500 170/500 1.19 0.92–1.54 0.19 Comment: The study had lost enough participants to follow-up to potentially nullify any conclusion of a significant treatment benefit based on the most extreme assumptions about event occurrence in lost subjects. A more plausible scenario might be to assume that the event rate among lost subjects was twice that seen in those with full follow-up. In this instance, up to 35 of the 60 lost subjects in the intervention arm and 17 of the 24 in the placebo arm could be anticipated to have experienced the event. The boundaries this yields (odds, 0.59; 95% CI, 0.45–0.77; P < 0.001 to odds, 0.96; 95% CI, 0.74–1.24; P = 0.74) again calls into question whether a definite effect of treatment can be concluded.

Burcu Cakir MPH · Val J Gebski MStat · Anthony C Keech FRACP, MSc(Epid)

Statistics 7 April 2003 Free

Managing the resource demands of a large sample size in clinical trials: can you succeed with fewer subjects?

To the Editor: Keech and Gebski recently discussed some strategies for answering randomised clinical trial (RCT) questions with fewer subjects.1 We would like to point out another alternative for addressing this important topic — adjustment for baseline characteristics.2-4 Heterogeneity among patients participating in RCTs is common. Prognosis may vary according to important baseline characteristics, which are commonly recorded in RCTs. Heterogeneity may lead to imbalanced treatment arms, even after proper randomisation.3 Covariate adjustment for baseline characteristics is a statistically efficient procedure. It leads to more individualised treatment-effect estimates, corrects for imbalance and improves statistical power.2,3 Hence, it may potentially reduce the necessary sample size of an RCT for the same power as unadjusted analyses. Nevertheless, covariate adjustment is not commonly performed in the RCTs reported in major medical journals.5 We recently performed a simulation study using logistic regression models in the context of RCTs with dichotomous outcomes and one simple dichotomous baseline characteristic in addition to the treatment indicator variable. Covariate adjustment was found to potentially reduce the sample size between 3% and 46%, in direct relation to the strength of the baseline characteristic (odds ratio, 2 to 30). Results of a simulation study in RCTs with survival outcomes, using Cox proportional hazards models, yielded similar results. Covariate adjustment for well-known and important predictors of patient prognosis is a useful tool for potentially reducing the sample size of RCTs, and should be considered more often in their design and analysis.

Adrián V Hernández · Ewout W Steyerberg

Statistics 7 April 2003 Free

In reply: Covariate adjustment for prognostic baseline characteristics may decrease the sample size

In reply: While we agree that covariate adjustment during analysis can be a potential mechanism for reducing sample size (even when there is no imbalance in the important covariate levels between the treatment groups), unless such analyses are prospectively planned then they will not allow valid statistical inference. This is because post-hoc adjustment is an exploratory procedure and may have involved examining any number of potential covariates. Further, to quantify any anticipated sample-size gains would depend on specifying likely maximum covariate imbalances, overall covariate distributions and plausible effects of treatment within the covariate levels during study design. In practice, the study would then have to meet these assumptions for the calculated sample-size gain to be achieved. Covariate adjustment is an accepted practice for subsidiary analysis in clinical trials, and can take account of differential effects in imbalanced subgroups. For example, see the case of an apparent chance imbalance in numbers of women in the treatment arms of the HERO-2 trial, where investigators presented both unadjusted and adjusted results.1 Where important predictors of the clinical outcomes are expected to be variable for the population under study, a particularly useful approach is to stratify the randomisation by those predictors.2 Such stratification allows for valid adjusted analyses.3

Anthony C Keech · Val J Gebski

Statistics 7 April 2003 Free

Determining the sample size in a clinical trial

To the Editor: Evidence-based medicine should be supported by randomised controlled trials (RCTs) that show the efficacy of interventions in producing clinically relevant outcomes, not by those that show statistically significant, but clinically irrelevant, differences. RCTs are designed to investigate whether an intervention in one homogeneous group results in a different outcome compared with no intervention, or a different intervention in an otherwise identical group. Statistical analyses are performed to estimate the probability that any difference in outcome has arisen as the result of chance alone, and sample size is determined to control the probability of a real difference in outcome being overlooked by chance alone. However, statistical analyses provide no information about the clinical relevance of any difference in outcome. There is no validated method for determining a minimum clinically relevant difference in outcome (minimum important difference). Kirby and colleagues suggest that, wherever possible, the minimum important difference in response should be determined from Phase II or pilot studies and expert opinion from colleagues.1 It is ironic that the clinical relevance of Level I evidence2 depends on determining the minimum important difference based on Level IV or Level V evidence. Although the original CONSORT statement recommended describing the minimum important difference and indicating how the target sample size was projected,3 the most recent statement is less specific and only recommends describing how the sample size was determined.4 Neither statement requires investigators to specifically describe the method by which the minimum important difference was determined. The minimum important difference must be justified so others can determine if the study has the power to detect a clinically relevant difference in outcome as the result of a particular intervention. Similarly, the minimum important difference must be stated so that any statistically significant difference in outcome can be judged for clinical relevance. If the minimum important difference cannot be justified as being clinically relevant, the result of the study will be of statistical interest only, and valuable resources will have been wasted. While it is reasonable to suggest that sample size must be planned to ensure that research time, patient effort and support costs invested in any clinical trial are not wasted,5 manipulating the minimum important difference to allow an RCT to conform to these constraints cannot be justified unless the RCT can still detect a clinically relevant difference in outcome.

Owen D Williamson

Statistics 7 April 2003 Free

In reply: Determining the sample size in a clinical trial

In reply: A key message of our article is that the minimum possible difference that would render the intervention clinically worthwhile needs to be determined in the design phase of the study.1 This potential clinical difference (net advantage over standard care) must, of necessity, incorporate the potential trade-offs between any outcome advantages and associated toxicities and/or cost disadvantages. In determining the minimum clinically worthwhile difference, the intention is not to ignore Level I evidence (phase III studies and meta-analyses) when such evidence exists. However, in most cases of designing a phase III study, such evidence is not available. In this instance, the use of phase II information (which can often provide good estimates of potential side effects and toxicities) can help to inform the estimated advantage in clinical outcome which would be needed to justify widespread use of the intervention. Tools have been developed to help clinicians determine worthwhile benefit against toxicity trade-offs for individual patients.2 While there are instances of study results reliably showing statistical significance without the measured effect being sufficiently large to be considered clinically relevant, this occurs rarely. Far more commonly, statistically non-significant results may obscure clinically relevant outcomes because studies have been seriously underpowered.3 For example, 17 of the first 22 small trials of thrombolytic therapy compared with placebo for acute myocardial infarction reported non-significant results, although they showed a 19% reduction in early mortality (2P < 0.01) in subsequent meta-analysis.4 Individual doctors treating their patients ultimately determine whether net differences are sufficiently worthwhile to change their clinical practice. The fundamental principle is that clinical trials should be designed with sufficient power (through having appropriate sample sizes) to detect differences that doctors would consider clinically worthwhile to improve health outcomes.

Adrienne Kirby · Val Gebski · Anthony C Keech

Child health Viewpoint 17 March 2003 Free

A new longitudinal study of the health and wellbeing of Australian children: how will it help?

The Longitudinal Study of Australian Children (LSAC) is a major research endeavour to assess emerging health and developmental concerns and their determinants in children. Previous longitudinal studies of children in Australia and New Zealand have contributed significantly, but have had limitations, which LSAC will attempt to address. A new generation of longitudinal studies is needed to enable transnational and historical comparisons. Members of the LSAC (Longitudinal Study of Australian Children) Research Consortium John Ainley, Australian Centre for Educational Research; Donna Berthelsen, Queensland University of Technology; Michael Bittman, University of New South Wales; Dorothy Broom, Australian National University; Linda Harrison, Charles Sturt University; Bryan Rogers, Australian National University; Michael Sawyer, University of Adelaide; Sven Silburn, Curtin University; Lyndall Strazdins, Australian National University; Judy Ungerer, Macquarie University; Graham Vimpani, Newcastle University; Melissa Wake, Murdoch Children's Research Institute; Stephen Zubrick, TVW Telethon Institute for Child Health Research. In March 2002, the Commonwealth Department of Family and Community Services announced the commencement of the Longitudinal Study of Australian Children (LSAC).1 This study is being implemented by a large multidisciplinary research consortium led by the Australian Institute of Family Studies. It will track the health and development of two national, population-representative cohorts of children recruited in their first and fourth years of life. The study will assess a broad range of individual, family and environmental determinants of health and wellbeing, and will focus on identifying factors that influence good and poor life-course outcomes. The study aims to provide data that will inform the development of health, family and social policy and services within Australia (Box 1). LSAC represents a significant government investment ($20.2 million over nine years) in longitudinal research on children, and it is timely to consider the extent to which it will address the shortcomings of past studies and add to our knowledge of children's health and development. Longitudinal studies are essential to understand the causes of health problems and identify possible solutions.2 Using this type of study, Australian and New Zealand researchers have contributed significantly to our knowledge of health and development3-7 (Box 2). However, there are several reasons why Australia needs a new longitudinal study of children. Past studies typically recruited cohorts during the 1970s and 1980s8 and are therefore limited. First, the health profile of Australian children has changed significantly in recent decades. Emerging health concerns include increasing rates of atopic and chronic diseases and mental health problems, and unacceptably high levels of suicide, preventable injuries and harmful health behaviours during childhood and adolescence.9 Second, the environments in which children are being raised have altered profoundly in the last 30 years.9,10 There have been changes to children's immediate environments (family, childcare, schools and neighbourhoods) and the broader sociopolitical climate (including widening social disparities).10 Past study findings may not be relevant to modern childhood environments. In addition, past cohorts were typically drawn from confined geographical locations,8 and data on health and social services and policies may not be more widely applicable. Third, there have been considerable advances over the last 30 years in theory, measurement tools and analytic techniques.11 Current epidemiological models highlight the central role of individual genetic and pathobiological factors in the expression of poor health and the need to examine multiple levels of influence (Box 3).12 Past studies have been unable to explore these factors adequately, often lacking sufficient sample sizes or the appropriate measures to disentangle multilevel influences. LSAC addresses a number of these limitations. Specifically, it will: measure a wide range of outcomes and determinants; recruit cohorts nationally from rural–regional and urban settings; and have a sufficient sample size to explore multiple determinants and the occurrence of relatively rare events.1 However, it is unrealistic to expect this study to address all research needs. For example, LSAC will not capture sufficient children from minority groups to enable exploration of life-course pathways that may be unique to these populations. In addition, cost and methodological constraints are likely to result in the exclusion or under-representation of children from remote areas, and some forms of data collection will be too time-intensive or costly to use (eg, certain observational, biological or environmental measures). Also, LSAC will not involve the systematic provision of interventions. Given the growing evidence of the benefits to health that arise from prevention and early intervention,13 a strong case can be made for a coordinated program of longitudinal intervention trials to assess the impact of interventions delivered around key life-course transition times.14 Thus, studies of specific populations, studies addressing research questions that require costly data collection, and studies that involve the assessment of interventions could all add value. These could be designed in parallel with LSAC, or as studies nested within LSAC for creating efficiencies in research costs. Australia is unique in geographical distribution of the population, family structures, ethnic diversity, social structures, policies and service provision. International research may have limited applicability in the Australian context. Nonetheless, it is important to consider Australian longitudinal research in an international context. Other Western nations are establishing new longitudinal studies,8 with European countries notable for coordinating studies that enable cross-country comparisons.15 LSAC may provide a foundation upon which other studies could be built to facilitate transnational comparisons and comparisons of changes to health pathways over time. Use of common design and measurement tools will facilitate these objectives. Longitudinal studies are expensive and demanding, not to be embarked upon lightly.2 Coordinated research efforts that strategically build upon the substantial national investment in LSAC may further enrich the evidence base for policy development and service provision to facilitate our nation's future health and wellbeing. 1: Questions to be addressed by the Longitudinal Study of Australian Children1 How well are Australian children doing on key developmental outcomes? What are the pathway markers, early indicators, or constellation of behaviours that are related to different child outcomes? How are child outcomes interlinked with children's wider circumstances and environment? In what ways do features of children's environment (such as families, communities and institutions) affect their outcomes? What helps maintain an effective pathway, or change one that is not promising? How is a child's potential maximised to achieve positive outcomes for children, their families and society? What role can government play in achieving these outcomes? 2: Achievements of past Australian and New Zealand longitudinal studies of children's health and development Study (starting date) Achievements Australian Temperament Study (1983)3 Clarified the contribution of temperament, family and environmental factors to later life adjustment. Christchurch Health and Development Study (1977)4 Contributed to child health, family and mental health policy, safety regulations for swimming pools and bicycle riding, and the development of early intervention programs for high risk mothers of infants. Dunedin Multidisciplinary Health and Development Study (1972)5 Findings across a range of physical and psychosocial health areas contributed to the development of health interventions for substance use, safe driving and cycling practices. Port Pirie Cohort Study (1979)6 Assessed the effects of environmental lead exposure and provided the impetus for changes in regulations relating to lead in petrol. Tasmanian Infant Health Study (1988)7 Assessed possible causes of sudden infant death syndrome, resulting in changed recommendations for infant sleeping positions, with documented reductions in infant deaths. 3: Ecological model of health across the life-course (modified from Lynch12)

on behalf of the LSAC Research Consortium

Endocrinology EBM: Trials on trial 17 February 2003 Free

Does a combined program of dietary modification and physical activity or the use of metformin reduce the conversion from impaired glucose tolerance to type 2 diabetes?

Trial: Diabetes Prevention Program Research Group. Reduction in the incidence of type 2 diabetes with lifestyle intervention or metformin. N Engl J Med 2002; 346: 393-403. QuestionCan treatment with lifestyle modification (changes in diet and physical activity) or metformin reduce the conversion from impaired glucose tolerance (IGT) to type 2 diabetes? Do these treatments differ in effectiveness? Trial details Design: A three-arm multicentre, stratified, randomised controlled trial. Setting: 27 centres in the United States. Patients: 3234 (mean age, 50.6 years; 45.3% non-white; 67.7% female). Inclusion criteria were age 25 years or older; body mass index 25 kg/m2 or more if white, 24 kg/m2 or more if Native American, or 22 kg/m2 or more if Asian; fasting plasma glucose level of 5.3–6.9 mmol/L or < 6.9 mmol/L if Native American; a 2-hour plasma glucose level of 7.8–11 mmol/L after a 75 g glucose tolerance test (GTT); and no previous history of diabetes except gestational diabetes. Patients taking medication affecting glucose tolerance or with a disease which would affect life expectancy or participation in the activity recommendations were excluded. Intervention: The three groups were standard lifestyle recommendations plus twice-daily placebo (control group); intensive lifestyle program plus twice-daily placebo (lifestyle group); and standard lifestyle recommendations plus metformin (metformin group). The intensive program aimed to reduce patients' weight by at least 7% by dietary means and to have them engage in physical activity of moderate intensity for at least 150 minutes a week.1,2 The program was taught in 16 one-to-one lessons followed by individual and group sessions. The standard lifestyle intervention included similar information to the intensive program, but this was given as a written brochure and advice at the annual visit. Patients taking metformin were given one 850 g tablet plus one placebo per day for the first month and two metformin tablets daily thereafter. Main outcome measures: Progression from impaired glucose tolerance (IGT) to diabetes on the basis of six-monthly fasting plasma glucose measurements and an annual 75 g oral GTT. If a result met the 1997 American Diabetes Association (ADA) definition of diabetes,3 the test was repeated within six weeks. If the repeat result also met the ADA definition, the primary endpoint was reached. Otherwise, the patient continued in the assigned group. Main results: The trial was stopped one year early after an average follow-up of 2.8 years. Compared with the control group, the rate of type 2 diabetes was reduced by 58% (95% CI, 48%–66%) in the lifestyle group and by 31% (95% CI, 17%–43%) in the metformin group. The three-year cumulative incidence was 14.4% in the lifestyle group, 21.7% in the metformin group and 28.9% in the control group. The general pattern of the results did not vary by age, sex or race. At the final visit, 38% of the lifestyle group had reduced weight by 7% or more, average fat intake had declined by 6.6% from a baseline of 34.1% of total energy, and 58% were achieving the activity goal. More than 70% of patients took at least 80% of their medication. Hospitalisation and death rates were not different among the three groups. Conclusion: Both lifestyle modification and metformin reduce progression rates from impaired glucose tolerance to diabetes but lifestyle changes were more effective than metformin. The number needed to treat to prevent one case of diabetes in three years is 6.9 for lifestyle and 13.9 for metformin. CommentaryRationale for the trialThe prevalence of type 2 diabetes and its precursor stages is increasing. In 1999, the prevalence of diabetes was 7.4% and of impaired glucose tolerance (IGT) and abnormal fasting blood glucose (FBG) level was 16.4% in Australians aged 25 years and older.4 People with IGT have an increased risk of macrovascular disease, but not of microvascular disease. IGT is associated with obesity, sedentary lifestyle and increasing age, and is more common in some racial and ethnic groups. The only well conducted previous study of lifestyle modification was smaller and included only white people.5 Previous studies of pharmacological agents to reduce the conversion rate had been underpowered.1 Trial methodsThis was a well-conducted study with great attention to detail. For example, the requirement that an endpoint was diagnosed based on the results of two GTTs or two FBG tests was a potential source of bias and unblinding for participants and trial staff. This was addressed by retesting a sample of patients with normal results on these tests and not disclosing the result until progression to diabetes was confirmed.1 Patients were randomly assigned to groups only after completing an extensive run-in phase. An intention-to-treat analysis was done. Sample size for the trial was based on having a 90% power to detect a 33% reduction in the expected incidence of diabetes (6.5% per year). Greater methodological detail is available elsewhere.1,2 Although the focus was on progression to diabetes, reversion to normoglycaemia was also reported. At one year, more than 20% of the control group and 40% of the lifestyle group had normal values for both fasting and post-load glucose levels, and this declined to about 20% and 30%, respectively, at three years' follow-up. This highlights the need for a control group when evaluating interventions. Without this, most of the reversion in the intervention group might have been attributed to the intervention. Instead, it is clear that much of the reversion is related to "regression to the mean".6 That is, when a group defined using a cut-off point in a measure with substantial intra-individual variability is retested, the average value on the second test will be closer to the total population mean.6 New informationThis study is the first to test lifestyle against pharmacological prevention for type 2 diabetes and also the first to include groups that are often under-represented — the elderly, women and non-white people. The results confirm the findings of the previous lifestyle trial5 and extend their generalisability. Whether the interventions actually prevent diabetes, or simply delay its onset, or what happens if the interventions stop, cannot be answered by either study. The variability of glucose levels has been previously described. However, the size of this trial means that the magnitude of the reversion to normal is probably a good estimate of what would happen in the clinical setting in patients who were retested. The size of the variability also has implications for interpreting the results of national surveys such as AusDiab,4 as it means that the proportion of people with abnormal glycaemia found on a single test overestimates the proportion who would have had the abnormality confirmed on a later test. Implications for clinical practiceThis study shows that 5–6 kg weight loss combined with dietary modification to reduce fat intake to less than 30% of energy and increasing activity for 150 minutes per week will halve the conversion rate to diabetes in people with IGT. Because the mean baseline body mass index (BMI) was 34.0 kg/m2, losing 5–6 kg reduced this to about 31.5 kg/m2, which is still in the obese range. Patients should not be given the false impression that they must achieve a body weight in the healthy range (BMI of 18.5–24.99 kg/m2) before any benefits occur. As Native Americans and Pacific Islanders were included, it is reasonable to conclude that the results of this study would apply to Indigenous Australians, who have high levels of diabetes and cardiovascular mortality.7 At present, it is not possible to separate the effects of dietary and activity change on the outcomes, so both need to be recommended together. The intervention for the control group was similar to what a general practitioner might do during a consultation. A lot of support was given to help the lifestyle group patients achieve their lifestyle changes,2 which means that this is not a cheap intervention. The challenge for clinical practice is to provide this support either directly or by referral to community groups and to advocate for environmental changes that support beneficial lifestyle changes.

Dorothy EM Mackerras BSc, MPH, PhD GradDipNutrition

Statistics EBM: Trials on trial 17 February 2003 Free

Statistical methods in clinical trials

Appropriate statistical methods for analysing trial data are critical for the correct interpretation of the results. Item 12 of the CONSORT statement (Box 1) relates to the statistical methods used in the reporting of trials, together with scientific and statistical principles concerning analyses of subgroups, endpoints and appropriate statistical tests. These issues need to be carefully considered before beginning a study and should be outlined in a standard trial protocol, which may be supplemented by a more extensive statistical analysis plan.1 1: CONSORT checklist of items to report when reporting a randomised trial1 Section and topic Item no. Descriptor Statistical methods 12 Statistical methods used to compare groups for primary outcome(s); methods for additional analyses, such as subgroup analyses and adjusted analyses. Primary outcomesBoth primary and secondary endpoints should be clearly described in the objectives sections of the trial protocol.2 Statistical considerations appropriate to the design of the trial, including sample-size calculations, timelines for any interim analyses and a sketch of a proposed statistical plan for analysing these endpoints, should be detailed in the statistical section of the protocol and reported in subsequent publications.3 The analysis principle for the primary outcome must be that of intention-to-treat, where the data are analysed according to the treatment group to which they were randomised.4 Statistical analysis planKey components of the statistical analysis plan for the primary endpoint or endpoints include: Specifying how the outcome will be measured. Common measures are: Binary (whether or not an event has occurred) — for example, whether or not the subject has experienced a complete or partial response from cancer treatment at 12 months. Typical measures of the event are proportions (risk), rates or odds, and measures of treatment effect include odds ratios and differences in the proportions (or rates) between the intervention and control groups. Count (the frequency of an event in a set time period) — for example, the number of episodes of epilepsy experienced by patients in a 30-day period. A typical unit of measurement would be the rate (count per unit time), and measures of treatment effect include incidence density ratios (similar to odds ratios) or differences between the rates in the groups being compared. Time to event (how long it takes to observe the outcome of interest) — for example, the survival time of patients with advanced breast cancer. Endpoints of this type usually contain censored data (ie, the event of interest has not been observed by the end of the follow-up period), and analyses would involve comparing "averaged" relative risks or hazard/risk ratios (pooled across the time period of the study) between the groups. Measurement on a continuous scale. Examples include blood pressure and temperature measurements, and analyses involve comparing the difference between the means of the intervention and control groups. Other measurements include ordinal scales (eg, quality-of-life ratings, 5-point trauma scales) and non-ordered scales (eg, patient preferences between oral, intravenous or combination treatment delivery). Outcomes measured on these scales require specialised statistical methods. Any transformations on the data likely to be required before analysis. This includes possible groupings or classifications of data (eg, into good, acceptable and poor quality of life), as well as mathematical transformations (logarithms, square root, etc) needed to "normalise" the outcome variables. Typically, these transformations are used if the distribution of the outcome exhibits skewdness, and where, after transformation, this distribution is symmetrical and thus satisfies the assumptions of the statistical method being used to make comparisons.5 If statistical or graphical methods will be used to examine the distribution of the outcome, such as boxplots, histograms and scatterplots,6 these should be detailed. Appropriate statistical tests which will be used to analyse the data. While the underlying assumptions of common statistical tests vary, underpinning all these tests is the assumption that either the outcome (or some transformation of the outcome) or other calculated measures (such as correlation coefficients, hazard or odds ratios) will be "normally" distributed. The normal distribution underlies most statistical inference for most continuous outcomes (and is the basis of χ2, F and t-tests, as well as comparison of odds, hazard and incidence density ratios). If the assumptions of proposed statistical tests do not apply (eg, the data are known to be bimodal), then alternative statistical approaches (eg, classifying the outcome into categories) for analysing such outcomes should be described. How missing data will be accounted for in the analyses (both scientifically and statistically). For example, missing data are sometimes omitted, assigned the baseline value or the group average, or imputed using statistical theory.7 Whether statistical inference be will drawn using one-tailed or two-tailed tests (with appropriate justification) and if any statistical adjustments for multiple comparisons will be performed. In reporting the results of randomised trials, an unadjusted analysis for the primary outcome will provide a consistent, unbiased estimate of underlying treatment differences; this is guaranteed by the randomisation process. This analysis should usually be the primary comparison. However, if the randomisation was stratified, a primary analysis stratified by the stratification factors may be equally appropriate. Subsidiary analyses, which adjust for stratification factors, other potential confounders, or both, can further define the effect of treatment and may provide more efficient statistical comparisons. Parametric tests are based on specific distributional assumptions such as the normal distribution.8 Common misconceptions in analysing clinical data are that a non-parametric analysis (eg, Wilcoxon rank-sum test) is appropriate if the sample size is small (< 30), the data appear skewed (ie, may not be normally distributed) or that the medians are being compared. Whether the distribution of the data departs significantly from the normal distribution may be formally tested; if no departure from normality is indicated, comparisons based on the normal distribution are usually still preferable. Tests based on the assumption of normally distributed data can also be statistically valid for small sample sizes (as low as three per arm). Of course, if there is clear evidence that the data are not normally distributed, the appropriate statistical tests (eg, "exact" tests or non-parametric tests) or appropriate data transformations are required. Finally, even non-parametric tests require some assumptions with respect to the underlying populations from which the samples are drawn.8 If there is a choice of statistical method (ie, assumptions of a parametric test are satisfied), non-parametric methods are generally not as powerful (ie, do not have the same ability to detect a significant difference if it actually exists) as their parametric counterparts. A checklist for a statistical analysis plan is provided in Box 2. 2: Checklist for a statistical analysis plan for clinical trials Provide a detailed description of the primary and secondary endpoints and how they are to be measured. Provide details of the statistical methods and tests that will be used to analyse the endpoints. The analysis of the primary outcome must follow the principle of intention-to-treat. Describe the strategy to be used (eg, alternative statistical procedures) if the distributional or test assumptions are not satisfied. Detail whether comparisons will be one-tailed or two-tailed (with appropriate justification if necessary) and specify the level of significance to be used. Identify whether any adjustment to the significance level or the final P values will be made to account for any planned or unplanned multiple testing or subgroup analyses. Specify potential adjusted analyses with a statement of which covariates or factors will be included. Identify any planned subgroup or subset analysis along with justification for the relevance of this analysis (eg, biological rationale) before commencement of the trial. Specify planned exploratory analyses, justifying their importance. Support claimed differential subgroup effects with biological rationale and supporting evidence from within and outside the study. Provide statistical evidence of interaction between the overall treatment effect and that observed in the subgroup(s) of interest. Remember that prespecified subgroups will have more interpretive value than those defined on an ad-hoc basis or as a result of multiple comparisons. Changing the primary outcome during the conduct of the studyCircumstances can arise where, after a trial commences, the primary outcome is deemed to be suboptimal. This most commonly occurs when the observed rate of the primary outcome is substantially lower than anticipated, reducing the ability (power) of the study to evaluate the effects of treatment on this outcome. This could be the result of a recent change in non-trial background therapy or to recruitment of a more healthy subset of the patients of interest. In these instances, it is possible to modify the primary outcome, provided the reason for so doing is not based on knowledge of interim results of the effect of treatment in the study. Thus, if study data indicate that the rate of myocardial infarction (the primary outcome) is much lower in the intervention or control arm than originally anticipated, it would be highly inappropriate to modify the primary outcome to, for example, include stroke, as this choice is potentially influenced by a knowledge of interim results of the effects of treatment in the study. However, using the overall event rate for myocardial infarction in the whole study cohort (blinded — not differentiated by treatment) could provide justification for endpoint modification in a valid way. Any change in primary outcome during the study requires careful thought, planning and documentation. Secondary outcomesAnalysis of the secondary outcomes needs to be described in the same way as that for the primary outcome, with sufficient documentation in the analysis plan as to how they will be analysed. Where possible, further exploratory analyses should be identified before the completion of the study, with a clear scientific rationale for the reason and value of such analyses. Subgroup analysisIt is essential that potential subgroup analyses are specified before the commencement of a study to guard against data "dredging" or "trawling". Applying many different statistical tests to the same data (eg, on subgroups or different outcomes) has the effect of greatly increasing the chance that at least one of these comparisons will be declared statistically significant even if there is no real difference. This practice is often termed data dredging.9 However, simply specifying a subgroup analysis in advance does not necessarily add scientific legitimacy to the interpretation. A number of strategies exist to ensure the credibility of subgroup analyses, and a checklist proposed by Simes (personal communication) suggests that the following criteria should be satisfied. That there is a biological rationale for considering the subgroup separately from the rest of the patients in the study. Lack of strong biological or clinical evidence for why the treatment should have different effects in a particular subgroup would detract from support of a true underlying differential effect, even if a conventionally significant difference were found. That there is prior evidence or belief that a differential treatment effect in a subgroup is plausible. Lack of prior evidence suggests that differential treatment effects observed in subgroups become hypothesis-generating observations rather than firm conclusions. That there is statistical evidence (ie, a significant interaction) of a difference in the effect of treatment for the subgroup in question compared with the other patients. For example, if there is an apparent advantage of treatment in younger compared with older patients, then careful (clinical and statistical) examination of this difference is required before it can be confidently concluded that a true differential treatment benefit exists in the subgroup of younger patients. Studies are frequently underpowered to detect such interaction effects; nevertheless, lack of statistical evidence of such interaction should prohibit firmly concluding any differential treatment effect in the subgroup. That there is independent confirmation from other factors in the study of the possible differential treatment effect in the subgroup. For example, if, in a trial examining the effect of chemotherapy in gastric cancer, it is observed that women survive longer after an intervention than men, supporting evidence could be to observe that the response rate to treatment was higher and time to disease progression was also longer in women compared with men. Common pitfalls with subgroup analysis are focusing on the size of the P value and of the treatment effect in any subgroup, ignoring the play of chance. Other issues, such as the total number of subgroups examined, also play a major role in determining the credibility of any observed differential subgroup effect. Subgroups defined before initiating the study would be more credible in terms of true differences in effect on the study findings than those determined only at the time of analysis.

Val J Gebski MStat · Anthony C Keech FRACP, MSc(Epid)

Subscribe to MJA email alerts

No spam, you can unsubscribe anytime you want.

By providing your information, you agree to our Terms of Use and our Privacy Policy.

Thanks for Subscribing! Tell us more

Your email updates will use your name.

Good one! Your updates are coming

Thank you for subscribing to the MJA email alerts. Receive the latest content in your inbox.