Topics

Information science

The upsurge of interest in Indigenous health in the 1950s and 1960s. Barry Christophers' letters to the MJA editor about Indigenous health

Barry E Christophers Retired General Practitioner, 1/12 Tollington Avenue, East Malvern, VIC 3145 To the Editor: I write concerning the recent article about my letters to the MJA in the 1950s and 1960s drawing attention to Indigenous health issues.1 Mention is made in the article of the campaign waged by the Federal Council for the Advancement of Aborigines and Torres Strait Islanders concerning the exclusion of Queensland Aboriginal patients with tuberculosis from the generous allowance paid to other TB patients. This campaign was successful. The Tuberculosis Act was amended so that Aboriginal people were not excluded from receiving this allowance. The Australian Medical Association supported this campaign. Without its support it would have failed.

Barry E Christophers

Information science Correction 19 July 2004 Free

The Medical Journal of Australia — prospere, procede et regna

Re: “The Medical Journal of Australia — prospere, procede et regna”, the editorial by Martin B Van Der Weyden in the 1 July 2004 issue of the Journal (Med J Aust 2004; 181: 3-4) cites incorrect references. These should read: 7 Bhasale AL, Miller GC, Reid SE, Britt H. Analysing potential harm in general practice: an incident-monitoring study. Med J Aust 1998; 169: 73-76. 8 Kuhse H, Singer P, Baume P, et al. End-of-life decisions in Australian medical practice. Med J Aust 1997; 166: 191-196. 12 Armstrong R, Van Der Weyden MB. Indigenous health: tell us your story [editorial]. Med J Aust 2004; 180: 492. 14 Stelfox HF, Chua G, O'Rourke K, Detsky AS. Conflict of interest in the debate over calcium channel antagonists. N Engl J Med 1998; 338: 101-106. 18 Laporte RE, Marler E, Akazawa S, et al. The death of the biomedical journal. BMJ 1995; 310: 1387-1390. This error only affected the printed version of the article. The web version was correct when published and has not been changed.

Martin B Van Der Weyden MD, FRACP, FRCPA

Information science Editorials – 90th Anniversary 5 July 2004 Free

The Medical Journal of Australia — prospere, procede et regna

On July 4, exactly 90 years ago, The Medical Journal of Australia began its life “as the official organ of the British Medical Association in Australia”. Its purpose was clear — “to record the progress of scientific medicine, and to assist in rendering the practice of medicine in all its branches of the greatest benefit to the people of Australia”.1 In that first issue, the president of the Victorian branch of the British Medical Association warmly welcomed the Journal, noting that it symbolised “the intimate union of all the branches of the British Medical Association in Australia”, and that it would “continue every week to indicate and advocate the common aims, interests, and ideals of the profession”. He closed by wishing that the Journal prospere, procede et regna2 — “proceed prosperously and reign”! This bridging between the readers and the profession is the stuff of the Journal. The following 90 years have seen the formation of the Australian Medical Association in 1962,3 and, with the advent of the AMA Gazette in 1968, the disappearance of Federal and Branch news in the Journal. In the late 1980s, after nearly 60 years of living the Australian dream of being an owner/occupier, the Journal’s publisher — the Australasian Medical Publishing Company (AMPCo) — sold its Sydney premises to finance the AMA’s move to Canberra. This was the culmination of the Journal’s Sturm und Drang decade, with the destabilising turnover of editors — six in all — and tensions caused by AMPCo’s financial difficulties. However, after 90 years, the Journal’s purpose remains clear — to “be the recognised forum for information and commentary on all aspects of health care in Australia” through “original peer-reviewed clinical research of the highest standard”, “high level continuing medical education”, and “commentary and informed debate on standards of clinical practice, ethics, social, legal and other issues related to health care in Australia”.4 In this 90th anniversary issue, Gregory (page 9) surveys the “clinical research of the highest standard” published by the Journal during this time.5 Its ongoing commitment to “commentary and informed debate on standards of clinical practice” is exemplified by the Quality in Australian Health Care Study6 and the study of adverse events in Australian general practice,7 both of which played a part in the lead-up to establishing the Australian Council for Safety and Quality in Health Care. The Journal’s role as a forum for “ethics, social, legal and other issues” is reflected in our reports on end-of-life decisions,8,9 the health of asylum seekers in detention,10,11 and in our commitment to Indigenous health.12 Finally, pragmatic links to the world of medical research and to specialist and general practice were pursued through networking and the Journal’s Content Review Committee. A cursory review of the Journal’s progress over the past 90 years will readily identify broad changes which have come to pass. There has been a noticeable decline in the number of clinical studies, case reports and the more leisurely reviews, with a concomitant increase in studies of healthcare interventions and health system performance, as well as those on adverse lifestyles, substance misuse, mental illness and, more recently, consumer concerns. With the rise of evidence-based medicine came a barrage of evidence-based guidelines and further delineation of levels of evidence. The design and reporting of research itself adopted more rigorous formats, such as controlled trials, systematic reviews and structured abstracts. Significantly, the number of authors per article continues to multiply,13 and the international trend now is for authorship to involve a team of doctors, other healthcare professionals and scientists. Correspondingly, the number of Journal editors has also increased as the number of submissions continues to rise.13 In 2003 we received a record 917 submissions, compared with 856 in 2001 and 741 in 1999. On the downside, the blurring of the boundaries between commerce and research has spawned a culture of suspicion, particularly for research supported by pharmaceutical companies.14 It is interesting to note that all Journal articles are now accompanied by an item noticeably absent a decade ago — the competing interests statement. The Journal’s policy of safeguarding the integrity of research by exploring potential conflicts of interest of contributors and reviewers is detailed by Chew.15 What does the future hold? Just as Gutenberg’s printing press saw the demise of the monastic monopoly of manuscript production, electronic technology has changed both the essence of publishing itself, and ease of access to the latest research. The Medical Journal of Australia, like most other medical journals, simultaneously releases the electronic (eMJA) with the print Journal, and uses rapid online publication for selected articles. Future electronic developments are also anticipated. There are those who promote the notion that peer review and editing are things of the past.16 They believe science should simply be posted on the Internet, thus letting the world judge its quality. However, an editor’s first responsibility is to the readers, and they have signalled that they are too busy to separate the wheat from the chaff.17 They prefer that to be the function of quality filters — the editors, peer reviewers and editorial staff who ensure the clarity, brevity and non-exclusive language of the final product. This bridging between the readers and the profession is the stuff of the Journal. Despite enthusiastic predictions of its demise,18 the printed Journal will live on for some time. There is something reassuring about knowing where a journal’s contents will be revealed, its portability from bed to breakfast table, and the feel of something physical, that binds readers to the paper Journal.19 In any event, whatever changes the future may bring, as long as the Journal continues to add value to its core content of original articles, editorials, reviews and informed debate on contemporaneous healthcare issues in Australia, The Medical Journal of Australia will most certainly prospere, procede et regna.

Martin B Van Der Weyden MD, FRACP, FRCPA

Information science Editorials – 90th Anniversary 5 July 2004 Free

What conflict of interest?

“If in doubt, disclose” — Committee on Publication Ethics1 Under full glare of the media spotlight in February this year, the editor of The Lancet, Richard Horton, partially retracted an article published 6 years before.2,3 This act was triggered by allegations of research misconduct in the study, taken to The Lancet by a Sunday Times journalist.4 The Lancet’s investigations5 and the events that ensued involved the article’s authors (doctors at London’s Royal Free and University College Medical School), the institution’s ethics committee, the General Medical Council, the British Parliament, and even the Prime Minister, Tony Blair. Why the fuss? In 1998, this research report implied a link between autism, bowel disease and the combined measles– mumps–rubella (MMR) vaccine.3 A storm of controversy erupted, and MMR vaccination rates in England fell from 92% of children reaching the age of two in 1996/97 to 82% in 2002/03.6 Confirmed cases of measles rose from 112 in 1996 to 442 in 2003. When the allegations of misconduct were probed this year, it was found that, at the time the article was submitted and published, its lead author and senior investigator, Dr Andrew Wakefield, had not disclosed that he had been commissioned by the Legal Aid Board (for the sum of £55,000) to determine if there was evidence to support legal action by parents of children allegedly harmed by the MMR vaccine. Some of these children were also in Wakefield’s Lancet study. This led to a partial retraction of the article by 10 of its original 13 authors, but not by Wakefield.7 Richard Horton remarked, “If we had known the conflict of interest Dr Wakefield had in this work I think that would have strongly affected the peer reviewers about [its] credibility . . . in my judgement, it would have been rejected”.2 The Medical Journal of Australia has not had such sensational experiences (yet!). Conflict of interest in publication was first raised in a position statement by the International Committee of Medical Journal Editors, published in 1993.8 Our first conflict of interest statement appeared in the late 1990s, but declaring conflict of interest was deemed by many then to be optional — an exercise in political correctness. So, except in extreme situations, does conflict of interest really matter? Unfortunately it does. It is a principle long enshrined in the conduct of scientific research.9 Moreover, we can no longer ignore the growing body of evidence that conflict of interest can bias research outcomes. Systematic reviews have shown that results of sponsored studies are more likely to favour the sponsor when it is a pharmaceutical company,10 and industry-sponsored studies are not only associated with pro-industry conclusions, but with restrictions on publication and data sharing.11 As more research is funded by non-academic sources, we cannot ignore the potential for bias to affect its outcomes.12 Neither can we ignore an increasingly educated, involved 21st-century Western society that is calling for greater accountability in health research. What conflict of interest is notConflict of interest is not always present, and may be potential rather than actual. Such dual interests are better termed “competing” rather than “conflicting” interests (eg, commitment to a patient’s welfare and to a research project). Having a conflict of interest is, in itself, not wrong, and may be unavoidable. Disclosing competing interests should not be seen as an admission of wrongdoing, but as promoting transparency in the public record. What conflict of interest is“Financial or personal relationships that inappropriately influence (bias) . . . actions [of an author (or the author’s institution), reviewer, or editor] . . . The potential for conflict of interest can exist whether or not an individual believes that the relationship affects his or her scientific judgment”.13 Anything, be it personal, financial, academic, religious, or political, “which, when revealed later, would make a reasonable reader feel misled or deceived”.1 Current MJA policyAuthorsAll authors are required to provide a disclosure statement (Box). Funding sources are to be acknowledged, together with any role they played in study design, data collection, data analysis, interpretation of the data, their reporting and publication. A study may not be published if its sponsor asserts the right to control publication.13 We publish competing interests statements for all research, viewpoint and review articles, and, where such interests are declared to exist, for editorials and letters to the Editor. ReviewersPeer review is a useful but imperfect tool that we can refine by seeking to clarify potential biases: We do not ask those from the same institution(s) as the author(s) to review an article. We now ask all reviewers to provide a disclosure statement similar to the one authors provide (Box). Reviewers who disclose competing interests are not necessarily disqualified. Their reviews will be carefully considered by the editors, bearing their potential biases in mind, and in conjunction with comments from other reviewers. EditorsEditors commission articles, assess submitted articles, and ultimately decide their fate. Thus, MJA editors with competing interests relating to an article will exclude themselves from taking primary responsibility for it. ConclusionOur aim is not to exclude anyone with a potential conflict of interest from publishing or reviewing — to do so would disqualify virtually everyone (including editors). Conflicts of interest may occasionally be too extreme to allow publication of the article or involvement of someone in the decision-making process. However, our ultimate goal in advocating disclosure is to promote transparency, reduce bias, and maintain public trust in what we publish. Let the reader to be the judge! MJA disclosure statement for authors and reviewers (A) Authors are asked to indicate Yes or No to questions about affiliations with manufacturers of products mentioned in the article or of competing products: Ownership of stock or stock options or other financial instruments of companies whose products are mentioned in the article or who manufacture competing products (does not include mutual fund ownership) Ongoing paid consultancy with company or a competitor (actual or within the last 2 years) Employment with company or competitor (actual or within the last 2 years) Honorarium or other compensation for writing the article or for participating in the development of the article Honorarium or other compensation for conducting research related to material contained in the article Speaker fees and/or educational grants Travel assistance to attend meetings. (B) Authors are asked to provide details where the answer to any of the above questions is “yes”. (C) Authors are asked to declare any other (non-financial) competing interests.

Mabel Chew MB BS(Hons), FRACGP, FAChPM

Information science Editorials – 90th Anniversary 5 July 2004 Free

Ninety years young — the changing covers of the MJA

1914 2004 1956 1978 1982 1989 1993 For more covers, see the pdf version of this article. As the eyes are said to be the window to a person’s soul, so is a journal’s cover a window into the ethos of its editors and readers. Over the past 90 years, the cover of The Medical Journal of Australia has changed many times, reflecting the national and international events of the times, the perceived desires of its readers and, occasionally, the whims of its editors. The Journal’s first issue was published on 4 July 1914. The MJA arose from the amalgamation of the Australasian Medical Gazette (published by the NSW Branch of the British Medical Association since 1881) and the Australian Medical Journal (published by the Victorian Branch of the BMA since 1856). This union was not without some opposition, but there was a clear need for a national journal that would unite the six Australian branches of the BMA and reach the whole medical profession in Australia. The cover changed little in the first 40 or so years, featuring only the title, issue details and a large monochrome advertisement. This was a time of war (World War I, with Britain declaring war against Germany exactly a month after the first issue of the fledgling journal appeared), financial difficulties as the cost of paper and printing suddenly rose, the Great Depression, another war (World War II, when many of the Journal’s contributors were in the armed forces), and editors in for the long haul (Henry William Armit [featured on this issue’s cover] served 16 years [1914–1930], Mervyn Archdall served 27 years [1930–1957], and Ron Winton, whose obituary is published on page 26 of this issue, served 20 years [1957–1977]). In the post-war years, medicine experienced a technological explosion and the Journal took on a more modern layout with a bright blue cover, better-quality paper and some colour printing. The MJA’s content adopted a more global perspective, reflecting Editor Ron Winton’s Chairmanship of the Council of the World Medical Association for many years. The MJA also officially became the journal of the Australian Medical Association when the shackles of the BMA were cast off in 1962 and the Association became fully independent. At this time, the number of specialist journals increased greatly, with some fragmentation of the readership of the MJA. Winton countered this with an editorial policy of a mix in each issue to appeal to both generalists and specialists, a policy which continues today in recognition that our readership is almost equally divided between specialists and general practitioners. After Ron Winton, the Journal went through a turbulent time, with six editors in 10 years. These were also lean years, when the Journal changed from a weekly to a fortnightly publication (1978), and The Printing House in Glebe, Sydney, closed (after publishing and printing the Journal completely “in-house” for more than 60 years). During the late ’70s and early ’80s, the covers changed almost as often as the editors. Of particular note were Alan Blum (1982–1983), who came to the Journal from the United States and was responsible for some of its most controversial covers on topics such as smoking, nuclear war and AIDS, and Alistair Brass (1983–1985), described as a liberal thinker and an outspoken critic, who put Australian artworks on the cover. The Journal subsequently returned to its more conservative roots in the late ’80s under Kathleen King with, for the first time, the full contents displayed on a bright green cover. This was initially a space-saving device, but it proved popular with readers. In the early ’90s, joint editors Laurel Thomas and Jill Forrest continued to appease the scientific purists with the full contents on the cover, but softened it with a picture and a less stark grey–blue background. Martin Van Der Weyden took the helm in 1995 and the cover slowly evolved, with only minor changes in colour and typeface. With this issue, however, we present a major overhaul. We’ve aimed for cleaner lines, a less cluttered look and a Journal that’s generally easier on the eye. Our ideas are perhaps most succinctly stated by Joseph Pulitzer: “Put it before them briefly so they will read it, clearly so they will appreciate it, picturesquely so they will remember it and, above all, accurately so they will be guided by its light.” We hope you approve.

Bronwyn Gaut

History and humanities The Research Enterprise – 90th Anniversary 5 July 2004 Free

Jewels in the crown: The Medical Journal of Australia’s 10 most-cited articles

According to data from the Institute for Scientific Information (ISI), the most-cited MJA article is Cade’s ground-breaking report on the effect of lithium in mania (1949; 888 citations), followed by Marshall et al’s reports on the role of Helicobacter pylori in gastroduodenal disease (1985; 766 and 523 citations, respectively). Others in the “top 10” span decades and disciplines; all have a common grounding in Australian data of global relevance. For the year 1995, shortly after the 80th anniversary of The Medical Journal of Australia (MJA), researchers from the Australian National University used citation analysis to determine Australia’s contribution to new knowledge in medical and health sciences. They found that Australians contributed 2.5% of all publications in the Science Citation Index — 18 390 publications, which had been cited over 88 000 times.1,2 In 2003, the MJA used citation data provided by Thomson ISI (www.isinet.com) to identify the “top 10” articles published in the MJA — that is, the articles which had been cited most often. Data comprised citations within journal articles covered by the database of the Institute for Scientific Information (ISI) for the years 1945–2002; thus, articles published before 1945 could be cited. The MJA top 10 articles span more than 60 years (Box 1). They have in common a grounding in Australian data, but a global relevance. In addition, all provide evidence of the importance of basic as well as clinical research, and of that endangered species the physician–scientist. 1 The MJA’s top 10 articles, by citation analysis Number 1 (888 citations) Cade JFJ. Lithium salts in the treatment of psychotic excitement. Med J Aust 1949; 2: 349-352. Number 2 (766 citations) Marshall BJ, Armstrong JA, McGechie DB, Glancy RJ. Attempt to fulfil Koch’s postulates for pyloric campylobacter. Med J Aust 1985; 142: 436-439. Number 3 (523 citations) Marshall BJ, McGechie DB, Rogers PA, Glancy RJ. Pyloric campylobacter infection and gastroduodenal disease. Med J Aust 1985; 142: 439-444. Number 4 (299 citations) Derrick EH. “Q” fever, a new fever entity: clinical features, diagnosis and laboratory investigation. Med J Aust 1937; 2: 281-299. Number 5 (267 citations) Swan C, Tostevin AL, Moore B, Mayo H, Barham Black GH. Congenital defects in infants following infectious diseases during pregnancy. Med J Aust 1943; 2: 201-210. Number 6 (203 citations) George LL, Borody TJ, Andrews P, Devine M, Moore-Jones D, Walton M, Brandl S. Cure of duodenal ulcer after eradication of Helicobacter pylori. Med J Aust 1990; 153: 145-149. Number 7 (170 citations) Trautner EM, Morris R, Noack CH, Gershon S. The excretion and retention of ingested lithium and its effect on the ionic balance of man. Med J Aust 1955; 2: 280-291. Number 8 (169 citations) Bower C, Stanley FJ. Dietary folate as a risk factor for neural-tube defects: evidence from a case–control study in Western Australia. Med J Aust 1989; 150: 613-619. Number 9 (167 citations) Wilson RMcL, Runciman WB, Gibberd RW, Harrison BT, Newby L, Hamilton JD. The Quality in Australian Health Care Study. Med J Aust 1995; 163: 458-471. Number 10 (166 citations) Borody TJ, Cole P, Noonan S, Morgan A, Lenne J, Hyland L, Brandl S, Borody EG, George LL. Recurrence of duodenal ulcer and Campylobacter pylori infection after eradication. Med J Aust 1989; 151: 431-435. Simple cation holds promise as psychotropic agentJohn Cade (1912–1980), the author of our most-cited article, once described himself self-deprecatingly as “an unknown psychiatrist, working alone in a small chronic [sic] hospital with no research training, primitive techniques and negligible equipment”.3 Born in Murtoa, a small country town in Victoria, Cade seemed destined to enter psychiatry. His father was a psychiatrist, and as a child Cade lived in the grounds of various “lunatic asylums”. He entered psychiatry in 1936, shortly after graduating in medicine (with honours in all subjects), but spent much of the Second World War as a prisoner of war in Changi, Singapore, returning to Australia as a 40 kg “walking skeleton”.4,5 Cade’s interests included all the sciences, and his “enquiring mind” stayed with him throughout life. The coauthor of his first article (published in 1940), detailing the serological response to influenza virus infection, was none other than Frank Macfarlane Burnet.6 While Cade was investigating potential anticonvulsant agents in guinea-pigs, he came to suspect that the cation lithium had a sedative effect which might be useful in treating mania. He demonstrated this sedative effect in guinea-pigs, and then took lithium himself, before extending his study to patients.3 In his MJA article, Cade reported the Results of a study of the effect of lithium salts in 10 patients with mania (as well as six with schizophrenia and three with “melancholia”). Lithium had a clear effect in mania. He published no further research on lithium, but did search for other cations with psychotropic activity.3 In commenting on his research career, he said: “My own research efforts have been sporadic over many years. Most have ended in blind alleys. Some have been successful. All have been fun. In the process I have learned a greater deal . . . , and en passant something of the causes and effective treatment of manic–depressive illness.”7 Cade’s findings were not immediately accepted in the rest of the world (and not until the 1970s in the United States), so it is not surprising that further notable research on lithium was also conducted in Australia. Ranked seventh in the MJA top 10, The excretion and retention of ingested lithium and its effect on the ionic balance of man was published in 1955. The authors included Trautner, a physiologist at the University of Melbourne, and Noack, a psychiatrist at Melbourne’s Mont Park Hospital. Their research, conducted on themselves and on patients with mania, showed that lithium is retained during the acute phase of mania, necessitating higher doses. These can be reduced as the mania resolves. They also showed that intercurrent illness increases the risk of lithium toxicity. Spiral bacterium linked with gastritis and peptic ulcerToday, we know that Helicobacter pylori colonises the stomach and infects about half the world’s population.8 Further, it has infected people since the dawn of human history, and its geographic variation is being used to map the earliest human migrations, including the arrival of Europe’s first neolithic farmers.9 However, as recently as two decades ago, notwithstanding reports suggesting otherwise (“dispersed over 100 years, in journals of different languages and subspecialities”10), the prevailing dogma was that the human stomach was sterile, and that bacteria could not survive in gastric acid.10 Gastroenterologist Barry Marshall and pathologist Robin Warren first described the association of a campylobacter-like organism with gastritis in two letters to the editor of The Lancet.11,12 Marshall and colleagues later published two articles — in the 15 April 1985 issue of the MJA — providing evidence of a causal relationship. These articles rank second and third in the MJA top 10. 2 Illustrations from Marshall et al’s two “top 10” MJA articles A. Numerous Helicobacter pylori organisms in a gastric biopsy specimen from Barry Marshall, taken 10 days after he ingested a pure culture of the organism (Warthin–Starry silver stain; original magnification x 900). B. Heavy growth of H. pylori from an antral biopsy specimen from a patient with duodenal ulcer. (Larger white colonies are commensal flora of the mouth.) In the first MJA article, the researchers successfully fulfilled Koch’s third postulate by demonstrating that H. pylori (then known as pyloric campylobacter) could colonise histologically normal mucosa. After trying to infect animal models without success, Marshall used himself as a “guinea-pig”. About 5 days after drinking a pure culture of H. pylori (109 organisms), he became ill, with early-morning nausea, vomiting of acid-free gastric juice, and “putrid” breath. Although the illness resolved spontaneously after 14 days, culture and histological examination on the 10th day showed severe acute gastritis with many H. pylori organisms (Box 2A). The experiment allowed Marshall and colleagues to link H. pylori to epidemic gastritis with hypochlorhydria. In the second MJA article, Marshall and colleagues proposed that pyloric campylobacter infection was responsible for damage to the duodenal epithelium, as well as the gastric antral mucosa, based on gastroduodenal biopsy and culture findings from over 100 patients referred to their dyspepsia research clinic (Box 2B). Looking back on these discoveries, Marshall later wrote that early reports of an association between peptic ulcer and H. pylori were met with extreme scepticism by many doctors, who were convinced that psychic stress, cigarette smoking and hyperacidity were the causes of peptic ulcer. Reports of the first therapy ever shown to heal gastritis received “a cool reception at gastroenterological meetings”.10 Compared with Cade’s era, communication among the world’s scientific community had accelerated greatly, and, in 1991, the first convincing study of cure of duodenal ulcer through eradication of H. pylori was published in the United States. However, this was preceded by another pair of notable articles on the same topic in the MJA, from the Centre for Digestive Diseases in Sydney. In 1989, the study by Borody and colleagues, which ranks tenth in the MJA top 10, showed that “triple chemotherapy” with bismuth, tetracyline and metronidazole could lead to long-term eradication of H. pylori in most patients with duodenal ulcer or non-ulcer dyspepsia. Further, they suggested that this eradication could reduce recurrence of, or even cure, duodenal ulcer. The group subsequently reported such cure in their 1990 MJA article, which ranks sixth in the top 10. H. pylori infection is now recognised as the major cause of peptic ulcer disease and an important risk factor for gastric malignancy. For discovering its role in peptic ulcer disease, Marshall was awarded the 1995 Albert Lasker Clinical Research Award.13 Mystery abattoir fever confirmed as new disease 3 Edward Derrick, who first described Q fever In 1961, Derrick became director of the Queensland Institute of Medical Research. (Illustration courtesy of the Brisbane Courier-Mail.) In 1935, unexplained fevers in abattoir workers in Queensland were referred for investigation to Edward Holbrook Derrick, newly appointed director of the state’s Laboratory of Microbiology and Pathology14 (Box 3). He was unable to identify a cause but found that “abattoir’s fever” had a distinctive natural history. The clinical resemblance to murine typhus led him to inoculate patients’ blood into guinea-pigs, which became febrile. The agent could be transmitted serially from one to another, and, after recovery, the guinea-pigs remained resistant to infection. These findings were reported in the MJA in 1937, in an article that ranks fourth in the top 10. Thirty years later, Macfarlane Burnet wrote that “these findings provided a rather cumbersome, but perfectly adequate means of establishing that abbatoir’s fever was a specific entity definable immunologically, and also of allowing laboratory diagnosis in a doubtful clinical case”.15 Although Derrick described the disease and named it “Q” fever, it was Macfarlane Burnet who showed it was caused by a rickettsial agent, as described in his article, coauthored with Mavis Freeman, which followed Derrick’s in the same issue of the MJA.16 The causative organism is now known as Coxiella burnetii. Derrick (1898–1976) received international recognition for discovering not only Q fever, but also the form of leptospirosis caused by Leptospira pomona.17 When Derrick was made a Fellow of the Australian Postgraduate Federation, Macfarlane Burnet stated that “to have defined and elucidated the aetiology of two worldwide infectious diseases is something no other living scientist can claim.”17 German measles in pregnancy may damage the fetusIn 1941, the teratogenic effects of rubella (German measles) were uncovered by the Australian ophthalmologist Norman Gregg.18 At that time, it was generally believed that birth defects were inherited, and that the placenta was an absolute barrier to infectious diseases. Gregg’s suggestion that maternal rubella played a causal role in congenital cataract was considered revolutionary, and several years passed before overseas medical journals commented on the idea.19 However, in Australia only a year later, Charles Spencer Swan was appointed by the National Health and Medical Research Council to investigate the possible relationship.19 On 7 October 1942, a circular sent to all South Australian general practitioners informed them of Gregg’s findings and asked them to complete a form for all children born to women who had an acute exanthem during pregnancy. From these data, covering the years 1939–1943, Swan and colleagues identified 49 infants whose mothers had been exposed to rubella during pregnancy; 31 had congenital malformations, including cataract, deaf-mutism, heart disease, microcephaly and mental retardation. In all but two of the 31 cases, rubella had been contracted in the first 3 months of pregnancy. Further, Swan and colleagues suggested that the type of congenital malformation depends on the stage of pregnancy at which the mother acquired rubella. Their MJA report ranks fifth in the top 10. The rubella virus itself was not identified for about another 20 years. Although Gregg’s landmark article on congenital cataract and maternal rubella was formally published in the Transactions of the Ophthalmological Society of Australia,18 he had presented his observations at the annual meeting of the society in October 1941. A description of the proceedings was published with permission in the MJA in December 1941,20 before the formal article appeared. In defence of the rapid publication, the MJA stated: “The series [of cases] is so striking and the sight of the children is so seriously affected that the facts must be made known without undue delay to the general body of the medical profession.”20 Folic acid in pregnancy can prevent spina bifidaFiona Stanley graduated in medicine from the University of Western Australia and trained in epidemiology at the London School of Hygiene and Tropical Medicine and the National Institutes of Health in the United States. In 1977, she returned to Perth for family reasons and, although a researcher at heart, became Senior Medical Officer in Child Health.21 Yet, this chance worked in both her and our favour, as it allowed her to establish, with colleagues, the Western Australian Congenital Malformations Registry. The registry provided the data for her landmark MJA article, coauthored with Carol Bower, which showed that dietary intake of folate in early pregnancy protects against the occurrence of isolated neural-tube defects in infants. It ranks eighth in the MJA top 10. 4 Fiona Stanley Since her discovery of the role of folate in preventing neural-tube defects, Stanley continues to investigate the epidemiology of childhood and maternal illness. In 1990, a year after her top 10 MJA article was published, Stanley became founding director of the Telethon Institute for Child Health Research in Perth. She continues to explore the promise of epidemiology and other scientific disciplines in tracking trends and preventing major childhood and maternal illnesses.22 Stanley (Box 4) was Australian of the Year in 2003. Healthcare can harm patientsThe Quality in Australian Health Care Study (QAHCS) arose from the Tito Review of Professional Indemnity Arrangements for Health Care Professionals, established by the Australian government in 1991. The review was to examine the adequacy of compensation and funding arrangements for healthcare misadventures in Australia, but lacked the data to answer the fundamental questions: How many adverse patient outcomes arise from healthcare services? How severe are they? What impact do they have on those services? A consortium of the University of Newcastle, the University of Adelaide and Sydney’s Royal North Shore Hospital was awarded the contract to provide these data, led by intensive care physician Ross Wilson. The QAHCS, based on the Harvard Medical Practice Study, was set up to measure preventability rather than negligence. Nevertheless, it provided a national measurement of the safety of healthcare, a measurement many other countries still lack. The most-cited report from the QAHCS was published in the MJA in 1995 and ranks ninth in the top 10. It found that 16.6% of hospital admissions in Australia in 1992 were associated with an “adverse event” to patients, that those events meant patients were injured by their healthcare, and that the injury had caused them some disability. About half the adverse events were considered preventable. In 1999, also in the MJA, the consortium reported further on the preventability of these events.23 “The spirit of the researcher”The MJA’s 10 most-cited articles are testimony to the power of clinical research to revise our understanding of disease and treatment methods, and to enhance prevention of disease and adverse events. Despite the refinements in clinical research methods over the decades, which will continue to evolve, these top 10 articles and the pioneering spirit of their authors should inspire new generations of doctors to make the most of any opportunities or insights that come their way. Derrick, in his address to the inaugural meeting of the Queensland Branch of the Australian Society for Medical Research in 1969, quoted the American physiologist Walter Cannon: Phenomena, no matter how mysterious they may appear to be, have a natural explanation and will yield their secrets to the persistent, ingenious, and cautious efforts of the investigator.24 Cade, in his presidential address to the Seventh Annual Congress of the Australian and New Zealand College of Psychiatrists in 1970, said: Almost everyone can and should do research, both because almost everyone has a unique observational opportunity at some time . . . and also because the intellectual discipline and technical training that it imposes is an essential prerequisite to expertise in a professional field.7 Derrick was said to have had a feeling for the historical context in which his research was done, an awareness of the stepwise progress of knowledge to which all, “however ill-equipped”, might hope to add. He was said to be fond of quoting the wisdom of Descartes: The last should commence where the preceding had left off, and thus by joining together the lives and labours of many, we should collectively proceed much further than anyone in particular would succeed in doing.17

Ann T Gregory MB BS, GradDipPopHealth

Information science Obituary 5 July 2004 Free

Ronald Richmond WintonOAM, MB BS, FRACP, FRACMA

Ronald Winton was Editor of The Medical Journal of Australia from 1957 to 1977. He would be the first to agree that his life had been a good one. Departing it, on 13 February 2004, at the age of 90, he would have thanked God for its variety, its achievements, and the love he received from his family, friends, colleagues and staff, as well as the many students for whom he became a surrogate father. Ron was born in 1913 in Campbelltown, New South Wales. His medical career began at Brisbane Hospital in 1935 and continued during Army service. He described his war as “quiet”, but managed to have some adventures in Palestine and New Guinea, retiring with the rank of Lieutenant-Colonel. Ron came to the Journal in 1947 as assistant to Mervyn Archdall, succeeding him as Editor in 1957. In those days, the Journal was wholly produced within The Printing House, at Glebe, from subedited manuscript through typesetting in hot metal to printing and binding. Ron’s term saw many changes — in medicine itself, in publishing procedures, in practice management and in medical insurance. He also saw the transition from the branches of the British Medical Association in Australia to the independent national and state bodies of the Australian Medical Association (AMA), and the controversial negotiations with government during the early days of Medibank. His editorials on medicopolitical topics were always insightful and balanced. Ron left his stamp on medical journalism in Australia in many ways. He edited the Australasian Annals of Medicine for some years and helped to formulate the Code of Conduct for pharmaceutical advertising for the National Medical Media Council. Through all the changes he steered the MJA, the “flagship” of the AMA, enabling it to keep honourable company with the great journals of the world. On his retirement he was awarded the Gold Medal of the AMA. Outside work, his passions included music, literature, theology and history (he lectured in the history of medicine at the University of Sydney). These topics often flavoured staff morning teas, to everyone’s great delight. On the world scene, Ron was a member of the World Medical Association, chairing its Committee on Medical Ethics, as well as the International Congress of Christian Physicians. His love affair with books led him to write several of his own.1-12 The great guiding force in his life was his deep Christian faith, exemplified in all his activities, especially his 20 years as Honorary Warden of “Wingham”, an Anglican hostel for country and overseas students. The Journal had been fortunate in its editors. Henry William Armit established its tradition of scrupulous accuracy, Mervyn Archdall brought wit and flair, and Ron Winton injected his own brand of wisdom and grace. As an editor, he was a hard act to follow. As a man, he inspired loyalty and love.

Laurel Thomas

Indigenous health: tell us your story

Announcing the Dr Ross Ingram Memorial Essay Competition (entry details below) Not so long ago, we at The Medical Journal of Australia realised that, when it came to Indigenous health, we were great at publicising the problems. Most of the articles we publish are observational studies confirming that, yes, in health, as well as in almost every other area, Indigenous Australians are worse off than other Australians and, indeed, Indigenous populations worldwide. Ross Ingram (16 Feb 1967 – 15 May 2003) Ross Ingram was an Indigenous doctor who died last year, aged 36, of cardiovascular disease. At the time of his sudden death he was working as a GP in the New South Wales rural town of Leeton. Ross grew up in the Leeton area, where he was educated at the local primary and high schools. In 1984 he was named Young Citizen of the Year for Leeton, and in 1985, while vice-captain of Leeton High School, he received a Rotary Citizenship Award. In 1987 he was awarded a National Aboriginal Islander Day Observance Committee (NAIDOC) Award for Aboriginal Youth of the Year. Ross was the first Indigenous person from NSW to be accepted into the University of Newcastle’s Medical School. He enrolled in 1986 and graduated in 1993, the first Wiradjuri person to become a doctor. Life and medicine took him to an internship and residency in Gosford, then general practice on the NSW central coast and in Tasmania, and finally back to practise in Wiradjuri Country (central western New South Wales). His death is the first among the small community of Indigenous doctors who have been graduating from Australian medical schools since 1984. A keen practitioner of softball, football and cricket, as well as medicine, Ross was proud of his achievements both as a man and an Indigenous man. He is remembered by a loving family, including his wife, Julie, three children and three stepchildren. We also realised that the Journal was missing an important “voice”, telling us the story of Indigenous health. Many of the people working in Indigenous healthcare do not publish in academic journals. Also, more than in some other sectors of the population, social, cultural, political and economic issues influence the health and wholeness of Indigenous people. Some of these factors cannot be explored in strict academic style. Essays, on the other hand, leave room for the writer to analyse and interpret, often from a personal perspective and possibly including some form of narrative — “telling a story”. With this in mind, we are delighted to announce the annual Dr Ross Ingram Memorial Essay Competition for the best essay relating to Indigenous health. The competition is open to any Indigenous person who is working, researching or training in a health-related field; we are looking for essays that present original and positive ideas aimed at promoting health gains and health equity for Australia’s Indigenous peoples. After all, real insights and solutions come from within, not from without. The essays should be no more than 2000 words long, and must be submitted by Monday, 10 January 2005. A panel, including external experts and MJA editorial staff, will judge finalist essays, and judges will be blinded to the identities of the authors. The judges’ decision will be final. The winning entry will be published in the 2005 Indigenous Health issue of the Journal (the second issue in May), and the author will receive $5000. Other essays of high merit may also be published. We asked the members of the Australian Indigenous Doctors’ Association (AIDA) to help us name the prize and they chose to name it after Dr Ross Ingram (see Box). Ross’s story of premature death from natural causes is not an unusual one. More than half the deaths in Indigenous men occur before they reach the age of 50, compared with 13% of deaths among non-Indigenous men. The members of AIDA chose Ross not just because he was the first known Indigenous doctor to die, but because his plight typified that of many of the people currently working in Indigenous health. The human reality of statistics like those mentioned above is that Indigenous Australians inhabit a world of sickness, death and tragedy. Many of the seeds of future ill health are present from before birth. To a greater extent than most of their non-Indigenous colleagues, Indigenous doctors risk becoming a part of the problem they are trying to treat. “As Indigenous doctors, the fraternity of medicine has always accepted us wholly, and without question, and yet we are very different from so many of our non-Indigenous colleagues. Many doctors, when they look into the eyes of an Indigenous child, get a glimpse of a world they never knew existed; when we look into the eyes of that child, we see ourselves, and are reminded of the toll taken by unending stress and anxiety, and cycles of grief. For Indigenous doctors, the loss of our dear brother Ross reminds us that the privilege we enjoy as doctors does not remove our responsibilities to our people.” — Louis Peachey, President, AIDA We are hoping that the Dr Ross Ingram Memorial Essay Competition will provide a forum for some of the stories and ideas of Indigenous people working in Indigenous healthcare. Ross Ingram will not be able to contribute in this way, but he is a silent reminder of both the problem and the struggle of those who are working to find a solution. We look forward to receiving your entries.

Ruth M Armstrong BMed · Martin B Van Der Weyden MD, FRACP, FRCPA

Information science EBM: Trials on trial 3 May 2004 Free

Subgroup analysis: application to individual patient decisions

Clinical trials provide evidence of effectiveness of treatments as an average for a group of patients, yet, in clinical medicine, we usually wish to apply these results to individuals. Can we simply apply the overall trial result for each patient, or can the result be tailored to individual patients in some way? Consider a hypothetical example: a randomised trial comparing treatments A and B shows that treatment A is more effective than B among men (P < 0.001), but not among women (not signficant). Does this mean men should receive the new treatment, but women should not? Box 1 illustrates these results from three different studies. In Study 1, the estimated treatment effect in men and women is the same — a 25% reduction in mortality associated with treatment A — but the much smaller number of women in the study gives rise to wider confidence intervals for this subgroup. In this case, there is no basis to consider that the treatment is any less effective in women (no heterogeneity; ie, non-significant test for interaction). Treatment could be considered effective for any patient regardless of sex. In Study 2, the treatment effect in women is less than that in men (8% v 25% relative reduction, or 0.92 v 0.75 relative risk, respectively), but the effects in both are still consistent with the overall result of a 20% relative reduction, and there is no evidence of significant heterogeneity between groups (test for interaction, P = 0.20). Here, the different results between men and women could be simply due to chance, and it would still be appropriate to apply the overall estimate to both men and women (unless there was additional evidence).1 In Study 3, the observed effects for men and women are sufficiently different to suggest that this difference is unlikely to be due to chance (test for interaction, P = 0.01), and it is reasonable to conclude that the treatment effect differs between men and women. In this instance, the trial evidence should be considered separately for these subgroups. However, even here, a test of interaction can still give a low P value simply on the play of chance if many subgroups have been evaluated.2 A practical approachConsider applying the overall trial treatment effect to each subgroupHow should we decide in practice whether to consider the treatment effects for these subgroups separately? A practical approach is shown in Box 2. A controlled trial is usually designed with a sample size large enough to show an overall treatment effect, but not necessarily adequate to show significant effects in each subgroup separately. The overall treatment effect is considered the best estimate for each subgroup of patients in the trial (this is sometimes referred to as the effect domination principle1). Hence, the treatment-effect results should only be applied differently for different subgroups if there is evidence of heterogeneity (a significant difference between subgroups, sometimes called interaction or treatment-effect modification).3-7 In considering whether there is evidence of heterogeneity, it is also worth reviewing the other questions outlined in the checklist for subgroup analyses in the previous article in this series.1 If there is clear and reliable evidence of heterogeneity, using the treatment effects for each subgroup may be appropriate. However, as these may be unreliable (based on smaller numbers of patients) they may still be considered exploratory and motivate further trials rather than necessarily leading to different treatment guidelines. Further, evidence of heterogeneity may also lead to a search for underlying factors which may be linked to the particular subgroup and provide a more plausible biological explanation for such variation. For example, an apparent difference in treatment effect between men and women may truly relate to differences between these groups in age or smoking status (so-called confounding). Seek confirmatory evidenceIf there is still uncertainty whether differences between the subgroups in treatment effect are real, the following steps should be taken: seek confirmation from the results of an independent trial, a meta-analysis, or both; determine whether the effect is also present for a composite (expanded) endpoint, or surrogate endpoints; and establish whether independent evidence exists of a-priori biological plausibility of differences in the treatment effect. Without a good a-priori rationale for subgroup differences, the overall treatment effect provides a reasonable estimate for each subgroup, unless confirmatory evidence of treatment differences becomes available. Estimate treatment effect according to baseline riskIf the same (or similar) relative treatment effect applies to different subgroups of patients, then those with a greater baseline risk of an event will derive a larger treatment effect. The absolute risk reduction associated with treatment is simply the absolute baseline risk multiplied by the relative risk reduction (Box 2).8 For example, for a patient group with a 20% baseline risk, a treatment with a relative risk reduction of 25% would translate into a 20 × 0.25 = 5% reduction in absolute risk. For a patient group with a 10% baseline risk, this would translate into a 10 × 0.25 = 2.5% absolute risk reduction. The number of patients needed to treat (NNT) to avoid one event can be calculated as one divided by the absolute risk reduction (Box 2).8 This corresponds to 1 ÷ 0.05 = 20 NNT for a patient with a baseline risk of 20%, and 1 ÷ 0.025 = 40 NNT for a patient with a baseline risk of 10%. Smaller numbers needed to treat will result for patients at higher baseline risk. A practical exampleA 75-year-old woman who has had a previous myocardial infarction (MI) presents within 4 hours of symptom onset with suspected acute MI and ST elevation on electrocardiogram (ECG); she is being considered for aspirin and reperfusion therapy. Data from randomised trials of aspirin, thrombolytic therapy and immediate coronary angioplasty are considered. The patient has no known contraindications to these treatments. Based on evidence from the ISIS-2 trial,9 there is strong evidence that aspirin reduces the risk of short-term mortality, by about 23%, both overall and within most subgroups examined, including women and patients aged over 70 years. However, in this trial, among patients with a prior MI, no significant treatment benefit was observed. While the treatment effect in this subgroup apparently differed from patients without a prior MI, the interaction may have been a chance finding owing to the many subgroups examined. (For example, the chance of at least one significant result at the 5% level among 20 independent tests is over 50%.1) Consequently, the trial evidence still strongly supports the use of aspirin therapy in this patient. Randomised trials of thrombolytic therapy in the FTT overview10 have also demonstrated clear evidence of a reduction in mortality from such treatment for patients with acute ST elevation presenting within 12 hours of symptom onset. This overview also suggested diminished effectiveness of such treatment in the elderly (test for trend with older age, P = 0.01). However, in this case, much of the heterogeneity could be explained by the fact that older patients more often presented later (after 12 hours) and without the specific diagnosis of ST elevation on ECG. Lack of ST elevation and late presentation to hospital relate directly to the underlying biology and are linked to diminished effects of treatment. Once these confounding factors have been taken into account, there is much less rationale for considering different treatment or withholding thrombolytic therapy simply on the basis of the age of our patient.11 Next, the role of immediate coronary angioplasty in such a patient could be considered. Randomised trials, particularly in specialised centres, have suggested an additional treatment benefit for immediate coronary intervention compared with thrombolysis. An individual patient data overview of earlier randomised trials suggests a relative reduction in death or reinfarction of about 50%, with similar relative effects in each of the subgroups examined.12 However, the absolute benefits of treatment (absolute risk reductions) were estimated to be much greater in the patients at high baseline risk, particularly those aged over 70 years (see Box 3). Consequently, if treatment with immediate angioplasty is considered appropriate in the particular hospital setting, it would be likely to have greater absolute benefit for this older patient than the average patient. Finally, the role of long-term treatment in this patient could be considered. Should statin therapy be considered on the basis of the evidence from such trials as the LIPID and CARE studies?13,14 Both of these had insufficient evidence to show reductions in mortality with treatment for women separately. In the LIPID trial, older and younger patients had similar relative reductions in events (Box 3), but older patients at higher baseline risk had greater absolute benefit. Fewer women than men were studied in these trials, and yet the results for women were not inconsistent with those for men (Box 3). The effect of statin therapy for women with prior CHD is illustrated further by the results of the 4S, CARE and LIPID trials.14 The combined results of these three trials show an overall significant reduction in coronary events; the estimates from the separate trials vary but are still consistent with the overall result. Evidence of a similar relative treatment effect from statin therapy in both women and men has recently been confirmed by the results of the Heart Protection Study.15 Finally, for some of these decisions, different recommendations for treatment may still apply even when a similar relative treatment effect seems valid and patients are at the same baseline risk. Circumstances in which different recommendations will be appropriate include: Where the importance of different outcomes (of benefit and harm) varies for different patients Where patient preference varies for other reasons Where there are limitations in applying the trial results in a particular setting, related to such factors as the skill or experience of practitioners and access to technologies. Principles for applying subgroup analysis to decisions about individual patients are summarised in Box 4. Cautious interpretation of the results of subgroup analyses is generally advisable. 1: Treatment effects in subgroups of men and women in three hypothetical trials Overall relative risk: 0.75 for Study 1; 0.80 for Study 2; 0.87 for Study 3; represented by the vertical dashed line in each case. 2: Interpreting treatment effects in different subgroups within a controlled clinical trial* * The decision pathway assumes there was a reasonable basis to consider the subgroups to have the same underlying condition. 3: Absolute risk reduction (ARR) and numbers needed to treat (NNT) in age and sex subgroups In none of these cases is there evidence of treatment effect modification (all P values for interaction are non-significant). ARRs and NNTs are derived from the overall relative treatment effect. 4: Principles for using subgroup evidence for making decisions about individual patients Use the subgroup-specific result only when there is (unconfounded) evidence of interaction and, ideally, confirmatory evidence. Use the estimated overall treatment effect if there is no evidence of heterogeneity (no interaction or treatment-effect modification). Adjust the size of the treatment benefit (and harm) according to the patient’s baseline risk. Consider patient preferences regarding each outcome when there are significant trade-offs in benefit and harm.

R John Simes MD, SM, FRACP · Val J Gebski BA, MStat · Anthony C Keech MB BS MScEpid FRACP

Information science EBM: Trials on trial 15 March 2004 Free

Subgroup analysis in clinical trials

Clinical trials represent a major investment by investigators, sponsors and participants, and it is reasonable to attempt to gain the maximum information from them. Practitioners and regulatory agencies are keen to know whether there are subgroups of trial participants who are more (or less) likely to be helped (or harmed) by the intervention under investigation, and a recent survey of trials published over 3 months in four leading journals found that 70% included subgroup analyses.1,2 Furthermore, regulatory guidance documents (such as the Committee for Proprietary Medicinal Products September 2002 document Points to consider on multiplicity issues in clinical trials3) strongly encourage appropriate subgroup analyses. The results of subgroup analyses can also drive changes in practice guidelines. For example, the United States National Institutes of Health issued a clinical alert following the unexpected finding in the BARI (Bypass Angioplasty Revascularisation Investigation) trial that mortality after angioplasty in patients with diabetes was nearly double that after bypass-graft surgery (P = 0.003).4 Meaningful information from subgroup analyses within a randomised trial is restricted by multiplicity of testing and low statistical power. There is therefore a tension between our wish to identify heterogeneity in the responses of trial participants to trial interventions and our technical capacity for doing so. Surveys on the adequacy of the reporting of clinical trials consistently find the reporting of subgroup analysis to be characterised by poor practice.2,5-7 Item 18 of the CONSORT checklist (Box 1) deals with the multiplicity issues that arise in subgroup analysis.8 Problems in subgroup analysisThe problem of multiple testingStatistical investigation of large numbers of subgroups inevitably shows significant interactions with the effectiveness of the trial intervention. By definition, testing at the 5% level of significance will erroneously report a statistically significant difference between subgroup categories in about 5% of the tests performed (so-called false-positive results). Trials with multiple comparisons to assess the comparability of randomised groups at baseline confirm this prediction.1,9 In subgroup analysis, where a plethora of factors (eg, sex, age, race, centre, smoking status, stage of disease, and coexistent disorders) may influence outcome, the risk of false-positive results is high.10 Overly enthusiastic analysis of subgroups can reveal statistically significant differences in outcome between subgroups even where neither arm of the study receives any intervention.11 In some cases, such as in the ISIS-2 study, which found a slight adverse impact of aspirin therapy on patients born under the star signs Gemini and Libra, and that aspirin helped after the first, but not subsequent, infarctions,12 the results of the subgroup analysis may be dismissed as contrary to current understanding of biological mechanisms. In other cases, such as the BARI trial,4 whether the finding was valid could only be established by additional studies.13,14 The problem of statistical powerMost studies enrol just enough participants to ensure that the primary hypothesis can be adequately tested. Therefore, statistical tests on subgroups will have only power to detect substantially larger effects on the same endpoint. Loss of compliance, together with adjustments for multiple testing, will exacerbate this lack of power.6 In consequence, when tested separately, many of the subgroups will fail to show the statistically significant treatment effect that was shown in the main population; at the same time, genuine differences in response to treatment (so-called heterogeneity) between study subpopulations may also go undetected. Can the problems be overcome?Despite subgroup analyses generally lacking statistical power, when used repeatedly to look for differences across many factors (eg, sex, age, smoking status, blood pressure) they have a proclivity to detect spurious effects. We are thus forced to reconcile our wish to find genuine differences between subgroups with the need to minimise the risk of accepting and publishing false positives.2,6 One solution to this dilemma is to accept that the results of subgroup analysis are hypotheses. Guidelines such as those given in Box 2 are intended to help readers identify which hypotheses are strong and which are weak. However, even among experts, opinions range from only accepting pre-specified subgroup analyses supported by a very strong a priori biological rationale15 to a more liberal view in which subgroup analyses, if properly carried out and carefully interpreted, are permitted to play a role in assisting doctors and their patients to choose between treatment options.16 Trial designAre the subgroups appropriately defined?Subgroups based on characteristics measured after randomisation, such as compliance, should be avoided, as allocation to the subgroup may be influenced by the intervention. Similarly, it is preferable to use the intention-to-treat population, as reasons for withdrawal may not be balanced between treatment arms. For example, adverse drug events may be a more important reason for withdrawals from an active treatment arm, whereas lack of efficacy may be more important in a placebo-controlled arm.17 Were the subgroup analyses planned before commencement of the study?In general, subgroup analyses should be defined a priori and purposely on the basis of known biological mechanisms or in response to findings in previous studies. Ideally, the choice of the subgroups and the expected direction of the subgroup difference should be justified in the trial protocol. Where a particular subgroup analysis is of great interest, adequate power to show the results can be designed into the trial, for example by using an expanded endpoint for the subgroup analysis. At the other extreme, subgroup analyses that are decided on once the dataset has been examined should be treated with scepticism. Intermediate between these two extremes are cases, such as occurred in the BARI trial, in which the subgroup analysis, although not originally planned, was decided on during the course of the trial in response to findings in other studies (with the investigators remaining blinded to the interim results of BARI).4 ReportingThe study report should include all the information required to assess the validity of subgroup analyses reported. In particular, the number of subgroup analyses should be declared, as this will enable readers to assess whether the issue of multiple testing is being dealt with. Analyses planned a priori, and the rationale for choosing them, should be clearly stated. Summary data, including event numbers and denominators for all the subgroup analyses, even the uninteresting ones, should be included, as this will facilitate future meta-analyses of the data and help prevent publication bias.18 The impact of multiple tests on the chance of declaring as statistically significant at least one false-positive result is shown in Box 3. Statistical analysisSome investigators avoid the issue of multiplicity of testing by tabulating the observed outcomes for the subgroups of interest without undertaking any formal statistical analysis. The data become available for meta-analysis,18 but there is the disadvantage that the investigator may fail to detect and draw attention to an important heterogeneity in the population. The statistical methods used should be appropriate for the hypothesis being tested. The common practice of performing subgroup-specific tests of treatment effect is flawed in that it is testing the wrong hypothesis.19 The hypothesis that should be tested is whether the treatment effect in a subgroup is significantly different from that in the overall population.19 Testing for a statistically significant treatment effect in a subgroup is hindered by a small sample size. The appropriate tests to use when analysing heterogeneity of responses among subgroups are interaction tests,2,10 for which worked examples are available.19,20 One study found that these were used in only 43% of 35 trials which reported subgroup analyses in their sample.2 Finally, the article should state whether the statistical tests used included adjustments for multiplicity. InterpretationBecause subgroup analyses have less power to detect a therapeutic effect than the main study, the trial report, especially in the Abstract or Conclusions, should emphasise the overall result. Given the risks of false-positive findings when multiple subgroup analyses are performed, it is not surprising if a subgroup-specific test shows a significant (P < 0.05) or suggestive (P = 0.05 to P = 0.10) effect of treatment, even when the trial failed to do so overall.2,7 Investigators are often tempted to highlight a particular subgroup analysis.2,7 For example, in one trial the suggestion that a psychosocial nursing intervention following myocardial infarction was harmful for women (P = 0.064), but not men (P = 0.94), was highlighted, even though the intervention did not affect survival in the overall population21 (and a test for interaction was not significant2). A number of arguments may be used to support the validity of a claimed subgroup effect (see, for example, the BARI trial4 and Rathore et al22): replication in another independent study; the presence of a dose–response relationship; reproducibilty of the observation in independent samples within the study, such as within individual sites; and the availability of a biological explanation. Of these, the first is the strongest evidence. For example, even though the BARI study found no difference in survival following bypass surgery or angioplasty in the overall population, the validity of the subgroup findings was supported by other studies.4 On the other hand, the report by Rathore et al that digoxin use is associated with a significantly increased risk of death among women (P < 0.014)22 is weakened by the fact that it was a post-hoc analysis which was motivated by “biological suspicion” rather than by suggestive findings in earlier trials. Biological justifications for the findings of a posteriori (exploratory) analyses, on the other hand, carry little weight6,23 — the reports that diabetes is more common in boys born in October,24 and that lung cancer is more common in people born in March,25 included (in)credible biological explanations after the findings had been revealed. The strategies for overcoming some of these difficulties in interpreting subgroup analyses will be explored in a forthcoming article in this series. 1: CONSORT checklist of items to include when reporting a trial8 Selection and topic Item no. Descriptor Ancillary analyses 18 Address multiplicity by reporting any other analyses performed, including subgroup analyses and adjusted analyses, indicating those pre-specified and those exploratory. 2: Checklist for subgroup analyses Design Are the subgroups based on pre-randomisation characteristics? What is the impact of patient misallocation on the subgroup analysis? Is the intention-to-treat population being used in the subgroup analysis? Were the subgroups planned a priori? Were they planned in response to existing trial or biological data? Was the expected direction of the subgroup effect stated a priori? Was the trial designed to have adequate power for the proposed subgroup analysis? Reporting Is the total number of subgroup analyses undertaken declared? Are relevant summary data, including event numbers and denominators, tabulated? Are analyses decided on a priori clearly distinguished from those decided on a posteriori? Statistical analysis Are the statistical tests appropriate for the underlying hypotheses? Are tests for heterogeneity (ie, interaction) statistically significant? Are there appropriate adjustments for multiple testing? Interpretation Is appropriate emphasis being placed on the primary outcome of the study? Is the validity of the findings of the subgroup analysis discussed in the light of current biological knowledge and the findings from similar trials? 3: Probability of at least one significant result at the 5% significance level given no true differences Number of tests Probability 1 0.05 2 0.10 3 0.14 5 0.23 10 0.40 20 0.64

David I Cook MD, FRACP · Val J Gebski BA, MStat · Anthony C Keech MScEpid, FRACP

Lessons from early large-scale adoption of celecoxib and rofecoxib by Australian general practitioners

Timothy H J Florin Director of Gastroenterology, Department of Medicine, University of Queensland, Mater Health Services’ Adult Hospital, South Brisbane, QLD 4101. t.florinATuq.edu.au To the Editor: In their article on adoption of celecoxib and rofecoxib by Australian general practitioners, Kerr et al noted that “the increase in COX-2 [cyclooxygenase-2] prescribing coincided with a period of energetic marketing to the medical profession, which promoted the message that the new C2SNs [COX-2-selective non-steroidal anti-inflammatory drugs] were ‘safer’ than traditional NSAIDs [non-steroidal anti-inflammatory drugs].”1 The implication is that the decision of Australian GPs to prescribe these new drugs may have been less than independent or rationally informed. The reason for prescribing C2SNs is that, like traditional NSAIDs, they relieve arthritic pain and so promote mobility, although, unlike traditional NSAIDs, they do not inhibit cyclooxygenase-1. While the power of advertising is undeniable, the simple message about C2SNs is that there is an approximate 50% reduction in clinically significant gastrointestinal (GI) complications compared with traditional NSAIDs.2 There are over a dozen articles to support the better GI side-effect profile of C2SNs. Most data support a non-cumulative, reversible, but constant, risk of peptic and other, more distal GI bleeds, or perforation, with the coefficient of risk being significantly greater for NSAIDs.3 Although the data from the CLASS study did suggest that the higher GI morbidity of NSAIDs seemed to diminish with time,4 that of celecoxib remained at a constantly lower rate.5 This is clinically important for all our patients, and especially for our ageing population with their comorbid conditions and polypharmacy. I suggest that it is for this single reason that many doctors have been quick to take up C2SNs for their patients. The better GI safety enfranchised patients who previously could not take NSAIDs safely, and could explain why the overall anti-inflammatory market increased by 20%.1 However, no one suggests that the C2SNs are free of non-GI side-effects. To the best of my knowledge, COX-2-specific NSAIDs have not been promoted as being free of non-GI side-effects or better than COX-1-specific NSAIDs in this regard. While agreeing with the last sentence of Dowden’s editorial that “new is not always better”,3 the opposite — that “new is sometimes better” — is also true. Thus, we persuaded the accountants in our hospital, who rightly participate in the determination of which drugs are available on its formulary, to accept one of the COX-2 drugs because of its better GI complication profile. While this will not reduce its pharmacy budget, it is anticipated to reduce overall hospital costs in this area,5 which should allow a reapportionment of its budget to other areas of need. Of course, the hospital is watching carefully for any unforeseen “serious adverse effects which sometimes only emerge after marketing”.3

Timothy H J Florin

Lessons from early large-scale adoption of celecoxib and rofecoxib by Australian general practitioners

Stephen J Kerr,* Andrea Mant,† Fiona E Horn,‡ Kevin McGeechan,§ Geoffrey P Sayer¶ * Decision Support Coordinator, National Prescribing Service, PO Box 1147, Strawberry Hills, NSW 2012; † Area Advisor, Quality Use of Medicines, South East Health, Sydney Hospital, Sydney, NSW; ‡ Research Analyst, § Senior Research Analyst, ¶ General Manager — Research, Health Communication Network, St Leonards, NSW. skerrATnps.org.au In reply: The advantage of the COX-2-selective NSAIDs (C2SNs) is the reduction in clinically significant gastrointestinal complications compared with conventional NSAIDs, but, as Florin agrees, other toxicities, including the risk of renal failure and heart failure, are similar for C2SNs and the older drugs.1 We speculated that doctors may have been more aware of the differences between the new and the conventional anti-inflammatories rather than the similarities: between 4.7% and 7.9% of patients in our study cohorts were treated with a combination of drugs which placed the patient at risk of renal complications. Florin points out that there is an approximate 50% reduction in clinically significant gastrointestinal (GI) complications with C2SNs compared with conventional NSAIDs. If, in a population, the annual incidence of serious GI complications with NSAID use is around 1.4%,2 then the absolute risk reduction is 0.7%. This means that about 140 patients would need to be treated with a C2SN for one year to prevent one serious GI complication. Messages conveyed in this way may be more pertinent to clinical decision making than a statement about relative risk reduction. Florin also notes the problems with elderly patients who often take multiple medications, and are probably at increased risk of upper-GI events with NSAIDs. Our data demonstrated very high prescribing rates in patients who were not elderly. Over 20% of patients in our cohorts were aged less than 50 years, and over 50% were aged less than 65 years. Furthermore, between 34.5% and 61.3% had no pain medication prescribed in the 12 months before the first C2SN prescription, suggesting that C2SNs may have been used as a first-line pain medication in these patients. Quality use of medicines advocates prescribing which is safe, judicious, effective and cost-effective. Recent pharmacoeconomic studies suggest the cost-effectiveness may only be realised when prescribing of C2SNs is confined to patients who are at high risk of GI complications.3,4

Stephen J Kerr · Andrea Mant · Fiona E Horn · Kevin McGeechan · Geoffrey P Sayer

Lessons from early large-scale adoption of celecoxib and rofecoxib by Australian general practitioners

John S Dowden Editor in Chief, Australian Prescriber, Suite 3/2 Phipps Close, Deakin, ACT 2600. jdowdenATnps.org.au In reply: The general practitioners’ decision to prescribe COX-2 inhibitors was rational, but the information underpinning their decision was less than independent. A big reduction in short-term relative risk can be persuasive, even if the absolute benefit is small. General practitioners deal with whole patients, so they consider the overall risks of treatment, and not just one adverse effect. While COX-2 inhibitors may have gastrointestinal advantages, they may have cardiovascular disadvantages. Treatments for chronic conditions should be based on long-term data. The observation that most of the ulcer complications in the second half of the CLASS study were in patients taking celecoxib is therefore important.1 Undoubtedly, some patients who could not take non-selective non-steroidal anti-inflammatory drugs (NSAIDs) were able to tolerate COX-2 inhibitors. However, Kerr et al found that up to 61% of patients given a COX-2 inhibitor had not previously been prescribed any analgesia.2 It seems unlikely that so many people suddenly required analgesia that only a COX-2 inhibitor could provide. The Pfizer-funded study by MacDonald et al shows that UK general practitioners tended to prescribe COX-2 inhibitors for patients at risk of gastrointestinal haemorrhage.3 This follows the advice of the National Institute of Clinical Excellence (NICE). However, NICE also recommended against the routine use of COX-2 inhibitors.4 A review by the Canadian Co-ordinating Office for Health Technology Assessment has also concluded that COX-2 inhibitors may have no significant safety advantage over diclofenac.5 Solomon et al conclude that the cost of adverse effects of NSAIDs in low-risk elderly patients is modest.6 However, there is no comparison with COX-2 inhibitors, so we do not know if they reduce this cost. Hospital accountants may be interested to know that researchers at the Mayo Clinic concluded that, in terms of averting gastrointestinal events, the most cost-effective analgesic is paracetamol.7

John S Dowden

Call for MJA submissions

Have you got a paper burning a hole in your desk, waiting to be sent to your favourite journal? Or an idea for an article, needing just a little more inspiration to push it into print? Now could be your big chance! MJA theme issues 2004 17 May — Indigenous Health Closed to research manuscripts 5 July — MJA’s 90th Anniversary Closing date 19 April 19 July — General Practice Closing date 15 March 16 August — Doctors’ Health and Lifestyle Closing date 10 May 18 October — Adolescent Health Closing date 5 July Contact details: email editorialATampco.com.au phone (02) 9562 6666 This year, the MJA will be publishing five special theme issues, and we are inviting submissions in all categories of articles (Research, Viewpoints, For Debate, Clinical Updates, Snapshots, and even Editorials, although these last should be discussed with an editor first to check suitability). Last year’s big hits were Women’s Health (16 June 2003) and Chronic Illness (1 September 2003). This year’s special issues are shown in the Box and the closing dates for submissions. Please remember that a lot of work is involved in preparing a manuscript for publication, including editorial assessment, peer review, revision, editing and finally page layout — it doesn’t happen overnight, so the sooner you can submit your material, the better. July will be a very special month for us because it is our 90th birthday — we hope our readers will be able to help us celebrate in style. In the decade since our last anniversary issue, there have been major changes to the Journal and to how readers prefer to receive and use their medical information. We would like to know what you think of the Journal, how medical journals and the MJA have changed over your years of reading, and what kind of journal you would like the MJA to be on our 100th anniversary. We would also be very interested to know which MJA articles have made a difference to the way you practise medicine or have affected the way you treated a particular patient. Finally, we would be interested in your views on changes in medicine in the last decade. In all issues and all categories, we’re looking for originality, innovation, topicality and positivity, with a focus on “health” rather than disease. Do consult our Advice to authors for submission requirements and details of article categories (www.mja.com.au/public/information/instruc.html). We look forward to receiving your submissions.

Bronwyn Gaut

Australian healthcare reform: in need of political courage and champions

Ron J Lord Editor, Healthcover, 28 Hereford Street, Glebe, NSW 2037. hcoverATihug.com.au To the Editor: The Editor’s article on health reform and the Australian Health Care Summit,1 in which he expressed sentiments with which I agree, included a Box setting out the “egalitarian and socially cohesive principles underpinning Australia’s healthcare” reaffirmed by the Summit. However, the Box contained a Christmas tree and an invitation to readers to enter a poem in the MJA’s Christmas Competition 2003. Among the lines were: “’Tis Christmas, the season to be kind”. While obviously the result of a glitch in the production process, you managed — much to the envy of other editors and publishers seriously wounded by such glitches (to the extent that entire print runs have had to be pulped and then reprinted) — to fall on your feet. I could not think of a better (or more comprehensive) set of principles to underpin our healthcare system than those embodied in the message and spirit of Christmas. Perhaps God moves in mysterious ways.

Ron J Lord

Australian healthcare reform: in need of political courage and champions

Robert A Jones Specialist Gynaecologist, Adelaide Private Menopause Clinic, Memorial Medical Centre, 8/1 Kermode Street, North Adelaide, SA 5006. robjonesAT senet.com.au To the Editor: 9/15 was disaster day at the MJA.1 Not only was the Editor guilty of printing perseveration, but his “Box” seems to have been transmogrified from . . . “(the) socially cohesive principles underpinning Australia’s healthcare” to an invitation to “expose” the readers of the Christmas journal to some “witty prose”. Perhaps the “healthcare dialogue” has indeed been reduced to rhyming couplets, possibly accompanied by the health ministers fiddling while the rest of us burn?

Robert A Jones

Australian healthcare reform: in need of political courage and champions

Martin B Van Der Weyden Editor, The Medical Journal of Australia, Locked Bag 3030, Strawberry Hills, NSW 2012. editorialATampco.com.au In reply: Fate (or God) moves in both mysterious and wondrous ways. Maybe the manoeuvring and machinations of our health ministers in the consummation of the 2003–2008 Australian Health Care Agreements are worthy of: If you have a little ditty You would like to expose, Send it to the Journal We’ll publish your witty prose. All I can say is that the faux pas in the production process shows that, despite its high technology, it is still a human process. To err is human, so let’s not make a very public faux pas all consuming.

Martin B Van Der Weyden

Information science Medicine and the media 1 December 2003 Free

An analysis of newspaper reports of cancer breakthroughs: hope or hype?

Objective: To assess the importance of cancer “breakthroughs” reported in the popular media 10 years after their publication.Study design: Questionnaire-based survey in 2003 of expert opinion on the importance of all alleged cancer “breakthroughs” in cancer research or treatment reported in news articles in The Sydney Morning Herald between 1992 and 1994.Main outcome measures: Assessment of each “breakthrough” by an expert in the relevant cancer subspecialty on seven measures of current importance.Results: 31 unique reports of alleged cancer “breakthroughs” were identified, and experts responded to questionnaires on 30. Thirteen of these 30 reports (43%) were judged as not having been supported by further research in the following decade, with three (10%) having been refuted, while 16 (53%) were judged to remain potential breakthroughs, but more research was required. Eight “breakthroughs” (27%) had, or would soon be, incorporated into practice.Conclusion: Cancer research findings reported in newspapers as “breakthroughs” are often not true breakthroughs but may be important for ongoing research. Consumers are likely to be receiving an overly optimistic picture of progress in understanding and treating cancer.

Ethel S Ooi MB BS · Simon Chapman PhD

Information science Quotable quotes 1 December 2003 Free

Between the sounds of silence . . .

Quotes from MJA contributors in 2003 Ever since the Medical Journal of Australia was first published in 1914, the library-like atmosphere that usually pervades our premises (reflecting the quiet industry within) is occasionally shattered by an exclamatory outburst. Such fractures of the usual peace and quiet can be perpetrated by any of our editorial staff, and can signify joy, indignation, solid agreement, pure amazement or other sundry emotions. Outbursts usually occur on reading a submission to the Journal — be it a manuscript, a peer reviewer’s report or other form of correspondence. The general effect is to make the working day all the more enjoyable! This year, the Journal’s new, modern open-plan office has facilitated a sharing of these moments. Here, we share with you a selection from our discerning collection of putative, causative agents in the hope that you, too, will gain some measure of sonorous pleasure in the reading of them. The quotes are real life and presented in raw form, although we could not curb our habit of arranging (and rearranging) material being considered for publication in the Journal. Further, in deference to our journalistic colleagues, we have chosen not to reveal our sources. Some of you may recognise that a few of these excerpts are already in the public domain. In acknowledgement of the good sense of the penultimate quote, and our admiration for all statisticians, we leave any analysis of these data to the experts. Lastly, we wish to thank all of you for your contribution to the Journal. You are all esteemed by us, whether you be an author, reviewer or reader. What’s in a title?“Adverse event reporting in clinical trials: regulatory tail wags the research ethics committee dog, distracting the latter from more useful activity” Gender issues“Most women live in an environment that is also populated by men.” “Seven of the nine patients who developed neurological sequelae were female and the rest of them were male.” All in a day’s work“My apologies for the delay in getting this [economist’s review] back to you but it really is about time that you guys worked out a cure for the common cold.” “I am not sleeping until I get a draft revision to you. I have not heard from my co-authors as yet. I will keep you posted. Time: 3.30 am.” Making a statement . . .“Cardiac arrest is more successfully treated in Chicago or Heathrow airport, or an American Airlines or Qantas jet, or in a Boston post office, than in the vestibules, corridors or general wards of Australia’s premier hospitals.” “Declaring war and prescribing drugs are decisions dependent on information. If that information is incomplete or inaccurate there can be calamitous consequences.” “The Academic Clinician is well recognised to be a breed on the verge of extinction internationally, and this sort of program is the last great hope for ensuring its survival.” ArbitrationOn occasion, the Journal’s editorial committee asks a reviewer to help us out when an author and responder are at odds . . . the exercise generally proves fruitful! “In both of his letters, the author uses analyses which aren’t correct. In fact, his second letter misses the [critic’s] point entirely, since he repeats his error. The responder [critic] adopts the correct method. The story may be recast this way: Author: 1 + a = apple. Responder: The correct method is to compare like with like. If you compare dissimilar things, you get a third, uninterpretable thing; 1 + 1 = 2. Author: Don’t look at me. When I contacted my source, I was given the information I got. Oh, and by the way, I also found that 1 + b = orange. I will venture to say that arguments made on the basis of faulty statistics are themselves faulty. I would suggest that the author seeks professional statistical assistance to clarify his analysis.” End-note“The conclusions end on a note that is remote from the key of the paper (to use a musical analogy)”.

Ann T Gregory MB BS, GradDipPopHealth

Information science EBM: Trials on trial 21 July 2003 Free

Baseline data in clinical trials

Although reporting baseline data seems simple, it is crucial information for readers in judging the validity of a trial. Knowing the baseline characteristics of the trial participants allows readers to assess how closely these match patients seen in their own clinical practice, and therefore how generalisable the results of the trial will be (so-called external validity). Baseline characteristics also allow the success of randomisation to be assessed. In studies where important baseline factors appear well balanced, it is likely that any differences in outcome between the intervention and control groups are a real effect of treatment (one component of internal validity). For these reasons, the reporting of baseline demographic and clinical characteristics of each group is a requirement of the CONSORT statement.1 The item and its descriptor as they appear in the CONSORT checklist are shown in Box 1, and a checklist for baseline data is provided in Box 2. ContentBaseline data should adequately describe the population in the trial. This means including demographic variables, known factors that influence the outcome (including medications being taken by participants), factors that are likely to modify any benefit of treatment, and those that may predict adverse reactions. These factors are called potential "confounders", because, if they are imbalanced between the treatment groups at baseline, they may result in an apparent treatment effect when none exists, or mask an effect that does exist. Baseline data should also include any factors (especially known potential confounders) that have been used as strata for randomisation. Stratified randomisation, described in detail earlier,2 is used when a baseline characteristic, such as tumour stage, is known to affect outcome risk; the characteristic is therefore included in the randomisation algorithm to minimise imbalances between treatment groups. This is particularly useful in small studies. If the study population contains subgroups of particular interest, the characteristics defining these subgroups, and numbers or proportion in each group, should be stated. For example, in a long-term trial of a new medication for preventing heart attack, diabetes mellitus would be a potential confounder (as people with diabetes have a much higher risk of heart attack than similar people without diabetes). Those with diabetes in this study would also be an interesting subgroup in whom the effects of the intervention might be different. Similarly, concurrent therapy with aspirin (which would substantially reduce the risk of heart attack) could confound the trial results if there was an imbalance between trial groups in the proportions of patients taking aspirin; aspirin therapy might also influence the likelihood of adverse reactions to study therapy. Baseline factors can be determined from interviews, physical examination, laboratory measures or imaging studies. MeasurementBaseline data are measured as close as possible to the time that participants are randomly allocated to study groups, and in all cases, should be measured before the allocated treatment commences (information collected after the commencement of trial treatment may have been altered by the treatment itself, and is generally not regarded as baseline data). Ideally, baseline data should be collected on all patients screened for eligibility, as this would provide further information about the generalisability of the trial population. However, this is not always practicable or affordable, so some variables (eg, tissue biopsy, measurement of genetic markers, expensive imaging tests) are measured only in actual participants randomly allocated to a trial group. For factors that are not constant, the conditions under which the baseline data are collected should be stated in the methods section of the study report. For example, it should be clear whether blood pressure recordings were measured sitting or supine, or after a specified rest period; also whether a single reading, the average of several readings, the highest of two, or the first of two or more, was used. Baseline data as entry criteriaIn some circumstances, threshold levels of one or more baseline variables will form part of the entry criteria for the study. In this case, if an extreme value of a baseline factor, such as high blood pressure, is required to qualify a person for entry into a study, potential participants whose value on the day of screening is more extreme (higher) than their usual level will be more likely to qualify for entry. A second baseline reading of the average blood pressure for this group will be lower and more accurately reflect their usual blood pressure; this is known as regression towards the mean.3 For this reason, remeasuring factors required for entry is desirable, to establish a more realistic group average value of the characteristic at baseline. PresentationThe baseline characteristics are usually presented in the first table in a report. Care should be taken to include the necessary descriptive information without overwhelming readers with unnecessary details. For example, in the recent AFFIRM trial comparing rate control with rhythm control of atrial fibrillation, the published first table has 16 baseline characteristics, each with a mean and percentage value for the overall group, and for both treatment groups separately, together with P values.4 The resulting table of 107 values and four footnotes may make it difficult for some readers to extract the key information.5 A simpler presentation appears in the FRISC II study of invasive compared with non-invasive treatments for unstable coronary artery disease.6 This presents more baseline characteristics (20), but by minimising detail (omitting overall group and P values), allows a more rapid comparison of the characteristics between groups. Comparability between groupsIf randomisation has been performed correctly, the groups should be similar in baseline characteristics, except for the play of chance. Stratification in the randomisation process further restricts the extent of chance imbalances.2 For continuous variables (such as blood pressure, age, cholesterol level), the similarity of the treatment groups should be assessed by comparing relevant summary measures (mean and standard deviation, or median and range). For categorical factors (such as sex, disease stage), the numbers and proportions in each category level should be shown for each treatment group. The more similar the treatment groups, the more credible are the trial results as reflecting a true result of treatment, especially if unadjusted analyses are presented.5,7 Use of P values to assess randomisationUse of statistical tests to compare the balance and/or values of baseline characteristics between the study groups and the presentation of P values are not uncommon. However, many authors assert that this is inappropriate.3,5,8-10 If randomisation has been performed correctly, chance is the only explanation for any observed difference between groups at the outset of the study, in which case statistical tests become superfluous. Consequently, only if it is suspected that the randomisation process has failed or was flawed, can performing significance tests on the baseline data be readily justified.8 It is worth remembering that, if 20 baseline characteristics are presented from a trial using simple randomisation, it is more likely than not that at least one characteristic will show a significant imbalance between groups at two-sided P < 0.05 by chance alone (actual likelihood, 64%). In any case, providing P values is not a substitute for carefully describing, in the results section, any imbalances between study groups that may be clinically important. For example, in a trial of a thrombolytic drug, a 1% baseline difference in history of previous intracranial haemorrhage may not be statistically significant, but could still affect haemorrhagic stroke rates after treatment (an outcome of the study), and hence could be regarded as potentially clinically significant. If there are imbalances that are considered important to the final study results, they should be accounted for by an adjusted analysis of the data, not simply noted with a P value in the first table.7 Other uses of baseline dataA longer-term benefit of collecting comprehensive baseline data is that, after outcome data become available, it allows the estimation of risk of the outcome in the control group, related to various baseline characteristics. This effectively uses the control group as an epidemiological cohort study, providing contemporary information about predictors of disease outcomes. In summary, careful planning and collection of baseline data enables performance of a high-quality trial and allows readers to clearly see the internal and external validity of the study. 1: CONSORT checklist of items to report when reporting a trial 1 Section and topic Item no. Descriptor Baseline data 15 Baseline demographic and clinical characteristics of each group 2: Checklist for baseline data Measurement Consider all important baseline variables to be measured and how they are to be measured before treatment starts: Demographic characteristics (age, sex, height, weight, etc) Known factors that predict the outcome (potential confounders) Factors that predict or alter the risk of adverse reactions Stratification factors Pre-specified subgroups. Reporting Tabulate relevant summary measures (eg, mean and standard deviation). Include all important baseline characteristics while keeping the table readable. Wherever possible, avoid displaying P values. Analysis In the results section, discuss the similarity of the two groups, highlighting any clinically important differences that may influence the outcome. Discussion Discuss the effect of the baseline data balance on the internal validity of the study and the comparability of the study population to patients seen in wider clinical practice.

David C Burgess BMed · Val J Gebski BA, MStat · Anthony C Keech MSc(Epid), FRACP

The "omnipotent" Science Citation Index Impact Factor

John H T Ellard Psychiatrist, Medical Specialist Centre, 710 Military Road, Mosman, NSW. To the Editor: I read with interest the article in the Journal on ranking medical journals and the fallacies to be found therein.1 I have a simpler method. I subscribe to two classes of journals: those specialising in psychiatry, and more general journals. The psychiatry journals I keep entire. However, as my house is of modest size I cannot do that with the general journals, so I tear out and file the articles that I find interesting and informative. You will be interested to know that in the past month I have filed away one article from the Lancet, one from the New England Journal of Medicine and three from the Medical Journal of Australia. What better measure of merit could there be?

John H T Ellard

Epitaph for the EBM in action series

This issue of the Journal (page 575) features the last article of the EBM in action series,1 conceived to show how clinicians can effectively look for the best available evidence to answer clinical questions. In the current medical climate, clinicians clearly need systems to obtain the best available evidence, and the responsibility for creating these systems falls on both individual clinicians and the organisations for ...

Christopher B Del Mar MD FRACGP · Jeremy N Anderson MD FRANZCP

The "omnipotent" Science Citation Index Impact Factor

The IF is a poor measure of the worth of journals, journal articles and authors Tell me the number; what is the ranking? All of us seem to love ratings. Whether it is the standings in the Rugby World Cup, the box office success of Harry Potter or the melting rate of Arctic ice, we all want numbers. So, why would it be any different for medical journal articles or even medical journals themselves? Who attaches importance to medical journal ratings? The owners/publishers of the journals, readers, advertisers, librarians and journalists may all be interested in journal ratings to varying degrees. Likewise, authors have a need to discern just how a publication is valued before deciding where to send the products of their labours. How can we evaluate the quality of an article or a journal? Properties of a medical journal that can be assessed include total circulation; readership numbers and surveys; quality of the editorial board, staff and peer reviewers; number of manuscripts received, percentage accepted, and turnaround; Science Citation Index (SCI) raw numbers, Immediacy Factor and Impact Factor (IF); number of paid subscribers; advertising revenue; listing on Medline; international distribution; cost to the reader; and page or peer-review charges to the author.1 But what do authors most value? Frank and colleagues have surveyed the Stanford University School of Medicine faculty regarding the factors that influenced their decisions about where to send manuscripts. The top attribute selected was "prestige".2 Impact factors are also used to adjudicate on academic performance. Some universities, especially in certain European countries, have decided that the IF of journals in which a faculty member publishes will enter into personnel decisions such as appointment, promotion and rate of pay.3 One would like to think that intelligent deans, chairs of departments and administrators, who work daily with faculty members, would have a better way to ascertain quality of performance than an arbitrary number. Seglen, of Norway, was an early critic of the IF, drawing attention to its narrow worth, and calling for its application to be reined in3 — but apparently to no avail. My belief is that the IF has one specific meaning: it is a clear measure of the extent to which a given journal functions as a connector of researchers in a specific field. This is one (but only one) critical function of medical journals. When I began as the editor of JAMA in 1982, JAMA's IF was in the range 3–4. Some considered this an embarrassment, so we set out to raise the IF as part of our efforts to improve the quality of the journal. We succeeded, to the extent that by the time I left the journal in 1999 its IF was in the range 10–11. Strange as it may seem, during the mid-1990s I deliberately tried to slow the growth of JAMA's IF. I was afraid that we were changing the character of the journal away from its fundamental purpose — to be useful to all doctors in their practices — and too far towards a research journal, used by researchers to communicate with each other. In this issue of the Journal, Walter and colleagues4 (page 280) criticise the IF, clarifying what it is and what it isn't. They describe an alternative way they have devised to judge the quality of articles (and presumably journals, if article scores are aggregated and averaged), using a five-person voting method guided by six criteria. It would have been interesting to see a side-by-side comparison between the article rankings of the selection panel and the SCI IF scores for each article. Walter and colleagues' form of post-publication peer review is now into its second year. The authors invite others to try it, and I hope there will be some who take up the challenge. In 1982, when I and my colleagues were developing plans to celebrate the JAMA Centennial, we tried an approach to evaluating medical articles somewhat like that of Walter et al. We wished to identify and republish the best 50 articles from the first 100 years of JAMA as "landmark articles". A list of prospective articles for inclusion was compiled from three sources: nominations by JAMA editorial board members and staff, entries in the 1976 edition of A medical bibliography (Garrison and Morton), and the most-cited JAMA articles from the Institute for Scientific Information. A total of 150 articles were nominated. The editorial board and staff then ranked the articles by a Delphi process and the top 50 were named "landmark articles".5 The article publication dates ranged from 1884 to 1968, with representatives from each decade. A subsequent analysis of the landmark articles by Eugene Garfield, founder of the Institute for Scientific Information (and father of the noted [or notorious] IF), demonstrated that of the 100 JAMA articles most cited by SCI up to 1983 only 13 were among the top 50 landmark articles, garnering from 174 to 506 citations by 1987.6 Thus, 37 landmark articles were not included in the top 100 JAMA articles ranked by total citations alone. Notably, such hugely important articles as those of Salk7 and Sabin et al8 had only received 39 and 90 citations, respectively, by 1987. So, number of citations and the derived IF are connected, but only to a limited degree. I would hesitate to suggest that the post-publication peer review process described by Walter et al could supplant the IF as the way that academic institutions, or even governments, decide on the merit of a publication or an author. But I can say with conviction that man (and academia) should not live by numbers alone.

George D Lundberg MD

Information science Viewpoint 17 March 2003 Free

Counting on citations: a flawed way to measure quality

The journal Impact Factor and citation counts are misconstrued and misused as measures of scientific quality. Articles must be read in order to judge their quality. We have introduced a system, which may be easily replicated, to identify the best articles published in a journal. Gloom or glee? Each September, journal editors and publishers anxiously await news of a particular figure from the Institute of Scientific Information (ISI) in Philadelphia, USA. The figure's value is promptly met with despair or delight. We are referring, of course, to the Impact Factor (IF) and the ritual surrounding its release that has come to dominate the editors' and publishers' calendar. Two of us, as editors ourselves, participate in this practice, albeit reluctantly and with rising apprehension that scientific publishing is being undermined by "numerology". What was introduced as an aid to librarians more than four decades ago to guide their selection of scientific journals has become, in the view of some people, an inappropriate means of scrutinising an applicant's "track record" when allocating research funds or considering academic promotions.1-4 Why inappropriate? Because the IF (defined as "the number of citations to a journal's articles published in the previous two years divided by the number of articles published by that journal during those two years"4-6) is conceptually and technically flawed, on a number of grounds: the quality of published material cannot be constrained by time — the two-year period set by the ISI for citations is arbitrary;7 the number of journals in the ISI's database is a minute proportion of those published;4 reviews are cited more frequently than original research, thus favouring journals that opt for these articles as part of a publishing strategy;1 the IF does not take into account self-citations, which amount to a third of all citations;8 errors are common in reference lists (occurring in up to a quarter of references), inevitably affecting IF accuracy;9 and the assumption of a positive link between citations and quality is ill-founded, in that we cite articles for diverse reasons, including to refer to research judged suspect or poor.10 If these flaws were not enough to instil scepticism about the IF's validity and precision, then the arbitrary assumption that the quality of a specific article correlates with the IF of the journal in which it appears is entirely ill-founded and, in itself, sufficient to warrant concern about its continuing use. As editors and researchers, we are duty-bound to analyse this assumption rigorously. What do we find? Consider two psychiatric journals published for a general readership: the Australian and New Zealand Journal of Psychiatry (ANZJP) and the Canadian Journal of Psychiatry (CJP). Applying the ISI's own "Web of Science" database,11 we can examine the purported link between IF, citations and the worth of a particular article. For instance, if we calculate the proportion of citations to all articles in each journal that is accounted for by the most-cited 50% of papers, a striking pattern emerges. In the case of the ANZJP, the most-cited 50% of papers published between 1990 and 1995 account for 94% (range, 91% [1992] to 98% [1990]) of all citations to articles in that journal. Figures for the CJP are virtually identical: 94% (range, 91% [1991] to 96% [1994]). A blunt summary of these findings is that half the articles published in both journals receive virtually no citations. We can conclude from the data (and from comparable findings from other journals, such as those in cardiology10) that to determine the academic worth of a paper from the IF of the journal in which it appears is ill-conceived and misleading. In an era of evidence-based medicine, its proponents avow that scientific progress can only be achieved by dint of diligent scrutiny of available data. Is it not incongruous, then, that the scientific community continues to cling to such an inadequate tool as the IF? What's more, the "parent" IF has spawned a range of flawed offspring, including "Scope-adjusted IF", "Discipline-specific IF", "Journal-specific influence factor", "Immediacy index" and "Cited half-life". As if that were not disconcerting enough, lo and behold, ISI recently faced a new rival, albeit short-lived (the venture collapsed in the wake of a threat from ISI to sue for violation of intellectual property rights). "PrestigeFactor.com" was launched in 2001, enticing us to ditch the IF and supplant it with another measure of journal quality, the "Prestige Factor" (PF).12 Its proponents boldly asserted that the PF provided "truer value" than the IF.12 There was, however, a fly in the ointment. Despite minor refinement (eg, the PF separated review articles from research reports and included citations to journal articles over the previous three years versus two), the underlying premise of both measures — that quality and number of citations are inextricably linked — was identical. It is also worth reporting the contemporary practice of open-access "e-journals" tracking their most popular articles through "hit rates".13,14 Again, we doubt that a popularity poll can indicate academic merit and fear that it may be misconstrued in this way. A watershedWe have reached a watershed — either we persevere with the notion that citations lie at the heart of scientific quality or we make a clean break. The latter option is attracting growing support. For instance, Richard Frackowiak, Dean of the Institute of Neurology in London, asserts that current measures are crude and reliance on them in making hiring-and-firing decisions is counterproductive.15 Zach Hall, a leading figure in US research, sees numerical methods as "excuses for not thinking".15 David Adam, a writer for Nature, highlights the absurdity of the situation in Finland, where government funding of university hospitals utilises a sliding scale corresponding to the IF of journals in which researchers publish their work.16 A notable development is a similar questioning among scientific bodies. The Deutsche Forschungsgemeinschaft, Germany's central research organisation, has promulgated innovative guidelines emphasising qualitative criteria in evaluating published material.17 As they posit, "Publications must be read and critically compared with the relevant state of the art and to the contributions of other individuals and working groups".17 Admirable but vexing. How are we to judge quality objectively? A more appropriate option?We have grappled with this challenge and devised an option for the ANZJP. Five international members of the journal's advisory board were invited in 2001 to identify that year's "top" articles. The quintet, selected on the basis of scholarship, professional integrity and knowledge of scientific psychiatry, were asked to select three publications per issue that satisfied one or more of the following criteria: adds consequentially to the field through original, innovative research findings; expands or challenges current knowledge; opens additional areas for new research activity; opens a pathway to advance knowledge; integrates discoveries obtained by different approaches and/or disciplines through creative synthesis, thus bringing new insights to bear on original research; and reflects critically on research findings to guide the direction of further research. The data were collated, and the titles of the nine articles gaining the most votes were announced in the April 2002 issue of the journal and posted on the websites of the Royal Australian and New Zealand College of Psychiatrists and Blackwell Publishing. This procedure is familiar in that it is, in essence, an extension of peer review. While by no means foolproof, we proffer this approach as a fresh way to establish the quality of the individual article. We sought feedback from the judges and learned that the task is feasible and the clarity and utility of the assessment criteria are satisfactory. The judges also found the assignment personally rewarding, even enjoyable! We are examining the method's reliability. Testing validity, of course, is more taxing given the lack of an objective yardstick. Interestingly, our experiment has been echoed by another initiative to highlight meritorious papers, namely the "Faculty of 1000" (F1000). Launched in November 2001 by the publishers of BioMed Central (a collection of wholly electronic biomedical "journals"),18,19 F1000 aims to identify the best papers in the basic biological sciences through the eyes of a "faculty" of over 1000 selected scientists who are experts in their fields. Thus, for instance, cell biology is divided into 18 categories, and the faculty members for each category select two to four papers each month from any journal, ranking them as "recommended", "must read" or "exceptional". The experts also briefly explain their choices. We applaud this initiative and see our own effort as complementary. Indeed, it is commonsensical that more than one method of identifying outstanding papers should be instituted, as there cannot be an absolute consensus. An invitationWe invite editors, publishers and authors to consider trying our experiment. We selected a new set of judges, again all distinguished figures in international psychiatry, to undertake a similar task for articles published in 2002. After this replication, we hope to be well placed to determine whether the method needs change. It would be disingenuous of us to conceal our fantasy that the annual "gloom or glee" ritual will be supplanted by published lists of "best quality articles", with their authors duly acknowledged. We may even witness the ISI collating the results, including the names of the judges and the criteria applicable for participating journals (consensually agreed criteria across all journals would be ideal). The implications are clear: successful authors will be able to cite articles that have made it to the "top" when documenting their academic track record, whatever the purpose (eg, applying for research grants). Depending on a measure devoid of any rational link to the appraisal of academic worth will be but a hazy memory.

Garry Walter PhD, FRANZCP · Karen Fisher MB BS · Sidney Bloch PhD, FRANZCP · Glenn Hunt MSc, PhD

Information science Letters 17 February 2003 Free

Screening mammography and mortality

Comment: The expression "give the lie to" has shifted its emphasis over the centuries, from the very direct "accuse (someone) of lying" to the much more abstract "show or imply (something) to be false". Some modern dictionaries, such as the Macquarie Dictionary (1997) and Merriam-Webster (2000), still give both meanings; others, such as the New Oxford Dictionary (1998), only the second. Large British and American databases, such as the British National Corpus, show that the phrase is usually used abstractly: one "gives the lie to" propaganda/a claim/an argument/a theory — whether in the context of academic discussion or political debate. The validity of an intellectual position is questioned, not the integrity of the person(s) associated with it. Yet, the simplicity of the phrase "give the lie to" probably gives the lie to the complexity of the challenge it expresses.

Pam Peters

Subscribe to MJA email alerts

No spam, you can unsubscribe anytime you want.

By providing your information, you agree to our Terms of Use and our Privacy Policy.

Thanks for Subscribing! Tell us more

Your email updates will use your name.

Good one! Your updates are coming

Thank you for subscribing to the MJA email alerts. Receive the latest content in your inbox.