This too shall pass: the performance of ChatGPT‐3.5, ChatGPT‐4 and New Bing in an Australian medical licensing examination
Authors: Oliver Kleinig, Christina Gao and Stephen Bacchi
Published online: 4 September 2023
To the Editor: Following the release of the generative pre‐trained transformer (GPT) ChatGPT in November 2022, a wide range of large language models (LLMs), including ChatGPT‐3.5 (GPT‐3‐derived), ChatGPT‐4 and New Bing (GPT‐4‐derived), have been made publicly available. There is suggestion that ChatGPT‐4 outperforms ChatGPT‐3.5 in answering questions from medical exams,1 but it is unknown whether GPT‐4‐derived LLMs consistently outperform GPT‐3‐derived LLMs.
Despite being GPT‐4‐derived, New Bing was fine‐tuned independently to ChatGPT‐4. Fine‐tuning provides additional training to LLMs for a specific task, such as responding to user queries, which can include human feedback on LLM‐generated responses.2 New Bing and ChatGPT may differ in performance due to such variations in fine‐tuning, and, at the time of testing (16–18 March 2023), only New Bing incorporated basic web searches.
We tested ChatGPT‐3.5, ChatGPT‐4 and New Bing against all 50 publicly available Australian Medical Council licensing examination practice questions.3 The questions had a five‐option multiple‐choice format, and we copied each one in full to the LLM. We omitted images during testing because only text inputs were accepted. Two medical student investigators (OK and CG) decoded responses into answer options, which the Australian Medical Council website graded. Each algorithm was tested three times in independent sessions.
ChatGPT‐3.5 and ChatGPT‐4 answered 49/50 questions in every session, and New Bing answered 49, 46 and 43 questions in separate sessions (). ChatGPT‐4 provided 46/50 answer options that were identical in each session. Conversely, ChatGPT‐3.5 gave only 37/50 identical answers, and New Bing, 33/50 (Box). ChatGPT‐4 scored the highest mean (39.7; standard deviation [SD], 0.6), and the means for New Bing and ChatGPT‐3 were 36.0 (SD, 2.6) and 33.0 (SD, 2.6) respectively.
GPT‐4‐derived LLMs (ChatGPT‐4 and New Bing) appear to exceed the medical multiple‐choice performance of their GPT‐3‐derived predecessors (ChatGPT‐3.5). Such models have previously been shown to outperform PubMedGPT,4 GPT‐3 without fine‐tuning, and InstructGPT.5 Despite improvements in accuracy, limitations in consistency remain. These results indicate the potential influence of fine‐tuning and web‐search access for LLM performance. Caution is advised when choosing whether to use an LLM, and which one to use, for a specific task, as it cannot be assumed that all same‐generation LLMs perform equally.
Box – Large language model performance on the Australian Medical Council's licensing examination
Large language model |
ChatGPT‐3.5 |
ChatGPT‐4 |
New Bing | ||||||||||||
Total number of questions |
50 |
50 |
50 |
||||||||||||
Raw score |
|||||||||||||||
Repeat 1 |
36 |
39 |
35 |
||||||||||||
Repeat 2 |
31 |
40 |
34 |
||||||||||||
Repeat 3 |
32 |
40 |
39 |
||||||||||||
Correct answers, mean (SD) |
33.0 (2.6) |
39.7 (0.6) |
36.0 (2.6) |
||||||||||||
Total correct answers |
99/150 |
119/150 |
108/150 |
||||||||||||
Overall accuracy |
66.0% |
79.3% |
72.0% |
||||||||||||
SD = standard deviation. | |||||||||||||||
Competing interests
Acknowledgements
We thank Joshua Kovoor for providing editorial and statistical support to this piece.
References
- Nori H, King N, McKinney SM, et al. Capabilities of GPT‐4 on medical challenge problems [preprint]. arXiv 230313375; 20 Mar 2023. https://doi.org/10.48550/arXiv.2303.13375 (viewed Mar 2023).
- OpenAI. GPT‐4 technical report [preprint]. ArXiv 2303.08774; 15 Mar 2023. https://doi.org/10.48550/arXiv.2303.08774 (viewed Mar 2023).
- Australian Medical Council Limited. MCQ trial examination [website]. Canberra: AMC, 2022. https://www.amc.org.au/assessment/mcq/mcq‐trial/ (viewed Mar 2023).
- Kung TH, Cheatham M, Medenilla A, et al. Performance of ChatGPT on USMLE: potential for AI‐assisted medical education using large language models. PLOS Digit Health 2023; 2: e0000198.
- Gilson A, Safranek CW, Huang T, et al. How does ChatGPT perform on the United States Medical Licensing Examination? The implications of large language models for medical education and knowledge assessment. JMIR Med Educ 2023; 9: e45312.