Artificial intelligence and medical imaging: applications, challenges and solutions
Authors: Meng Law, Jarrel Seah and George Shih
Published online: 7 June 2021
AI‐based tools can help with image acquisition, reconstruction and quality; interpretation, diagnosis and decision support; and manual tasks
Artificial intelligence (AI) is having a disruptive impact in many areas, including health care. In medicine, machine learning (ML) techniques have existed for decades but were mostly not adopted. New deep learning techniques, along with copious medical imaging and digital health data, now provide standardised, reproducible, dependable and accurate diagnostic reports. These can only improve patient care and safety, enhancing the practice of clinical medicine. However, a number of challenges have arisen, hindering progress and more widespread application. In this article, we describe current AI/ML tools in medical imaging, discuss the major challenges facing the field, and offer some potential solutions.
Currently approved tools
Currently, there are 106 commercial ML applications which are approved by the relevant authorities in the United States, Australia, Korea, Japan and Europe. Almost half of these apply to neuroimaging or chest imaging, with more than triple the number of approvals granted in 2018 compared with 2017.
Diagnosis and therapy in medicine is still relatively subjective, with a human doctor providing a diagnosis often supported by a human radiologist generating a radiology report. This is why we often seek multiple opinions and host multidisciplinary meetings to discuss the large amount of digital health data available to us, suggest the most likely diagnoses, and recommend what we think is the best therapy for the optimal outcome.1 Not only is this not always reproducible but the outcome for any given patient is variable. Using AI to perform time consuming review of images will allow radiologists to perform more value‐added tasks, becoming more valuable to patients and clinicians in multidisciplinary clinical settings. Publications on AI dramatically increased from 100–150 per year in 2007–2008 to 700–800 per year in 2016–2017.2
Medical imaging tools which employ AI can be divided into three broad categories: image acquisition, reconstruction and quality; image interpretation, diagnosis and decision support; and helping doctors with manual tasks.
Image acquisition, reconstruction and quality
AI for image acquisition, reconstruction and quality3,4 involves algorithms that accelerate the creation and improve the quality of medical images (often using decreased radiation, contrast dose or decrease acquisition time). Examples include:- reproducible patient positioning for a follow‐up scan;
- automatic magnetic resonance imaging (MRI) protocols based on feature detection;
- MRI acquisition acceleration methods and sampling of k‐space data to decrease scan times and improve signal‐to-noise ratios;
- deep learning reconstruction for computed tomography (CT) and MRI motion correction;
- dose and time reduction for positron emission tomography reconstruction; and
- reduction in intravenous contrast volume in MRI and CT.
Image interpretation and classification
AI for image interpretation typically involves pixel‐based algorithms that perform tasks such as lesion detection, localisation and segmentation/measurement. Some current applications include:- image segmentation (brain/hippocampal volumes), labelling of spinal levels, comparison of multiple and previous examinations (multiple sclerosis lesions, oncology/metastases follow‐up);
- quantification of pathological biomarkers (eg, liver iron quantification from MRI or B‐type natriuretic peptide from chest x‐rays);
- prioritisation of abnormal scans to expedite reporting, reducing report turnaround times;
- detection of spine and rib fractures, intracranial or other haemorrhage, pulmonary embolism, free abdominal air, bowel obstruction, pneumothorax, cardiomegaly, aortic dissection/aneurysm, and breast lesions;
- automated quality control of radiographer performance;
- natural language processing (NLP) to review reports for abnormality detected on imaging (eg, lung nodules);
- internal automated peer review; and
- generation of chest x‐ray and other imaging reports.
Helper tools
The third category comprises tools that help with currently manual tasks (eg, selecting the correct hanging protocol while displaying x‐rays, or finding similar pathology from the imaging archive) and do not directly assist interpretation or decisions. Examples include:- NLP tools for querying reports and other text‐based medical records, or for generating x‐ray report drafts from other AI model outputs;
- clinical decision support for ordering, interpreting and defining management;
- peer review of trainee (registrar/resident) reports;
- tracking of radiation dosimetry (CT), gadolinium dosage (MRI) and renal function (injection of contrast for CT and MRI); and
- NLP for clinical trials matching.
Deep learning algorithms
Deep learning algorithms for image acquisition and reconstruction have been developed and validated, and have received regulatory clearance. These are beginning to be widely implemented as a natural extension of existing reconstruction algorithms, requiring no alteration in workflow.
Triage algorithms
In contrast, triage algorithms that alter the diagnostic process (eg, by prioritising abnormal scans to be reported first or highlighting abnormalities on images) are novel and require workflow changes to realise their full potential. As a result, these have had slower uptake. Currently, most approved products are triage algorithms. They retain a narrow scope of findings, typically focusing on the common and urgent findings. This reflects the increasing scrutiny that these algorithms will face from end‐users as well as regulatory bodies, complicating the development of products claiming to identify a wide variety of findings.
Notably, in 2020, the US Centers for Medicare and Medicaid Services approved a new technology add‐on payment for the use of an AI algorithm looking for large vessel occlusion on CT angiograms of the head. While there are stringent criteria required for this payment (which make the actual financial implications of this decision complicated and beyond the scope of this article), this reimbursement pathway indicates recognition of the value that AI algorithms can add.
Image segmentation and quantification algorithms
Image segmentation and quantification algorithms promise non‐invasive determination of quantities previously only obtainable by biopsy or autopsy, such as liver iron concentrations. The greatest challenge faced by these algorithms is the scarcity of gold standard labels, as these are typically obtained via pathological analysis of tissue. With the development of Transformer neural network architecture in 2017 enabling training on huge datasets and models, and its subsequent popularisation with Google’s BERT (Bidirectional Encoder Representations from Transformers) model, the NLP field is set for a renaissance, and many algorithms are currently being developed to parse or even generate radiology reports. While none of these algorithms are likely to be cleared by regulatory authorities anytime soon, they are likely to have widespread future application and adoption. It is clear that the field is evolving quickly, and other applications will be developed and come to market over the next few years.
Current challenges
Development of AI/ML software for medical imaging requires a large dataset consisting of inputs as well as gold standard labels, split into training, testing and validation sets. A major challenge is the safety, efficacy and regulatory governance of software and devices. In 2018, the US Food and Drug Administration (FDA) cleared the first AI/ML‐based software (a program for diabetic retinopathy) that provides screening decisions without needing clinician interpretation (https://www.fda.gov/news-events/press-announcements/fda-permits-marketing-artificial-intelligence-based-device-detect-certain-diabetes-related-eye). Although such technologies hold promise, they also raise concerns regarding safety and effectiveness. In April 2019, the FDA announced that it was reviewing how to regulate AI/ML‐based software,1 which requires unique regulatory approaches that span the life cycle of the technologies, allowing necessary improvements while ensuring that the algorithm is safe.5 New FDA guidance requires careful review of the safety and effectiveness of such software, consideration of the allowable post‐approval modifications to the software, and review of unanticipated divergence in the software’s eventual performance from the originally approved product.6 In Australia, the Therapeutic Goods Administration has also recognised the increasingly important role that AI/ML‐based software will play in patient care, with up‐classification of software‐based medical devices in August 2020 (Therapeutic Goods Legislation Amendment (2019 Measures No. 1) Regulations). AI software is becoming increasingly important in medical devices and clinical adoption, so up‐classification of software into the device category may require more stringent regulatory review and approval (https://www.tga.gov.au/regulation-software-based-medical-devices), with regulatory changes in place from 25 February 2021.
Almost all AI/ML‐based software now requires human or near‐human performance, with increasing recognition by both deep learning researchers and medical professionals that such software can perform very differently in subpopulations such as inpatients versus outpatients, or in the presence of confounding factors. The prototypical case is a pneumothorax detection algorithm learning to identify the presence of a chest tube instead of an actual pneumothorax.
These regulatory considerations are evolving as they try to keep abreast of the advances in technology, computational power, ethical challenges of data ownership, data sharing, privacy and safety, and the commercialisation of the many products coming to market.5 Protected health information (PHI) has been a challenge in the development, implementation and dissemination of AI tools in medical imaging. Despite efforts to remove PHI from DICOM image data, PHI can be hidden in the metadata or pixel data (burned‐in annotations). Three‐dimensional reconstruction of facial features can potentially re‐identify patients, as can personal objects (eg, jewellery) on a patient’s image.7 An organisation must be extremely careful in curating and sharing these data in the development of AI algorithms.8 Another important challenge beyond the scope of this article is the ownership of these data. Patients may have provided consent for their data to be used for patient care, academic and research purposes. However, when these data are utilised to generate algorithms for commercial purposes, one could argue that the patient never consented to their data being used to generate revenue for one of the many AI start‐ups.
Other challenges facing the research, development and translation of AI in medical imaging include:- developing suitable performance metrics that align with clinical utility (typically based on the area under the receiver operating curve, positive and negative predictive value);
- preprocessing/curation of data and labelling/annotation of large training datasets (manually intensive and time consuming); and
- application and generalisability of the algorithm to different image protocols (slice thickness, resolution) not only between organisations but within institutions with multiple scanner types and differing clinical settings.7
Solutions
Over the past 4 years, the Radiological Society of North America (RSNA) has hosted public AI challenges on bone age (2017), pneumonia (2018), intracranial haemorrhage (2019) and pulmonary embolism (2020). In 2020, the Royal Australian and New Zealand College of Radiologists (RANZCR) hosted a similar event. The RANZCR catheter and line position challenge was completed in March 2021. These AI competitions have provided solutions to some of the challenges listed above. In an ongoing pandemic, the RANZCR challenge found AI solutions for detecting the location of central lines and nasogastric and endotracheal tubes on chest x‐rays, which are important for patients requiring hospitalisation and intensive care. De‐identification of medical images requires a thorough process to ensure that PHI is removed, and our industry will need to continue to improve these techniques to ensure the safe progress of AI development for both academia and industry. The RANZCR challenge employed a subset of the 112 000 chest radiographs from the National Institutes of Health Clinical Center.7 DICOM images were initially converted to PNG format (PHI removed) before release to the public and then converted back to DICOM format, a more familiar format for the competition. Another strategy to protect patient privacy is federated learning — retaining the datasets within each organisation’s firewall and only sharing the trained algorithms with other sites. The intent is to expedite model validation when applied to different settings or when there is a need for more datasets and more heterogeneous datasets.
For the RANZCR challenge, the data were curated and preprocessed. We labelled and annotated the abnormalities in the x‐rays, utilised NLP to curate the dataset, and engaged radiologists to perform the annotations. This required not only consensus definitions for the labelling but also annotation of x‐rays by two to three readers.
AI challenges advance the field by involving multiple groups working in parallel, rather than serial individual groups performing hypothesis‐driven research which is in turn reproduced and validated by the next group. In essence, this speeds up medical discovery and translation into the clinic. The contribution of the RSNA and RANZCR challenges to AI research and implementation has been enormous, not only in identifying some of the challenges facing AI but also through open sourcing the models for public validation and implementation.7,9
These large academic collaborations and AI challenges have many benefits:- creation of a new public annotated dataset (even if data was already public, annotations are new) with expert physicians;7,10
- ML experts (who may not otherwise be involved in medical AI) working on a medical AI problem;
- open source availability of winning solutions (code and model weights);
- validation of datasets that future algorithms can use to benchmark against as new AI algorithms are developed and improved upon; and
- potential commercialisation of algorithms for specific diseases and clinical scenarios that academic societies have deemed important.
In future, these new public datasets will need to undergo a rigorous process for de‐identification to ensure removal of PHI, including DICOM metadata as well as DICOM pixel data. The academic community will need to improve existing processes and tools to help ensure that these new open datasets are as private and non‐identifying as possible. Even with creation of many open annotated datasets, it will be impossible to allow sharing of digital health data globally. Other forms of collaboration will still need to occur, aided by new advancements in technology. Federated learning and other technological advances (including differential privacy and homomorphic encryption) will provide other options for creating high quality AI without necessarily making the data public, and new generative adversarial network approaches will augment training datasets to further improve clinical diagnoses.11,12
Competing interests
References
- Elshafeey N, Kotrotsou A, Hassan A, et al. Multicenter study demonstrates radiomic features derived from magnetic resonance perfusion images identify pseudoprogression in glioblastoma. Nat Commun 2019; 10: 3170.
- Pesapane F, Codari M, Sardanelli F. Artificial intelligence in medical imaging: threat or opportunity? Radiologists again at the forefront of innovation in medicine. Eur Radiol Exp 2018; 2: 35.
- Kim H, Irimia A, Hobel SM, et al. The LONI QC system: a semi‐automated, web‐based and freely‐available environment for the comprehensive quality control of neuroimaging data. Front Neuroinform 2019; 13: 60.
- Duffy BA, Zhang W, Tang H, et al. Retrospective correction of motion artifact affected structural MRI images using deep learning of simulated motion. Proceedings of 1st Conference on Medical Imaging with Deep Learning; 4‐6 July 2018; Amsterdam. https://openreview.net/pdf?id=H1hWfZnjM (viewed Apr 2021).
- Hwang TJ, Kesselheim AS, Vokinger KN. Lifecycle regulation of artificial intelligence‐ and machine learning‐based software devices in medicine. JAMA 2019; 322: 2285–2286.
- US Food and Drug Administration. Artificial Intelligence/Machine Learning (AI/ML)-Based Software as a Medical Device (SaMD) Action Plan. January 2021. Silver Spring, MD: FDA. 2021. https://www.fda.gov/medical-devices/software-medical-device-samd/artificial-intelligence-and-machine-learning-software-medical-device (viewed Apr 2021).
- Shih G, Wu CC, Halabi SS, et al. Augmenting the National Institutes of Health Chest Radiograph Dataset with expert annotations of possible pneumonia. Radiol Artif Intell 2019; 1: https://doi.org/10.1148/ryai.2019180041.
- Chervenak AL, Van Erp TGM, Kesselman C, et al. A system architecture for sharing de‐identified, research‐ready brain scans and health information across clinical imaging centers. Stud Health Technol Inform 2012; 175: 19–28.
- Prevedello LM, Halabi SS, Shih G, et al. Challenges related to artificial intelligence research in medical imaging and the importance of image analysis competitions. Radiol Artif Intell 2019; 1: https://doi.org/10.1148/ryai.2019180031.
- Filice R, Stein A, Wu C, et al. Crowdsourcing pneumothorax annotations using machine learning annotations on the NIH chest x‐ray dataset. J Digit Imaging 2020; 33: 490–496.
- Seah JCY, Tang JSN, Kitchen A, et al. Chest radiographs in congestive heart failure: visualizing neural network learning. Radiology 2019; 290: 180887.
- Xing Y, Ge Z, Zeng R, et al. Adversarial pulmonary pathology translation for pairwise chest x‐ray data augmentation. In: Shen D, Liu T, Peters TM, et al, editors. Proceedings of 22nd International Conference on Medical Image Computing and Computer‐Assisted Intervention 2019. Lecture Notes in Computer Science, Vol. 11769. Cham, Switzerland: Springer, 2019.
Provenance: Commissioned; externally peer reviewed.
Reducing Nitrous Oxide Emissions Across the Melbourne Biomedical Precinct
Ross Robertson, Andrew Downey, Daryl Williams, Bjorn Makein, Ben Dunne, Tugce Ozturk, Ying Gu, Rebecca McIntyre
Multimorbidity Clusters Among People Aged 65 Years and Over in Australia: A Nationwide Cross-Sectional Data Linkage Study
Weisi Chen, Christine Y. Lu, Sarah N. Hilmer, Alice A. Gibson, Edwin C. K. Tan
When ‘Liver Enzymes’ Are Not Hepatic: Late-Onset Pompe Disease
Shauna Madigan, Georgina England, Wayne Rankin
West Nile virus Kunjin subtype in rural NSW
Emily Gibson, Megan Whitley, Peter Murray, Linda Hueston, Jane Bennett, Raguharan Kathiresu, David N Durrheim
The impact of the Breast Screen NSW transition from film to digital mammography, 2002–2016: a linked population health data analysis
Rachel Farber, Nehmat Houssami, Katy J L Bell