Volume 209 - Issue 5

Whole genome sequencing provides better diagnostic yield and future value than whole exome sequencing

Authors:  John S Mattick, Marcel Dinger, Nicole Schonrock and Mark Cowley

Med J Aust 2018; 209 (5): 197-199. || doi: 10.5694/mja17.01176
Published online: 9 April 2018

The higher cost of whole genome sequencing is justified through better diagnostic yields and the potential for future analysis

The integration of genome sequencing with clinical records and data from the internet of things will transform health care

There is a great deal of optimism about the potential of genomics to transform medicine and health care. That optimism is justified. Indeed, it is hard to imagine a future where personal genomic information is not consulted routinely at the point of care. Every one of us is different, with personal genetic idiosyncrasies and risks — of cancer, cardiac arrest, blood clots, emphysema, diabetes, arthritis or toxic reactions to medications, among many others; the list will only continue to grow. Knowledge of individual genetic variation will change medicine from the art of crisis response to the science of health management, with huge benefits, both individually and systemically. It will also create new enterprises at a time of rapid change in the largest and fastest growing industry in the world.

This transformation is underway but raises several issues en route from the old world to the new. The first is whether to undertake whole exome sequencing (WES) or whole genome sequencing (WGS). For the uninitiated, exome is the term used to describe the protein-coding portion of the genome, which, while comprising less than 1.5% of the total genome, is the location of the majority of the mutations that cause severe developmental or cognitive disabilities and disorders such as cystic fibrosis and thalassaemia. Of course, proteins are the key components of cells and any damage to them causes serious, if not catastrophic, problems, which are sometimes referred to as monogenic or Mendelian conditions.

The rest of the genome has been traditionally thought to be (mostly) non-functional, despite the fact that it produces a plethora of RNA molecules shown to play important roles in organising human development and cell biology.1 The misunderstanding of the genetic programming of humans and other complex organisms that has its roots in the mechanical view of the world that held sway in the 19th century and in the first half of the 20th century is a story for another day. However, it is already clear that the genetic variation that underpins complex traits and diseases lies overwhelmingly outside of the protein-coding sequences,2 in the rest of the genome, where the regulatory information that controls gene expression is contained.

For individuals seeking diagnosis of acute monogenic disorders, and for those who hold that most of the genome is irrelevant, WES is favoured, mainly because of price: while technically more complex, WES is cheaper because much less data need to be generated and analysed, which is appealing in the context of limited research and clinical service budgets. In addition, testing laboratories in genetic pathology services and paediatric hospitals, such as the Victorian Clinical Genetics Service (www.vcgs.org.au) are challenged to acquire and operate the large scale equipment and develop the bioinformatic pipelines to undertake cost-effective WGS and analysis. In Australia, clinically accredited WGS (ISO15189) is only available through Genome.One (www.genome.one), and even internationally availability is limited mainly to the United States, through companies such as Illumina TruGenome Clinical Sequencing Services (https://sapac.illumina.com/clinical/illumina_clinical_laboratory/trugenome-clinical-sequencing-services.html) or Broad Institute Genomic Services (http://genomics.broadinstitute.org/products/clinical-research-sequencing).

On the other hand, while WGS is presently more expensive, it polls more of the exome and gives a higher diagnostic return because it is less susceptible to technical distortions and blind spots, such as those caused by deletions, rearrangements, copy number variations, repetitive elements and pseudogenes, all of which are involved in genetic disease, disease risk, quantitative trait variation and diagnostic assessment.3-5 Indeed, the major limitations of WES are the uneven capture or complete skipping of exons due to technical limitations of target-probe hybridisation and/or high G+C content. These limitations results in inadequate sequencing depth for some regions and subsequent poor accuracy scores of variants that are then dismissed or missed completely. In general, the more steps involved in the preparation of DNA for sequencing the more bias is introduced, thus increasing the rate of errors and false positives.6

WGS also makes no assumptions about what constitutes the exome. Indeed, deep RNA analysis shows that the human gene complement is far from being fully characterised.7 A recent high resolution study of the transcriptional output from human chromosome 21 RNA identified over 2000 unannotated isoforms of protein-coding mRNAs, including almost 300 previously unknown protein-coding exons.8 Therefore, given the incomplete characterisation of human gene structure, it is clear that WES is intrinsically inadequate and biased towards the currently incompletely known catalogue of protein-coding exons.

The Box shows that, over a large number of studies, the diagnostic rate achieved for WES is generally in the range of 25–35% (weighted average, 28%). In contrast, WGS, albeit on lower numbers, is generally in the range of 40–60% (weighted average, 49%), nearly doubling the rate of diagnosis.

A caveat is that the range of conditions and cohorts analysed to date by WES and WGS were different and may have intrinsically different diagnostic rates, as well as having employed different protocols and diagnostic criteria, which makes strict comparison difficult. Nonetheless, the single intrastudy comparison on intellectual disability showed an about two-fold better performance of WGS (42–62%) over WES (27%),9 similar to the differences in the weighted average diagnostic rate. It is to be acknowledged that the technology underpinning WES has improved a little, giving better exon coverage. This improvement is reflected in the WES publications from 2017–18 compared with 2013–14, which show an increase in the weighted average diagnostic yield from 26% to 31% (Box).

The comparative value proposition of WES and WGS approaches for the diagnosis of serious genetic disorders is debated, but the value proposition needs to consider not only the numbers of successful diagnoses per unit cost, but also the lifetime savings and social benefits achieved by successful diagnoses. On the latter basis alone, WGS is likely to provide more value for money than WES, particularly in those cases where the diagnosis leads to an effective treatment or reduction of the risk of reoccurrence in affected families.

The crucial difference and major upside is that WGS generates a complete dataset of an individual’s genetic compendium. This is a complete dataset that can not only be interrogated and re-interrogated for clinical use throughout an individual’s lifetime in a far more comprehensive way than exome data but also serves as an enormously valuable resource to grow our understanding of the relationship between genomic information and our biology, especially in relation to complex traits and disorders. This advantage accrues immediately even if WGS technology improves, as it inevitably will, to allow longer reads, haplotype phasing, and assessment of epigenetic marks. That is, in any case, an exome enables an analysis that is intrinsically limited to present-day knowledge, whereas WGS is an investment in science and in the future.

For these reasons the whole genome, not just the protein-coding portion, was sequenced in the human genome project. One cannot hope to understand genomic information or its influence on human biology without a complete picture, especially when considering that most of the genetic variation underpinning both simple and complex disease lies in non-coding regulatory sequences.2

The issue is whether the payers for genetic testing (public and private health systems and insurers) perceive better value in WGS than WES, when all factors are taken into account, as we suggest they should. We do not suggest that current clinical genetic budgets should be stretched to enable future research, but rather that increased resources for genome sequencing in clinical contexts should be provided to allow both improved diagnostic rates, with consequent better lifetime outcomes,38 and the ability for productive re-analysis, including for reasons that may be independent of the motivation for doing the sequencing in the first place.

For the same reasons, the 100 000 genomes project being undertaken by Genomics England, owned by the National Health Service, employs WGS. While the cost of WGS might currently be three times higher than exome sequencing, the difference will inevitably converge due to continual advances in sequencing technology. Integration of WGS with clinical records and data from the internet of things, including wearable devices to monitor physiological indices and environmental exposures, will form a new information ecology that will transform the management of health, both individually and societally. As a result, health care will evolve from the last of the great cottage industries to the most important of the data-intensive industries of the 21st century.

Box – Results of whole exome sequencing (WES) and whole genome sequencing (WGS) studies in patients with a variety of disorders

Test

Disorder

n

Diagnosis rate

Study, year

References


WES

Severe non-syndromic ID

51

50%

Rauch et al, 2012

10

Severe ID (IQ < 50)

100

16%

de Ligt et al, 2012

11

Consecutive patients: 85% neurological disorders in patients aged under 18 years

250

25%

Yang et al, 2013

12

Consecutive patients: DD in 37%

814

26%

Lee et al, 2014

13

Severe ID (IQ < 50)

100

27%

Gilissen et al, 2014

9

Consecutive patients: neurodevelopmental disorders

2000

25%

Yang et al, 2014

14

Moderate or severe non-syndromic ID

41

29%

Hamdan et al, 2014

15

Neurodevelopmental disorders

78

48%

Srivastava et al, 2014

16

Consecutive patients, ID or DD in 64%, 84% in patients aged under 18 years

500

30%

Farwell et al, 2015

17

Developmental disorders

1133

27%

Wright et al, 2015

18

Neuromuscular disorders

45

47%

Todd et al, 2015

19

Consecutive patients with any suspected Mendelian disorder

75

29%

Lazaridis et al, 2016

20

Solid tumours

121

27%

Parsons et al, 2016

21

Consecutive adult patients with any suspected Mendelian disorder

486

18%

Posey et al, 2016

22

Various paediatric disorders

80

58%

Stark et al, 2016

23

Consecutive patients with any disorder, ID or DD in 52%

3040

29%

Retterer et al, 2016

24

Paediatric neurological disorders

50

48%

Nolan, Carlson, 2016

25

Consecutive patients with any disorder, after extensive diagnostic odyssey

362

29%

Sawyer et al, 2016

26

Paediatric neurological disorders

57

49%

Kuperberg et al, 2016

27

Intellectual disability

17

29%

Monroe et al, 2016

28

Consecutive patients, nervous system abnormality in 77%

1000

31%

Trujillano et al, 2017

29

Limb girdle muscular dystrophies

104

37%

Harris et al, 2017

30

Paediatric neurological disorders

150

29%

Vissers et al, 2017

31

Neurodevelopmental, neurometabolic disorders and dystonias

72

35%

Evers et al, 2017

32

Fetal structural abnormalities

196

24%

Fu et al, 2017

33

Any suspected Mendelian disorder

61

52%

Tan et al, 2017

34

Paediatric dilated cardiomyopathy

15

50%

Long et al, 2017

35

Chronic kidney disease

92

24%

Lata et al, 2018

36

Neurological disorders

40

40%

Cordoba et al, 2018

37

Weighted average, mean (SD)

 

28% (11%)

 

 

WGS

Intellectual disability

50

42–62%

Gilissen et al, 2014

9

Neurodevelopmental disorders

15

73%

Soden et al, 2014

38

Autism spectrum disorders

170

42%

Yuen et al, 2015

39

Any suspected Mendelian disorder, ICU care

35

57%

Willig et al, 2015

40

Autosomal dominant polycystic kidney disease

28

86%

Mallawaarachchi et al, 2016

4

Inherited retinal disease

46

52%

Ellingford et al, 2016

41

Any suspected Mendelian disorder

22

36%

Bick et al, 2017

42

Various paediatric disorders

103

41%

Lionel et al, 2017

43

Hypertrophic cardiomyopathy

41

54%

Cirino et al, 2017

44

Weighted average, mean (SD)

 

49% (15%)

 

 


ICU = intensive care unit. ID = intellectual disability. DD = developmental delay. SD = standard deviation.


Authors


Competing interests


Acknowledgements


References


Provenance: Not commissioned; externally peer reviewed.