An overview of the GRADE approach and a peek at the future
Authors: Waleed Alhazzani and Gordon Guyatt
Published online: 1 October 2018
Worldwide, over 100 organisations are using the Grading of Recommendations Assessment, Development and Evaluation approach; it is thus essential that clinicians using formal guidelines become familiar with the GRADE approach
The Grading of Recommendations Assessment, Development and Evaluation approach provides a clear interpretation of the quality of evidence for key stakeholders
In 2004, a group of international experts in methodology and practice guidelines first published the Grading of Recommendations Assessment, Development and Evaluation (GRADE) approach1 to assess the quality of evidence supporting medical interventions and develop recommendations. Since then, the group has published a six-part series for clinicians using GRADE guidelines,2 and a series of articles supporting systematic review authors and guideline groups using GRADE in their work.3
Since its introduction, over 100 organisations worldwide — including the World Health Organization, the Cochrane Collaboration, the Joanna Briggs Institute, the American College of Physicians, DynaMed Plus and UpToDate — have endorsed or adopted GRADE. It is now difficult for any clinician using formal guidelines, or recommendations in online texts such as UpToDate, to avoid confronting GRADE recommendations.
The enthusiasm for the GRADE approach represents a response to its considerable strengths:
-
providing carefully developed and comprehensive criteria to rate the quality of evidence (also known as certainty or confidence in evidence);
-
making this quality rating on the basis of systematic summaries of the entire body of relevant studies rather than selected individual studies;
-
making the quality rating for each relevant outcome;
-
moving from evidence to recommendations, evaluating the importance of outcomes from a patient perspective, including both benefit and harm related to the interventions under consideration;
-
using a transparent process to move from evidence to recommendations that considers the magnitude of benefits and harms, the associated quality of evidence, patient values and preferences and costs;4 and
-
providing clear interpretation of the strength of recommendations for key stakeholders (clinicians, patients and policy makers).
The Box outlines the GRADE approach to rating the quality of the evidence summarised in systematic reviews of the best available evidence. This approach classifies evidence quality into high, moderate, low and very low.5 In this framework, randomised controlled trials (RCTs) start at high quality, and observational studies (such as cohort and case-control studies) start as low quality. There are, however, a number of limitations that can result in rating down evidence quality from a summary of relevant RCTs.
First, the design and execution of the RCTs may be flawed — a problem GRADE labels as risk of bias.6 For instance, the RCTs may have failed to conceal randomisation, failed to blind patients and clinicians, or failed to follow a large proportion of their patients to the study’s conclusion.
Second, the studies may differ in their results (some showing clear benefit, others no benefit or harm for the same outcome), and systematic review authors may have failed to find an explanation for the difference. GRADE characterises this limitation as inconsistency of results.7
Third, the studies may have included only a small number of patients, and the confidence intervals around pooled results may, therefore, be wide — a limitation GRADE labels as imprecision.8 Fourth, the patients included, the interventions administered or the outcomes measured may differ from those of primary interest — a limitation GRADE labels as indirectness.9 Finally, GRADE notes the problems of failure to publish RCT results, and includes a provision for rating down for publication bias (Box).
GRADE also allows for rating up the quality of evidence from observational studies. We have no RCTs of insulin for diabetic ketoacidosis, dialysis for terminal renal failure, or hip replacement for severe osteoarthritis. Nor should we, since the response to these treatments is sufficiently large, rapid and consistent that it leaves no doubt as to the very considerable associated benefits. Therefore, GRADE would classify the evidence as high quality. It also provides for rating up quality from observational studies for a clear dose–response gradient, or if all plausible confounders, if present, would only support an inference of benefit, harm or lack of effect.10
Thus, following the GRADE approach, we will have high confidence in the evidence when estimates are based on well designed RCTs, with consistent, precise results, addressing our question of interest, and free from publication bias. On the other hand, we will lose confidence in the evidence if it comes from RCTs at high risk of bias; from RCTs with imprecise or inconsistent results; from RCTs that enrolled patients, tested interventions or measured outcomes that differ from those in which we are primarily interested; or if we have a high suspicion of publication bias (Box).
GRADE provides two categories of recommendations: strong and weak (also known as conditional or contingent). It is logical that guideline panels will be much more inclined to recommend an intervention when the evidence is high or at least moderate quality than if it is low or very low. Similarly, guideline developers are much more likely to make a strong recommendation if the benefits far outweigh the harms than if they are closely balanced. That judgment of balance requires an understanding of patients’ beliefs and preferences, cost, and often of effects on equity, feasibility and acceptability.11
Strong and weak recommendations come with clear interpretations. If a guideline panel has made appropriate judgments, all or almost all fully informed patients will choose a strongly recommended intervention. Strong recommendations therefore represent “just do it” situations. If, however, the recommendation is weak, it signifies that although the majority of fully informed patients will choose the recommended intervention, an appreciable minority will not — an “it depends” situation. Weak recommendations, therefore, demand shared decision making to ensure that the course of action is best for that particular patient.
Anyone producing a systematic review should have the expertise necessary to apply GRADE. Indeed, the Cochrane Collaboration requires reviewers to produce GRADE summary of findings tables that outline the quality of evidence and pooled results for all patient-important outcomes, and GRADE inter-rater reliability is high when reviewers are appropriately trained.12 While the process of gathering and summarising the evidence — that is, producing a rigorous systematic review — can be cumbersome and time consuming, once appropriately summarised, reviewers can make GRADE quality of evidence judgments easily. Therefore, clinicians are entitled to expect GRADE quality of evidence ratings from the systematic reviews that guide their practice.
GRADE’s initial focus has been on the methodology of assessing the quality of evidence from RCTs and observational studies of interventions, and on moving from evidence to recommendations. More recently, GRADE has provided substantial guidance for questions of diagnosis,13 and has begun to apply the GRADE approach to issues of prognosis14 and to qualitative evidence.15 More work in the latter two areas remains, as does the development of GRADE guidance in rare diseases and animal or laboratory studies.
In addition to further developments in areas of prognosis, summarising qualitative evidence, rare diseases and animal studies, we can anticipate additional advances in GRADE methodology. The GRADE Working Group includes a number of project groups, each focusing on a specific area to develop the appropriate GRADE guidance. Of particular interest is GRADE’s work in network meta-analysis (NMA). NMA is a burgeoning area that extends conventional meta-analysis to the simultaneous consideration of multiple treatments. Initially, NMA practitioners more or less ignored the issue of the quality of the evidence emerging from their analyses. GRADE has now begun to fill the gap, and has produced initial guidance in rating quality of evidence from NMAs.16 Doing so continues to present challenges that the GRADE NMA project group will continue to address.
GRADE is likely to play a variety of roles in the broader evidence ecosystem. Because of the high demand for trained users of GRADE and to ensure adherence to quality standards, the GRADE Working Group is in the process of developing credentialling criteria and process. To enhance the efficiency of systematic reviews and GRADE application, innovative approaches in the future could include crowd outsourcing (http://crowd.cochrane.org/index.html), using artificial intelligence,17 and conducting rapid reviews with the production of guidelines that facilitate rapid incorporation of practice-changing evidence. An example of this last initiative is the Rapid Recommendations collaboration between The BMJ and the Making GRADE the Irresistible Choice (MAGIC) group, which is using GRADE methodology to rapidly produce a small number of recommendations addressing a specific clinical context in response to potentially practice-changing RCTs.
Other initiatives that are part of the GRADE context include patient engagement in the process of assessing the evidence and developing recommendations. Current efforts focus on developing rigorous methodology to engage patients in this process.
Lastly, new knowledge translation platforms (eg, https://metaclinician.com) can help measure clinicians’ behaviour and understanding of recommendations, identify barriers to implementations and facilitate clinicians’ use of best GRADE evidence.
Box – Quality assessment criteria
|
Study design |
Confidence in estimates |
Lower if |
Higher if |
||||||||||||
|
|
|||||||||||||||
|
Randomised controlled trials |
High |
Risk of bias:
|
Large effect:
|
||||||||||||
|
Moderate |
Inconsistency:
|
Dose response:
|
|||||||||||||
|
Observational studies |
Low |
Indirectness:
|
All plausible confounding:
|
||||||||||||
|
Very low |
Imprecision:
|
|
|||||||||||||
|
Publication bias:
|
|
||||||||||||||
|
|
|||||||||||||||
|
|
|||||||||||||||
Competing interests
Waleed Alhazzani and Gordon Guyatt are owners of MetaClinician. Gordon Guyatt is a consultant for UpToDate.
References
- Atkins D, Best D, Briss PA, et al; GRADE Working Group. Grading quality of evidence and strength of recommendations. BMJ 2004; 328: 1490.
- Guyatt GH, Oxman AD, Vist GE, et al; GRADE Working Group. GRADE: an emerging consensus on rating quality of evidence and strength of recommendations. BMJ 2008; 336: 924-926.
- Guyatt G, Oxman AD, Akl EA, et al; GRADE guidelines: 1. Introduction-GRADE evidence profiles and summary of findings tables. J Clin Epidemiol 2011; 64: 383-394.
- Andrews JC, Schünemann HJ, Oxman AD, et al; GRADE guidelines: 15. Going from evidence to recommendation-determinants of a recommendation’s direction and strength. J Clin Epidemiol 2013; 66: 726-735.
- Balshem H, Helfand M, Schünemann HJ, et al; GRADE guidelines: 3. Rating the quality of evidence. J Clin Epidemiol 2011; 64: 401-406.
- Guyatt GH, Oxman AD, Vist G, et al; GRADE guidelines: 4. Rating the quality of evidence — study limitations (risk of bias). J Clin Epidemiol 2011; 64: 407-415.
- Guyatt GH, Oxman AD, Kunz R, et al; GRADE guidelines: 2. Framing the question and deciding on important outcomes. J Clin Epidemiol 2011; 64: 395-400.
- Guyatt GH, Oxman AD, Kunz R, et al; GRADE guidelines 6. Rating the quality of evidence — imprecision. J Clin Epidemiol 2011; 64: 1283-1293.
- Guyatt GH, Oxman AD, Kunz R, et al; GRADE Working Group. GRADE guidelines: 8. Rating the quality of evidence — indirectness. J Clin Epidemiol 2011; 64: 1303-1310.
- Guyatt GH, Oxman AD, Sultan S, et al; GRADE Working Group. GRADE guidelines: 9. Rating up the quality of evidence. J Clin Epidemiol 2011; 64: 1311-1316.
- Alonso-Coello P, Schünemann HJ, Moberg J, et al; GRADE Working Group. GRADE Evidence to Decision (EtD) frameworks: a systematic and transparent approach to making well informed healthcare choices. 1: Introduction. BMJ 2016; 353: i2016.
- Mustafa RA, Santesso N, Brozek J, et al. The GRADE approach is reproducible in assessing the quality of evidence of quantitative evidence syntheses. J Clin Epidemiol 2013; 66: 736-742; quiz 42 e1-5.
- Brozek JL, Akl EA, Alonso-Coello P, et al; GRADE Working Group. Grading quality of evidence and strength of recommendations in clinical practice guidelines. Part 1 of 3. An overview of the GRADE approach and grading quality of evidence about interventions. Allergy 2009; 64: 669-677.
- Iorio A, Spencer FA, Falavigna M, et al. Use of GRADE for assessment of evidence about prognosis: rating confidence in estimates of event rates in broad categories of patients. BMJ 2015; 350: h870.
- Lewin S, Glenton C, Munthe-Kaas H, et al. Using qualitative evidence in decision making for health and social interventions: an approach to assess confidence in findings from qualitative evidence syntheses (GRADE-CERQual). PLoS Med 2015; 12: e1001895.
- Brignardello-Petersen R, Bonner A, Alexander PE, et al. Advances in the GRADE approach to rate the certainty in estimates from a network meta-analysis. J Clin Epidemiol 2018; 93: 36-44.
- O’Mara-Eves A, Thomas J, McNaught J, et al. Using text mining for study identification in systematic reviews: a systematic review of current approaches. Syst Rev 2015; 4: 5.
Provenance: Not commissioned; externally peer reviewed.