Guides / systematic review
Systematic reviews: a step-by-step guide
A systematic review is defined by its method, not its length: a pre-specified question, a reproducible search, transparent screening and an honest appraisal of what was found. This guide walks through each stage in order, explains what has to be decided before you start, and covers the PRISMA reporting that journals now require.
A systematic review is a review that uses explicit, pre-specified and reproducible methods to identify, select, appraise and synthesise all studies relevant to a defined question. What distinguishes it from a traditional literature review is that the method is fixed in advance and documented in enough detail that another researcher could repeat it and reach the same set of included studies.
Definition
What makes a review systematic
A systematic review is defined by its method, not by how thorough it feels. The question, the search, the inclusion criteria and the analysis are all specified before the work begins, and documented so that someone else could reproduce them.
| Traditional review | Systematic review | |
|---|---|---|
| Question | Broad topic | Specific, pre-specified |
| Search | Whatever the author knows | Documented strategy across named databases |
| Selection | Author’s judgement | Explicit criteria, applied by two reviewers |
| Appraisal | Optional | Formal risk-of-bias assessment |
| Reproducible | No | Yes — that is the point |
Could another researcher follow your written method and arrive at the same set of included studies? If not, the review is not systematic, however comprehensive it is. This is why the protocol comes first and why the search strategy must be published in full.
The effort involved is routinely underestimated, and it is worth being realistic before committing. A search returning 1,500 records means 1,500 titles and abstracts to screen, twice over if you are following the standard method, then perhaps 80 full texts to obtain and read, then extraction and appraisal for those that survive. For a lone researcher on a dissertation timetable, that is a substantial share of the available time and it front-loads onto the months when supervisors expect to see progress. Scope the question so the yield is manageable, and run a scoping search before writing the protocol so you know what you are committing to rather than discovering it afterwards.
Review types worth distinguishing
| Type | Purpose | Typical length |
|---|---|---|
| Systematic review | Answer a specific question exhaustively | 6–18 months |
| Scoping review | Map what evidence exists on a broad topic | 3–9 months |
| Rapid review | Systematic method with documented shortcuts | 4–12 weeks |
| Meta-analysis | Statistical pooling within a review | Adds weeks |
| Narrative review | Expert overview | Variable |
Question
Step 1: frame the question
A systematic review can only be systematic if the question is specific enough that inclusion decisions follow from it. The usual scaffold is PICO.
| Framework | Elements | Best for |
|---|---|---|
| PICO | Population, Intervention, Comparator, Outcome | Effectiveness questions |
| PICOS | PICO plus Study design | When restricting to trials |
| SPIDER | Sample, Phenomenon, Design, Evaluation, Research type | Qualitative reviews |
| PEO | Population, Exposure, Outcome | Observational and epidemiological |
“The effectiveness of educational technology” will return tens of thousands of records and cannot be completed by one person. Narrow the population, the intervention or the outcome until the expected yield is manageable. Run a scoping search first to find out what the yield will be, before committing.
Protocol
Step 2: write and register the protocol
The protocol states what you will do before you do it, which is what prevents inclusion criteria drifting to suit the studies you happen to find.
Writing the protocol is also the point at which most scoping problems surface. If you cannot write inclusion criteria precise enough for two people to apply consistently, the question is not yet specific enough, and no amount of effort later will compensate. Testing the draft criteria against a handful of papers you already know should be included is a quick and effective check.
Register it prospectively. PROSPERO is the usual registry for health and social care reviews and is free; OSF serves other fields. Registration is increasingly a submission requirement, and it protects you: a documented prior plan is the strongest answer to any suggestion that criteria were adjusted after seeing results.
Protocols legitimately change — a search returns an unworkable volume, or an outcome turns out never to be reported. Record the change, the date and the reason, and report deviations in the paper. What damages a review is an undisclosed change, not a disclosed one.
Searching
Step 3: build the search strategy
The search is the part most often done badly, and it is the part that determines whether the review can claim completeness.
Sensitivity and specificity trade off against each other in a search exactly as they do in a diagnostic test. A broad search finds nearly everything relevant but returns a great deal of irrelevant material to screen; a narrow one is quicker but risks missing eligible studies. For a systematic review the convention leans heavily towards sensitivity, because a missed study is a flaw in the review while an extra thousand records is only labour. Where you have deliberately narrowed the search — excluding non-English publications, for example — state it as a limitation rather than leaving the reader to infer it.
The structure of a search
Each PICO element becomes a block. Within a block, every way of expressing the concept is joined with OR; the blocks are then joined with AND. For the example above, the intervention block might be (“retrieval practice” OR “testing effect” OR “practice testing” OR quiz*).
A specialist librarian will improve almost any search strategy, and many institutions provide this free. The PRESS checklist exists specifically for peer review of search strategies. An inadequate search is the single most common reason a systematic review is rejected, and it cannot be fixed at the writing-up stage.
Publish the full strategy for at least one database as an appendix, exactly as run, including line numbers and the count returned at each line. Describing it in prose is not sufficient for reproducibility.
Screening
Step 4: screen in two stages
| Stage | What you read | Typical retention |
|---|---|---|
| Title and abstract | Titles and abstracts only | 5–15% proceed |
| Full text | The whole paper | 20–40% of those proceed |
Screening is where most of the calendar time goes, and it is worth planning around that. A rate of roughly 200 to 400 titles and abstracts an hour is realistic once you are practised and the criteria are settled; full-text screening runs at perhaps eight to fifteen papers an hour depending on how easily eligibility can be judged. Multiply those against your expected yield before you commit to a timetable, because a search returning 4,000 records is several full working days of screening alone, twice over if it is done in duplicate.
Two reviewers should screen independently, with disagreements resolved by discussion or a third reviewer. Where resources do not allow full double screening, a common compromise is for a second reviewer to independently screen a random 20% and to report the agreement, usually as Cohen's kappa.
Every count in the PRISMA diagram has to reconcile: records identified minus duplicates equals records screened; screened minus excluded equals full texts; full texts minus exclusions equals included. Reconstructing these afterwards is painful and is where most diagram errors originate.
Enter your counts and get a PRISMA 2020 flow diagram as a downloadable SVG, with every stage checked for consistency before you export it.
Use the generatorAppraisal
Step 5: assess risk of bias
Including a study is not the same as trusting it. Risk of bias assessment is what separates a systematic review from a bibliography, and it must use an established tool rather than a checklist of your own.
| Study design | Tool |
|---|---|
| Randomised trials | Cochrane RoB 2 |
| Non-randomised intervention studies | ROBINS-I |
| Observational studies | Newcastle-Ottawa Scale |
| Diagnostic accuracy | QUADAS-2 |
| Qualitative studies | CASP qualitative checklist |
| Systematic reviews (for an overview) | AMSTAR 2 |
Note that risk of bias is judged per outcome, not per study, in the current Cochrane tools. A trial may be at low risk for its primary outcome, which was objectively measured, and at high risk for a secondary one that depended on self-report from unblinded participants. Assessing the study as a single unit obscures that distinction, and it is the outcome-level judgement that should feed your synthesis.
Assess independently in duplicate, as with screening. Present the results per study and per domain, usually as a traffic-light figure, and — crucially — use them. An assessment that is reported and then ignored in the synthesis has served no purpose.
Run a sensitivity analysis restricted to low-risk studies. If the conclusion holds, that materially strengthens it. If it does not, that is itself an important finding and should be stated prominently rather than buried in an appendix.
Extraction
Step 6: extract the data
Design the extraction form before you start and pilot it on three or four studies. Redesigning it midway means revisiting every study already done.
Build the form in a spreadsheet or a purpose-built tool rather than in a word processor, and give every study a stable identifier you use consistently across screening, extraction and appraisal. Reconciling three separate files that refer to the same study as “Ahmed 2021”, “Ahmed et al.” and “study 14” is a predictable and entirely avoidable waste of a week.
Extract in duplicate where possible; extraction errors are common and consequential, particularly for the numbers that feed a meta-analysis. Where a paper reports a median and range but you need a mean and SD, use a documented conversion method and record that you used it.
It works more often than people expect, particularly for recent papers. Record who you contacted, when, and whether they responded — PRISMA asks for it, and it demonstrates the search for data was genuinely exhaustive.
Synthesis
Step 7: synthesise
Meta-analysis is one option, not the default. The right synthesis depends on whether the studies are similar enough to combine and whether they report data that can be pooled.
| Approach | When appropriate |
|---|---|
| Meta-analysis | Studies comparable in population, intervention and outcome, with extractable effect sizes |
| Narrative synthesis | Studies too heterogeneous, or outcomes reported inconsistently |
| Synthesis without meta-analysis (SWiM) | Structured narrative following a reporting standard |
| Vote counting on direction | A last resort, and only with the limitations stated |
| Thematic synthesis | Qualitative studies |
Concluding that meta-analysis is inappropriate is a legitimate and often correct outcome. A pooled number derived from studies that measured different things in different populations is worse than no pooled number, because it carries false authority.
GRADE is the standard framework for rating how much confidence to place in the body of evidence for each outcome, taking account of risk of bias, inconsistency, indirectness, imprecision and publication bias. Journals in health increasingly expect a summary-of-findings table, and it is what turns a synthesis into a usable conclusion.
Send your extraction table. A named statistician advises whether pooling is appropriate, runs the meta-analysis if it is, and returns forest and funnel plots with the heterogeneity assessed.
See statistical consultancyReporting
Reporting with PRISMA 2020
PRISMA 2020 is the reporting standard, comprising a 27-item checklist and the flow diagram. Most journals require both, and the checklist should be submitted with page numbers.
Reviewers add up the numbers in the flow diagram before reading the discussion, because it is the fastest way to detect a review that was not conducted as described. Mismatches are common and they undermine confidence in everything that follows.
Pitfalls
Seven mistakes that get reviews rejected
1. An inadequate search
One database, no controlled vocabulary, no grey literature. This cannot be repaired later and is the most common reason for rejection.
2. No protocol or registration
Without a prior plan, there is no defence against the suspicion that criteria followed the findings.
3. Single-reviewer screening with no agreement check
Duplicate screening is the standard safeguard against error and bias. If resources genuinely prevent it, have a second person screen a random sample, report the agreement statistic, and state the limitation explicitly rather than leaving it unmentioned.
4. Risk of bias assessed but not used
If the assessment does not influence the synthesis or a sensitivity analysis, it was decorative.
5. A flow diagram whose numbers do not add up
Check every stage reconciles before submission.
6. Meta-analysing incomparable studies
A pooled estimate across different populations and outcomes has false precision. Narrative synthesis is the better answer.
7. A question too broad to complete
Run a scoping search first. If it returns 40,000 records, narrow the question before writing the protocol.
Answers
Frequently asked questions
What is the difference between a systematic review and a literature review?
A systematic review uses explicit, pre-specified and reproducible methods: a registered protocol, a documented search across multiple databases, defined inclusion criteria applied by two reviewers, and formal risk-of-bias assessment. A traditional literature review relies on the author's selection and cannot be reproduced.
How long does a systematic review take?
Typically six to eighteen months for a full review with a team. Screening alone can take weeks when the search returns thousands of records. A rapid review compresses this to four to twelve weeks by documenting specific shortcuts, such as single-reviewer screening or restricting the search.
How many databases should I search?
At least three, and usually more. For health topics MEDLINE, Embase and CENTRAL are standard; for social sciences and education, PsycINFO, ERIC and Scopus or Web of Science. Supplement with reference list checking, trial registries and grey literature, and record the date each search was run.
Do I need to register my systematic review?
It is strongly recommended and increasingly required for publication. PROSPERO is free and standard for health and social care reviews; OSF serves other fields. Registration protects you against any suggestion that inclusion criteria were adjusted after seeing the results.
Does every systematic review need a meta-analysis?
No. Meta-analysis is appropriate only when studies are similar enough in population, intervention and outcome to combine meaningfully, and report extractable data. Concluding that pooling would be inappropriate and presenting a narrative synthesis instead is a legitimate and often correct outcome.
What risk of bias tool should I use?
It depends on the study designs included: Cochrane RoB 2 for randomised trials, ROBINS-I for non-randomised intervention studies, the Newcastle-Ottawa Scale for observational studies, QUADAS-2 for diagnostic accuracy, and CASP checklists for qualitative research. Use an established tool rather than devising your own.
What is PRISMA?
The Preferred Reporting Items for Systematic Reviews and Meta-Analyses: a 27-item checklist and a flow diagram that together define the reporting standard. The 2020 version is current. Most journals require the completed checklist with page numbers at submission.
Can one person do a systematic review?
It is possible but methodologically weaker, since duplicate screening and extraction are standard safeguards against error. If working alone, have a second person independently screen a random sample — commonly 20% — report the agreement, and state the limitation explicitly.
Send the data. Get a fixed quote.
Attach your dataset or just describe the project. A named statistician replies with a price and a deadline, usually within one working day.