Zubair Khalid

Virologist/Molecular Biologist | Veterinarian | Bioinformatician

Conventional & Molecular Virology • Vaccine Development • Computational Biology

Dr. Zubair Khalid is a veterinarian and virologist specializing in conventional and molecular virology, vaccine development, and computational biology. Dedicated to advancing animal health through innovative research and multi-omics approaches.

Dr. Zubair Khalid - Veterinarian, Virologist, and Vaccine Development Researcher specializing in Computational Biology, Multi-omics, Animal Health, and Infectious Disease Research

Category: Guides

Steps in Conducting a Meta-Analysis: A Structured Workflow

A meta-analysis is a statistical procedure that combines results from multiple independent studies on the same question to produce a pooled effect estimate with greater precision than any single study can provide. For researchers in agriculture, animal science, and life sciences, this method offers a way to resolve conflicting findings across trials, quantify the magnitude of an intervention effect, and identify sources of variation that explain why results differ between studies. This article presents a structured workflow for conducting a meta-analysis, from formulating the research question through data extraction, statistical analysis, and interpretation. The workflow applies to observational studies and controlled trials, with attention to the specific decisions that determine whether a meta-analysis produces trustworthy conclusions or misleading summaries.

Defining the Research Question and Scope

The research question determines every subsequent step in a meta-analysis. A poorly framed question leads to ambiguous inclusion criteria, inconsistent data extraction, and results that cannot answer the question that motivated the work. Begin by specifying the population, intervention or exposure, comparator, outcome, and study design. This framework forces clarity about who the results apply to, what intervention or exposure is being evaluated, what the comparison group is, and which outcomes matter.

For example, a question about walking and health outcomes might specify adults aged 18 years and older, device-measured daily steps as the exposure, mortality as the outcome, and prospective cohort studies as the eligible design. A question about phototherapy in newborns would specify neonates with hyperbilirubinemia, intensive versus conventional phototherapy, reduction in total serum bilirubin as the primary outcome, and randomized controlled trials plus cohort studies as eligible designs. The level of specificity determines whether the search strategy can be executed and whether the included studies are sufficiently similar to justify pooling.

The scope also includes decisions about language restrictions, publication date ranges, and geographic coverage. Some meta-analyses restrict to English-language publications for practical reasons, but this can introduce selection bias if relevant studies are published in other languages. The search period should be justified by the research question. A meta-analysis of daily steps and health outcomes published in 2025 searched literature from January 2014 to February 2025, reflecting the period when device-measured step data became widely available in prospective cohorts. The rationale for any restriction should be documented in the protocol.

Assembling the Review Team and Registering the Protocol

A meta-analysis requires a team with complementary skills. At minimum, the team needs someone with content expertise in the topic area, someone with methodological training in systematic review methods, and someone with statistical expertise for the quantitative synthesis. For topics where study selection and data extraction require judgment, two reviewers should work independently to reduce errors and bias. The daily steps meta-analysis published in The Lancet Public Health used pairs of reviewers who independently performed study selection, data extraction, and risk of bias assessment using the Newcastle-Ottawa Scale. This dual-reviewer approach is standard practice because single-reviewer screening misses eligible studies and introduces inconsistencies in data extraction.

Protocol registration occurs before the search begins. Prospective registration in a registry such as PROSPERO documents the planned methods, including the research question, search strategy, inclusion criteria, outcomes, and analysis plan. Registration serves two purposes. It prevents selective reporting of outcomes that show significant results while omitting null findings, and it allows readers to compare the planned methods with the reported methods. The daily steps and depression meta-analysis registered with PROSPERO under number CRD42024529706, and the digital tools meta-analysis registered under CRD42024510602. Registration should occur after the protocol is finalized but before screening begins.

The protocol should specify the inclusion and exclusion criteria in operational terms. Inclusion criteria define the study designs, populations, interventions or exposures, comparators, and outcomes that qualify for the review. Exclusion criteria define what does not qualify, such as studies without original data, conference abstracts without full text, or studies that report outcomes in a format that cannot be converted to a common effect size. The protocol should also specify how disagreements between reviewers will be resolved, typically through discussion or adjudication by a third reviewer.

Developing the Search Strategy and Selecting Databases

The search strategy translates the research question into database-specific queries. A comprehensive search combines controlled vocabulary terms with free-text keywords. For medical and life science topics, the primary databases include PubMed, maintained by the National Center for Biotechnology Information, and other bibliographic databases relevant to the discipline. The daily steps and health outcomes meta-analysis searched PubMed and EBSCO CINAHL, while the walking training meta-analysis searched PubMed, Web of Science, EMBASE, Cochrane Library, and EBSCO. The depression meta-analysis searched PubMed, PsycINFO, Scopus, SPORTDiscus, and Web of Science. The choice of databases depends on the topic. A meta-analysis in animal science might search Agricola and CAB Abstracts in addition to PubMed. A meta-analysis in psychology would include PsycINFO. The protocol should justify the database selection.

Search strategies should be developed with input from a librarian or information specialist when available. The strategy should be tested against a set of known eligible studies to verify that it retrieves them. The final search strategy for each database should be reported in full so that readers can reproduce the search. The search should be supplemented by other methods, including scanning reference lists of included studies and relevant reviews, citation searching, and contacting experts in the field. The daily steps and health outcomes meta-analysis supplemented database searches with other search strategies, which is common practice for identifying unpublished or in-press studies.

The search results are exported to reference management software, and duplicate records are removed. The number of records identified, screened, assessed for eligibility, and included should be documented in a flow diagram. This documentation allows readers to assess whether the search was comprehensive and whether the screening process was systematic.

Screening Studies Against Inclusion Criteria

Screening proceeds in two stages. The first stage screens titles and abstracts against the inclusion criteria. Records that clearly do not meet the criteria are excluded. Records that potentially meet the criteria or where eligibility is unclear proceed to full-text review. The second stage screens full texts against the inclusion criteria. Each full text is assessed independently by two reviewers, and disagreements are resolved through discussion or adjudication.

The screening process should be piloted on a sample of records to ensure that reviewers apply the criteria consistently. The inclusion criteria should be applied strictly. A common error is including studies that report relevant outcomes but use ineligible designs or populations. For example, a meta-analysis of daily steps and depression in adults would exclude studies of children and adolescents, studies without objectively measured step counts, and studies that do not report depression outcomes. The digital tools meta-analysis restricted eligibility to healthy school-aged children aged 6 to 17 years and required objectively measured step counts or moderate-to-vigorous physical activity.

The screening process should track the reasons for exclusion at the full-text stage. Common reasons include ineligible study design, ineligible population, missing outcome data, duplicate publication, and insufficient data for effect size calculation. Reporting these reasons in a flow diagram allows readers to assess whether the exclusion decisions were appropriate.

Extracting Data and Assessing Risk of Bias

Data extraction converts the information in each included study into a standardized format suitable for analysis. The extraction form should capture study characteristics, participant characteristics, intervention or exposure details, comparator details, outcome data, and effect estimates. For the daily steps meta-analyses, extraction included the number of participants, age distribution, follow-up duration, step count measurement method, mortality events, and adjusted hazard ratios with confidence intervals. For the phototherapy meta-analysis, extraction included gestational age, baseline bilirubin levels, treatment protocol, and change in bilirubin levels.

Two reviewers should extract data independently, and discrepancies should be resolved by referring to the original study. The extraction form should be piloted on several studies to ensure that all relevant data fields are captured and that the definitions are clear. Data extraction errors are a major source of error in meta-analyses. A study that reports multiple effect estimates for the same outcome requires a decision about which estimate to extract, typically the most fully adjusted estimate for observational studies or the intention-to-treat estimate for randomized trials.

Risk of bias assessment evaluates the methodological quality of each included study. The choice of tool depends on the study design. The Newcastle-Ottawa Scale is commonly used for observational studies and assigns up to nine points across three domains: selection of study groups, comparability of groups, and ascertainment of exposure or outcome. The daily steps and health outcomes meta-analysis used the 9-point Newcastle-Ottawa Scale, and the COVID-19 daily steps meta-analysis used a modified version. The Cochrane Risk of Bias 2 tool is used for randomized controlled trials and assesses bias across domains including randomization, deviations from intended interventions, missing outcome data, measurement of the outcome, and selection of reported results. The digital tools meta-analysis used the revised Cochrane RoB 2 tool.

Risk of bias assessment should be conducted independently by two reviewers, with disagreements resolved through discussion. The results should be reported for each included study, and the potential impact of risk of bias on the pooled estimates should be examined through sensitivity analyses that exclude studies at high risk of bias.

Choosing the Effect Size and Statistical Model

The effect size is the metric used to combine results across studies. The choice of effect size depends on the outcome type and the data reported in the included studies. For binary outcomes, common effect sizes include risk ratios, odds ratios, and hazard ratios. For continuous outcomes, common effect sizes include mean differences and standardized mean differences. The walking training meta-analysis used mean differences with 95% confidence intervals for continuous outcomes such as glycated hemoglobin. The daily steps meta-analyses used hazard ratios for time-to-event outcomes such as all-cause mortality.

When studies report outcomes in different metrics, the effect size must be converted to a common scale. For example, studies of depression might report correlation coefficients, standardized mean differences, or risk ratios. The depression meta-analysis pooled correlation coefficients, standardized mean differences, and risk ratios separately because these metrics are not directly interchangeable. The decision about which effect size to use should be specified in the protocol and justified in the methods.

The statistical model determines how the pooled effect size is calculated and how heterogeneity is handled. A fixed effect model assumes that all studies estimate the same underlying effect and that differences between studies are due to sampling error alone. A random effects model assumes that the true effect varies across studies and that the pooled estimate represents the average of a distribution of effects. The random effects model is generally preferred for meta-analyses of observational studies and for clinical trials where interventions are delivered in different settings. The daily steps meta-analyses used random effects models with inverse-variance weighting, and the phototherapy meta-analysis used a random effects model. The depression meta-analysis used the Sidik-Jonkman random effects method.

The choice between fixed and random effects models should be guided by the expectation of heterogeneity instead of by the observed heterogeneity. A random effects model is appropriate when studies differ in populations, interventions, or settings. The random effects model produces wider confidence intervals than the fixed effect model when heterogeneity is present, reflecting the additional uncertainty from between-study variation.

Assessing Heterogeneity and Exploring Sources of Variation

Heterogeneity refers to the variability in effect estimates across studies beyond what would be expected from sampling error alone. Quantifying heterogeneity is essential because high heterogeneity undermines the validity of a pooled estimate. The I-squared statistic describes the percentage of total variation across studies that is due to heterogeneity instead of chance. The walking training meta-analysis reported moderate heterogeneity with an I-squared of 67 percent for the glycated hemoglobin outcome. The phototherapy meta-analysis reported I-squared values of 83.9 percent for bilirubin reduction and 94.7 percent for treatment duration, indicating substantial heterogeneity.

The I-squared statistic has limitations. It depends on the precision of the included studies, and it does not tell the analyst whether the heterogeneity is clinically or biologically important. Prediction intervals provide a more informative measure by estimating the range within which the true effect of a future study is expected to fall. The digital tools meta-analysis used tau-squared to calculate prediction intervals in addition to reporting I-squared. Prediction intervals are particularly useful for communicating the uncertainty around a pooled estimate when heterogeneity is present.

When heterogeneity is substantial, the analyst should explore its sources through subgroup analysis and meta-regression. Subgroup analysis divides studies into groups based on characteristics such as age, sex, intervention intensity, or study design, and compares the pooled estimates across groups. The phototherapy meta-analysis found that gestational age and baseline bilirubin level were potential sources of heterogeneity, while study design was not. The walking training meta-analysis found that the walking plus other interventions subgroup and the clear step target subgroup showed larger effects than the overall pooled estimate.

Meta-regression examines the relationship between study-level characteristics and effect size. It can assess whether continuous variables such as mean age, follow-up duration, or baseline severity explain variation in effects. Meta-regression requires a sufficient number of studies to produce reliable estimates, typically at least ten studies per covariate. The results of subgroup analyses and meta-regression should be interpreted cautiously because they are observational comparisons across studies and are susceptible to confounding.

Conducting Sensitivity Analyses and Assessing Publication Bias

Sensitivity analyses test whether the pooled estimate is robust to decisions made during the review process. Common sensitivity analyses include excluding studies at high risk of bias, excluding studies with small sample sizes, using a different statistical model, and using different methods for handling missing data. The COVID-19 daily steps meta-analysis performed sensitivity analyses by excluding studies with low methodological quality or small sample sizes to test the robustness of the findings. The phototherapy meta-analysis used sensitivity analysis to determine whether any single study had an influential effect on the pooled estimate.

Publication bias arises when studies with statistically significant results are more likely to be published than studies with null or negative results. If the available literature is a biased sample of all conducted studies, the pooled estimate will be biased. Publication bias is assessed through funnel plots and statistical tests. A funnel plot displays the effect size of each study against a measure of its precision, such as the standard error or sample size. In the absence of publication bias, the plot should resemble a symmetric inverted funnel, with smaller studies scattered more widely at the bottom and larger studies clustered near the top. Asymmetry suggests that small studies with null results are missing.

Statistical tests for funnel plot asymmetry include the Egger test and the Begg test. The phototherapy meta-analysis assessed publication bias using the funnel plot, Begg test, and Egger test. The COVID-19 daily steps meta-analysis used the Egger test. The R package metasens provides methods for testing and adjusting for funnel plot asymmetry, including the trim and fill method. The digital tools meta-analysis assessed small-study effects to detect potential publication bias.

Funnel plot asymmetry does not necessarily indicate publication bias. It can also result from true heterogeneity, where smaller studies have different effect sizes because they include different populations or use different interventions. The interpretation of funnel plot asymmetry should consider the clinical and methodological context of the included studies.

Reporting the Meta-Analysis Transparently

Reporting standards for meta-analyses ensure that readers can assess the validity of the methods and the interpretation of the results. The Preferred Reporting Items for Systematic Reviews and Meta-Analyses, known as PRISMA, provides a checklist of items that should be reported. The depression meta-analysis followed the PRISMA and Meta-analysis of Observational Studies in Epidemiology reporting guidelines. The phototherapy meta-analysis was conducted in accordance with PRISMA 2020 and MOOSE guidelines. The EQUATOR Network maintains a comprehensive collection of reporting guidelines for health research, including PRISMA and MOOSE.

The report should describe the search strategy in full, including the databases searched, the search dates, and the search terms. The study selection process should be documented in a flow diagram showing the number of records identified, screened, excluded, and included. The characteristics of each included study should be presented in a table. The risk of bias assessment should be reported for each study. The results should include the pooled effect estimate with confidence intervals, the heterogeneity statistics, the results of subgroup analyses and sensitivity analyses, and the assessment of publication bias.

The methods section should describe the statistical analysis in sufficient detail for replication. This includes the effect size metric, the statistical model, the method for estimating between-study variance, the software used, and the commands or procedures for each analysis. The R package meta is commonly used for standard meta-analysis, and the R package metasens is used for sensitivity analyses related to missing data and selection bias. The multilevel meta-analysis tutorial provides annotated R scripts and R notebooks to support transparency and reproducibility.

The interpretation should discuss the certainty of the evidence. The GRADE approach assesses the certainty of evidence across domains including risk of bias, inconsistency, indirectness, imprecision, and publication bias. The daily steps and health outcomes meta-analysis assessed certainty of evidence using GRADE. The interpretation should distinguish between the statistical significance of the pooled estimate and the clinical or practical importance of the effect. A statistically significant effect with a small magnitude may not justify a change in practice, while a non-significant effect with a wide confidence interval may be consistent with a clinically important effect.

At a Glance

Phase Key Actions Common Errors Quality Check
Question and protocol Specify population, exposure, comparator, outcome, design. Register protocol before searching Vague question, no registration, post hoc changes to methods Protocol specifies all eligibility criteria and analysis plan
Search and screening Search multiple databases, supplement with citation searching. Screen titles and abstracts, then full texts Single database search, single reviewer screening, no flow diagram Two reviewers screen independently, exclusion reasons documented
Data extraction and risk of bias Extract study characteristics and outcome data. Assess bias with design-appropriate tool Extraction errors, no dual extraction, wrong risk of bias tool Pilot extraction form, dual extraction with adjudication
Analysis Choose effect size and model. Assess heterogeneity, conduct subgroup and sensitivity analyses Fixed effect model with high heterogeneity, no sensitivity analysis, no publication bias assessment Report I-squared and prediction intervals, test robustness
Reporting Follow PRISMA or MOOSE. Report search, selection, characteristics, results, and certainty Missing search details, no flow diagram, overinterpretation of pooled estimate Full methods for replication, GRADE assessment for certainty

Practical Implementation Steps

The following sequence provides a practical workflow for conducting a meta-analysis. Each step should be completed before proceeding to the next, and the protocol should document the plan for each step.

Step 1. Form the review team. Include content expertise, methodological expertise, and statistical expertise. Assign roles for screening, extraction, and analysis. The team should include at least two reviewers for screening and extraction.

Step 2. Define the research question using the population, intervention or exposure, comparator, outcome, and study design framework. Write the question in a single sentence that specifies each element.

Step 3. Define the inclusion and exclusion criteria in operational terms. Specify the study designs, populations, interventions or exposures, comparators, outcomes, and minimum data requirements.

Step 4. Register the protocol in PROSPERO or an equivalent registry. The protocol should include the research question, search strategy, eligibility criteria, outcomes, risk of bias assessment plan, and analysis plan.

Step 5. Develop the search strategy. Identify the databases to search based on the topic. Develop the search terms with input from a librarian if available. Test the search against known eligible studies.

Step 6. Execute the search and export results to reference management software. Remove duplicates and document the number of records.

Step 7. Screen titles and abstracts against the inclusion criteria. Two reviewers screen independently. Resolve disagreements through discussion or adjudication.

Step 8. Retrieve full texts for potentially eligible records. Screen full texts against the inclusion criteria. Document the reasons for exclusion.

Step 9. Develop and pilot the data extraction form. Extract data independently with two reviewers. Resolve discrepancies by referring to the original study.

Step 10. Assess risk of bias for each included study using a design-appropriate tool. Two reviewers assess independently. Resolve disagreements through discussion.

Step 11. Choose the effect size metric and the statistical model. Specify these choices in the analysis plan before examining the data.

Step 12. Conduct the meta-analysis. Calculate the pooled effect estimate with confidence intervals. Quantify heterogeneity using I-squared and prediction intervals.

Step 13. Explore heterogeneity through subgroup analysis and meta-regression when the number of studies permits. Interpret these analyses cautiously.

Step 14. Conduct sensitivity analyses to test the robustness of the pooled estimate. Exclude studies at high risk of bias, exclude small studies, and use alternative models.

Step 15. Assess publication bias using funnel plots and statistical tests. Interpret the results in the context of the included studies.

Step 16. Report the meta-analysis following PRISMA or MOOSE guidelines. Provide the full search strategy, flow diagram, study characteristics, risk of bias results, pooled estimates, heterogeneity statistics, sensitivity analyses, and publication bias assessment.

Step 17. Assess the certainty of evidence using GRADE or an equivalent framework. Discuss the implications of the findings for practice and research.

Records and Measurements

The quality of a meta-analysis depends on the quality of the records maintained throughout the process. The following records should be maintained and made available to readers.

The search record documents the date of each search, the database searched, the search strategy, and the number of records retrieved. This record allows readers to reproduce the search and assess its comprehensiveness.

The screening record documents the number of records screened at each stage, the number excluded, and the reasons for exclusion. This record is typically presented as a flow diagram in the final report.

The data extraction record documents the data extracted from each included study, the reviewer who extracted each data point, and any discrepancies that were resolved. This record supports the accuracy of the data used in the analysis.

The risk of bias record documents the risk of bias assessment for each included study, including the domain-level judgments and the overall judgment. This record allows readers to assess the methodological quality of the included studies.

The analysis record documents the statistical methods used, the software and commands, the data files, and the results of each analysis. This record supports the reproducibility of the analysis.

The protocol and its amendments document the planned methods and any changes made during the review process. Changes to the protocol should be documented with the date and rationale for each change.

Common Failure Patterns

Several recurring problems undermine the validity of meta-analyses. Recognizing these failure patterns helps researchers avoid them and helps readers identify unreliable meta-analyses.

The first failure pattern is an inadequately defined research question. A question that does not specify the population, exposure, comparator, and outcome leads to ambiguous eligibility criteria and inconsistent study selection. The resulting meta-analysis may combine studies that are too different to justify pooling.

The second failure pattern is an incomplete search. Searching a single database or restricting to English-language publications can miss relevant studies. The daily steps and health outcomes meta-analysis searched multiple databases and supplemented with other search strategies, reflecting the need for comprehensive identification of eligible studies.

The third failure pattern is single-reviewer screening and extraction. A single reviewer may miss eligible studies or extract data incorrectly. The dual-reviewer approach used in the daily steps meta-analyses reduces these errors.

The fourth failure pattern is ignoring heterogeneity. Pooling studies with substantial heterogeneity without exploring its sources produces a pooled estimate that may not apply to any specific population or setting. The phototherapy meta-analysis reported high I-squared values and explored heterogeneity through subgroup analysis, identifying gestational age and baseline bilirubin level as sources of variation.

The fifth failure pattern is overinterpreting the pooled estimate. A pooled estimate from a meta-analysis of observational studies cannot establish causation, and a pooled estimate from studies with high risk of bias cannot overcome the limitations of the primary studies. The interpretation should acknowledge these limitations.

The sixth failure pattern is selective reporting. Reporting only the analyses that produce significant results, or changing the outcomes after seeing the data, undermines the validity of the meta-analysis. Protocol registration and adherence to the protocol reduce this risk.

The seventh failure pattern is inadequate reporting. Meta-analyses that do not report the full search strategy, the flow diagram, the risk of bias results, or the statistical methods cannot be evaluated by readers. Reporting guidelines such as PRISMA and MOOSE provide the framework for complete reporting.

Limitations and Context

Meta-analysis has important limitations that should be understood before undertaking or interpreting one. The method cannot overcome the limitations of the primary studies. If the included studies have methodological flaws, the pooled estimate will reflect those flaws. The daily steps and all-cause mortality meta-analysis noted that the evidence base consisted primarily of observational studies, which cannot establish causation.

Meta-analysis is subject to publication bias when the available literature is a biased sample of all conducted studies. Statistical tests for publication bias have limited power, particularly when the number of included studies is small. The absence of evidence for publication bias does not prove that publication bias is absent.

Meta-analysis requires that the included studies be sufficiently similar to justify pooling. When studies differ in populations, interventions, comparators, or outcome definitions, the pooled estimate may not apply to any specific context. The decision about whether studies are sufficiently similar requires judgment, and different analysts may reach different conclusions.

Meta-analysis of observational studies faces additional challenges. Observational studies are subject to confounding, and the adjusted effect estimates may not fully account for all confounders. The daily steps meta-analyses used adjusted hazard ratios from the primary studies, but residual confounding remains possible.

The statistical methods for meta-analysis continue to evolve. Multilevel meta-analysis addresses non-independence of effect sizes, such as when multiple studies come from the same laboratory. The step-by-step guide to multilevel meta-analysis describes methods for estimating overall effect sizes when effect sizes are nested within labs. Bayesian methods and the Knapp-Hartung adjustment provide alternatives for highly heterogeneous meta-analyses.

The certainty of evidence should be assessed systematically. The GRADE approach evaluates the certainty of evidence across domains including risk of bias, inconsistency, indirectness, imprecision, and publication bias. The daily steps and health outcomes meta-analysis used GRADE to assess certainty of evidence for each outcome.

Welfare and Safety Context

Meta-analyses in the life sciences often address questions with direct implications for human or animal welfare. The daily steps meta-analyses inform physical activity recommendations. The 2022 meta-analysis of 15 international cohorts found that compared with the lowest quartile of daily steps, the adjusted hazard ratio for all-cause mortality was 0.60 for the second quartile, 0.55 for the third quartile, and 0.47 for the fourth quartile. The 2025 meta-analysis found an inverse non-linear dose-response association for all-cause mortality, cardiovascular disease incidence, dementia, and falls, with inflection points around 5000 to 7000 steps per day. The umbrella review found a protective threshold at 3143 steps per day and a pooled hazard ratio of 0.91 per 1000 steps per day increment.

These findings have welfare implications because they identify physical activity levels associated with reduced mortality risk. The COVID-19 daily steps meta-analysis found that the proportion of studies with participants achieving optimal daily steps of at least 7000 steps per day declined during the pandemic, raising concerns about the health consequences of reduced physical activity.

Meta-analyses of clinical interventions have direct safety implications. The phototherapy meta-analysis found that intensive phototherapy achieved significantly greater bilirubin reductions than conventional phototherapy, with shorter treatment durations. The walking training meta-analysis found that walking training significantly reduced glycated hemoglobin levels in patients with type 2 diabetes, with a combined effect of negative 0.48 percent.

Researchers conducting meta-analyses should consider the welfare implications of their findings and communicate them responsibly. The interpretation should distinguish between statistical associations and causal effects, particularly for observational evidence. The implications for practice should be stated with appropriate caution, and the limitations of the evidence should be acknowledged.

Professional Escalation Criteria

Certain situations require escalation to a methodologist, statistician, or content expert before proceeding. The following criteria indicate that additional expertise is needed.

Escalate when the included studies are highly heterogeneous and the sources of heterogeneity cannot be identified through subgroup analysis or meta-regression. A statistician with expertise in meta-analysis methods can advise on alternative models, including the Knapp-Hartung adjustment, likelihood-based methods, or Bayesian approaches.

Escalate when the number of included studies is small, typically fewer than five, because the statistical methods for assessing heterogeneity and publication bias have limited power with small numbers of studies. A methodologist can advise on whether meta-analysis is appropriate or whether a narrative synthesis is more suitable.

Escalate when the included studies report outcomes in formats that cannot be converted to a common effect size. A statistician can advise on methods for transforming effect sizes or on whether the studies should be analyzed separately.

Escalate when the risk of bias assessment reveals serious methodological concerns in a substantial proportion of included studies. A methodologist can advise on whether the evidence base is adequate for meta-analysis and on how to handle studies at high risk of bias.

Escalate when the search identifies a large number of potentially eligible studies and the screening process becomes unmanageable. An information specialist can advise on refining the search strategy or on using screening tools.

Escalate when the results of the meta-analysis have direct implications for clinical or public health recommendations. A content expert should review the interpretation to ensure that the conclusions are consistent with the broader evidence base and that the limitations are adequately communicated.

Frequently Asked Questions

What is the difference between a systematic review and a meta-analysis?

A systematic review follows a predefined protocol to identify, evaluate, and synthesize all available evidence on a specific question. A meta-analysis is the statistical component that combines the quantitative results of eligible studies into a pooled effect estimate. A systematic review may or may not include a meta-analysis, depending on whether the included studies are sufficiently similar to justify pooling. When a review is performed following predefined steps and its results are quantitatively analyzed, it is called a meta-analysis.

How many studies are needed to conduct a meta-analysis?

There is no fixed minimum number of studies for a meta-analysis. In practice, meta-analyses with fewer than five studies have limited statistical power and unreliable estimates of between-study variance. The decision to pool should be based on the similarity of the included studies and the adequacy of the reported data, not solely on the number of studies. When few studies are available, a narrative synthesis may be more appropriate than a meta-analysis.

What is the difference between a fixed effect model and a random effects model?

A fixed effect model assumes that all studies estimate the same underlying effect and that differences between studies are due to sampling error alone. A random effects model assumes that the true effect varies across studies and that the pooled estimate represents the average of a distribution of effects. The random effects model is generally preferred when studies differ in populations, interventions, or settings, and it produces wider confidence intervals when heterogeneity is present.

How is heterogeneity measured in a meta-analysis?

Heterogeneity is commonly quantified using the I-squared statistic, which describes the percentage of total variation across studies that is due to heterogeneity instead of chance. Prediction intervals provide a more informative measure by estimating the range within which the true effect of a future study is expected to fall. The walking training meta-analysis reported moderate heterogeneity with an I-squared of 67 percent, while the phototherapy meta-analysis reported substantial heterogeneity with I-squared values above 80 percent.

What is publication bias and how is it assessed?

Publication bias arises when studies with statistically significant results are more likely to be published than studies with null or negative results. It is assessed through funnel plots, which display the effect size of each study against a measure of precision, and through statistical tests such as the Egger test and the Begg test. Funnel plot asymmetry can also result from true heterogeneity, so the interpretation should consider the clinical and methodological context.

What is the role of sensitivity analysis in a meta-analysis?

Sensitivity analysis tests whether the pooled estimate is robust to decisions made during the review process. Common sensitivity analyses include excluding studies at high risk of bias, excluding studies with small sample sizes, using a different statistical model, and using different methods for handling missing data. The COVID-19 daily steps meta-analysis performed sensitivity analyses by excluding studies with low methodological quality or small sample sizes to test the robustness of the findings.

What reporting guidelines apply to meta-analyses?

The Preferred Reporting Items for Systematic Reviews and Meta-Analyses, known as PRISMA, provides a checklist for reporting systematic reviews and meta-analyses. The Meta-analysis of Observational Studies in Epidemiology guidelines, known as MOOSE, applies to meta-analyses of observational studies. The EQUATOR Network maintains a comprehensive collection of reporting guidelines for health research. The depression meta-analysis followed PRISMA and MOOSE reporting guidelines, and the phototherapy meta-analysis was conducted in accordance with PRISMA 2020 and MOOSE guidelines.

What software is available for conducting a meta-analysis?

The R statistical environment provides free and flexible tools for meta-analysis. The R package meta is used for standard meta-analysis, and the R package metasens is used for sensitivity analyses for missing binary outcome data and potential selection bias. RevMan is another commonly used software for meta-analysis, as used in the walking training meta-analysis. The choice of software depends on the analyst's familiarity and the complexity of the analysis.

Related Articles

References and Further Reading

This article is educational and does not replace institutional policy, professional advice, or applicable safety and regulatory requirements.