How to Report Multivariate Statistical Methods in Biological Papers

By Dr. Zubair Khalid, DVM, MS, PhD ·

How to Report Multivariate Statistical Methods in Biological Papers

Key Takeaways

  • Reproducibility hinges on granular reporting of software and parameters: Specify exact software package versions (e.g., R version 4.3.1, mixOmics 6.22.0), computing environments, and all parameter settings (e.g., number of components, regularization values, convergence criteria) to enable exact replication of the analysis pipeline.
  • Preprocessing transparency is paramount for multivariate analysis: Detail every step of data transformation, including scaling (e.g., Pareto scaling), normalization, centering, and missing data imputation methods (e.g., half the minimum detected value), as these decisions critically influence multivariate results.
  • Biological questions must precede statistical model justification: Clearly articulate the research question (e.g., differentiating treatment groups vs. identifying natural clusters) before presenting the multivariate method, and provide citations for the chosen method's rationale or reporting guidelines (e.g., EQUATOR Network).
  • Validation strategies require explicit description: Detail the specific cross-validation scheme (e.g., five-fold), permutation test parameters (e.g., 100 permutations), or bootstrap procedures, including the evaluation metric, to allow readers to assess model overfitting and reliability.
  • Data availability and access conditions are non-negotiable for verification: Provide repository names, accession numbers, and clear access conditions for raw and processed data, moving beyond simple statements of availability "on request" to meet reproducibility standards.

Quick Answer

  • Report multivariate methods with software names, version numbers, and complete parameter settings so another laboratory can reproduce the exact analysis pipeline.
  • State the biological question before the statistical model, then justify each multivariate choice with a citation to a methods reference or reporting guideline.
  • A critical limitation is that multivariate results depend heavily on preprocessing decisions, so transparency about scaling, missing data handling, and validation procedures matters more than the final p value.

At a Glance

Reporting ComponentWhat to IncludeCommon OmissionConsequence of Omission
Software and versionPackage name, version number, and computing environmentListing only the software nameReviewers cannot verify that the analysis matches the stated method
Preprocessing stepsScaling, transformation, centering, and missing data handlingDescribing only the final modelResults cannot be reproduced because input data differ
Model parametersNumber of components, regularization values, and convergence criteriaReporting only the outputThe analysis cannot be rerun with identical settings
Validation strategyCross-validation scheme, permutation tests, or bootstrap detailsClaiming validation without describing itOverfitting cannot be be assessed by readers
Data availabilityRepository name, accession numbers, and access conditionsSaying data are available on requestThe study fails reproducibility standards

Why Multivariate Reporting Matters in Biological Research

Biological datasets routinely contain many measured variables per sample. Gene expression arrays, metabolomics profiles, microbiome counts, and imaging features all produce wide data tables where the number of variables can exceed the number of biological samples. Multivariate statistical methods reduce these high dimensional data into interpretable structures, identify clusters, and reveal relationships among variables that univariate tests cannot detect.

The challenge for a biologist writing a manuscript is not performing the analysis but describing it with enough precision that another laboratory can repeat it. Journals increasingly require detailed methods sections, and funding agencies expect data management plans that support verification. The National Library of Medicine Research Methods Resources provides access to authoritative biomedical research methods references that can guide the selection and description of statistical approaches. The EQUATOR Network maintains reporting guidelines for many study types and can help authors identify the specific checklist that applies to their design.

A methods section that says only principal component analysis was performed is not reproducible. The reader needs to know which software, which function, which scaling method, which number of components, and which validation procedure were used. This level of detail is not optional in modern biological publishing. It is the difference between a manuscript that can be verified and one that must be taken on faith.

Core Principles for Reporting Multivariate Methods

State the Biological Question Before the Statistical Method

The methods section should begin with the biological question that motivates the multivariate approach. A reader needs to understand whether the analysis is exploratory, hypothesis testing, or predictive. This context determines which methods are appropriate and how the results should be interpreted.

For example, a study that asks whether two treatment groups differ in overall gene expression profiles might use a supervised method such as partial least squares discriminant analysis. A study that asks whether samples cluster into natural groups without prior labels might use an unsupervised method such as principal component analysis or hierarchical clustering. The biological question determines the method, and the methods section should make that reasoning explicit.

Match the Method to the Data Structure

Multivariate methods make different assumptions about the data. Principal component analysis assumes linear relationships and works best with continuous variables. Correspondence analysis handles count data. Distance based methods such as permutational multivariate analysis of variance work with any distance measure and are common in ecology and microbiome research. The choice of method must be justified in relation to the data type, the research question, and the distribution of the variables.

The methods section should state why a particular method was chosen over alternatives. This justification does not need to be long, but it should show that the choice was deliberate and informed. A sentence that says the data were analyzed with principal component analysis because the variables were continuous and the goal was to visualize overall sample structure is sufficient.

Report the Software and Version

Every multivariate analysis should include the software package, the version number, and the citation for the package. R packages such as vegan, ade4, and mixOmics have specific functions and default settings that change between versions. The same analysis run in two different versions can produce different results. The methods section should name the package, the version, and the function used.

The computing environment also matters. The operating system, the R version, and the random seed used for any stochastic procedures should be reported. This information allows a reader to reproduce the analysis exactly or to understand why their results might differ.

Describe Preprocessing Steps in Order

Preprocessing is the most underreported part of multivariate analysis. The methods section should list every step applied to the raw data before the multivariate analysis. This includes data filtering, normalization, transformation, scaling, and missing value imputation. Each step should be described with its parameters.

For example, a metabolomics study might report that raw peak areas were log transformed, then Pareto scaled, and that missing values were replaced with half the minimum detected value. A microbiome study might report that counts were rarefied to a depth of 10,000 sequences per sample and then transformed to relative abundance. These details are essential for reproduction.

Specify Model Parameters and Settings

The methods section should list every parameter that was set by the analyst. This includes the number of components or factors retained, the distance measure used, the linkage method for clustering, the number of permutations in a test, and the number of folds in cross-validation. If the software has default settings, the methods should state that defaults were used and name the defaults.

The number of components retained in a principal component analysis is a decision that affects the interpretation. The methods should state how this number was chosen, whether by the Kaiser criterion, a scree plot, a cross-validation procedure, or a fixed threshold. The same applies to the number of clusters in a clustering analysis.

Report Validation and Permutation Procedures

Validation is the process of checking whether the multivariate model performs well on data that were not used to build the model. Cross-validation splits the data into training and test sets, fits the model on the training set, and evaluates it on the test set. Permutation tests shuffle the group labels to create a null distribution and compare the observed statistic to that distribution.

The methods section should describe the validation procedure in enough detail that a reader can repeat it. This includes the number of folds, the number of permutations, the random seed, and the criterion used to evaluate the model. A statement that says the model was validated with five fold cross-validation is not enough. The reader needs to know how the folds were created and what metric was used to evaluate the model.

Practical Workflow for Writing the Methods Section

Step 1: List Every Analysis That Was Performed

Before writing the methods section, list every multivariate analysis that appears in the results. This includes the main analysis, any sensitivity analyses, and any supplementary analyses. Each analysis needs its own description in the methods section.

Step 2: Record the Software and Version

For each analysis, record the software package, the version number, and the function or command used. This information should be recorded at the time of the analysis, not reconstructed later. A laboratory notebook or an electronic log is the best place to keep this record.

Step 3: Document the Preprocessing Steps

Write down every step that transformed the raw data into the data that was used in the analysis. This includes any filtering, normalization, transformation, scaling, or imputation. The order of the steps matters, so record them in the sequence they were applied.

Step 4: State the Parameters and Settings

For each analysis, record every parameter that was set. This includes the number of components, the distance measure, the linkage method, the number of permutations, and the number of folds. If defaults were used, record the default values.

Step 5: Describe the Validation Procedure

For each analysis that includes validation, describe the validation procedure in detail. This includes the number of folds, the number of repetitions, and the metric used to evaluate the model.

Step 6: Write the Methods Section

Use the recorded information to write the methods section. The section should be organized by analysis, with each analysis described in a separate paragraph. The description should follow the order of the workflow: data, preprocessing, method, parameters, validation.

Step 7: Check Against the Reporting Guideline

Before submitting the manuscript, check the methods section against the relevant reporting guideline. The EQUATOR Network maintains a searchable database of reporting guidelines for different study types. The guideline will list the items that should be reported and can help identify any missing information.

Example Paragraphs for the Methods Section

Example for Principal Component Analysis

Principal component analysis was used to visualize the overall structure of the metabolomics data and to identify patterns of sample separation. The analysis was performed in R version 4.3.1 using the prcomp function from the base stats package. The data matrix contained 120 samples and 450 metabolite features. Raw peak areas were log transformed and then Pareto scaled before analysis. The number of components retained was determined by the broken stick model, which indicated that the first three components explained 68 percent of the total variance. The analysis was performed on the scaled data matrix without centering because the data had already been centered by the scaling procedure. The results were visualized as a score plot of the first two components, with samples colored by treatment group.

Example for Partial Least Squares Discriminant Analysis

Partial least squares discriminant analysis was used to identify the metabolites that best discriminated between the control and treated groups. The analysis was performed in R using the mixOmics package version 6.22.0. The data was log2 transformed and Pareto scaled before analysis. The model was fitted with two components, which were selected by five fold cross-validation using the root mean squared error of prediction as the criterion. The model was validated by 100 permutations of the group labels, which produced a p value of 0.01 for the observed separation. The variable importance in projection scores were used to identify the metabolites that contributed most to the discrimination.

Example for Permutational Multivariate Analysis of Variance

Permutational multivariate analysis of variance was used to test whether the microbial community composition differed between the three treatment groups. The analysis was performed in R using the adonis2 function of the vegan package version 2.6.4. The community data was a matrix of operational taxonomic unit counts that was rarefied to 10,000 sequences per sample and then transformed with a Hellinger transformation. The distance matrix was calculated using the Bray-Curtis dissimilarity. The analysis was run with 999 permutations, and the p-value was calculated as the proportion of permutations that produced a pseudo-F statistic greater than the observed value. The analysis was stratified by the block variable to account for the experimental design.

Common Failure Patterns in Reporting Multivariate Methods

Failure to Report the Software Version

A methods section that says the analysis was performed in R without stating the version is not reproducible. The R version affects the behavior of the packages and the results of the analysis. The same analysis in R 4.0 and R 4.5 can produce different results because the underlying functions have changed.

Failure to Describe Preprocessing

Preprocessing is the most common source of irreproducibility in multivariate analysis. A reader cannot reproduce the analysis if they do not know how the data was normalized, transformed, or scaled. The methods section should describe every preprocessing step in the order it was applied.

Failure to State the Number of Components

The number of components retained in a principal component analysis or a partial least squares model is a critical parameter. The methods section should state how the number was chosen and what criterion was used. A statement that says the first two components were used without explaining why is not sufficient.

Failure to Describe Validation

Validation is the step that gives the reader confidence in the model. A methods section that says the model was validated without describing the validation procedure is not useful. The reader needs to know the number of folds, the number of permutations, and the metric used to evaluate the model.

Failure to Report the Random Seed

Many multivariate methods use random number generation for cross-validation, permutation tests, or bootstrap procedures. The random seed determines the sequence of random numbers and therefore the results of the analysis. The methods section should state the random seed that was used so that the analysis can be reproduced exactly.

Failure to Distinguish Exploratory and Confirmatory Analysis

An exploratory analysis that generates hypotheses should be reported differently from a confirmatory analysis that tests a hypothesis. The methods section should state which type of analysis was performed. An exploratory analysis should not be presented as a confirmatory test, and the results should be interpreted with the appropriate caution.

Observations and Measurements That Support the Methods

Record the Analysis Environment

The analysis environment includes the operating system, the software version, and the package versions. This information should be recorded at the time of the analysis. The R session information can be saved using the sessionInfo function, which records the R version, the platform, and the versions of all loaded packages.

Record the Random Seed

The random seed should be recorded for any analysis that uses random procedures. The seed can be set with the set.seed function in R. The seed value should be reported in the methods section so that the analysis can be reproduced exactly.

Record the Data Processing Steps

The data processing steps should be recorded in a script or a notebook. The script should be saved with the analysis so that the steps can be reviewed and repeated. The script should be commented to explain the purpose of each step.

Record the Model Parameters

The model parameters should be recorded in the script or in a separate log. The parameters include the number of components, the distance measure, the linkage method, and the number of permutations. The parameters should be recorded at the time of the analysis, not reconstructed later.

Record the Validation Results

The validation results should be recorded for each model. This includes the cross-validation error, the permutation p-value, and the metric used to evaluate the model. The validation results should be reported in the results section of the manuscript.

Quality Controls for the Methods Section

Check That Every Analysis Is Described

Before submitting the manuscript, check that every analysis that appears in the results is described in the methods section. A reader should be able to find the description of each analysis in the methods and the results in the results.

Check That Every Parameter Is Stated

Check that every parameter that was set by the analyst is stated in the methods section. This includes the number of components, the distance measure, the linkage method, and the number of permutations. If a default value was used, the default should be stated.

Check That the Preprocessing Is Complete

Check that every preprocessing step is described in the methods section. This includes any filtering, normalization, transformation, scaling, or imputation. The order of the steps should be stated.

Check That the Validation Is Described

Check that the validation procedure is described for every model that was validated. This includes the number of folds, the number of permutations, and the metric used to evaluate the model.

Check That the Data Availability Statement Is Complete

The data availability statement should state where the data can be accessed and how the analysis scripts can be obtained. The National Institutes of Health Data Management and Sharing Policy describes the expectations for data management and sharing for NIH funded research. The policy requires that data be shared in a way that is consistent with the scientific integrity and the protection of the research participants.

Common Failure Patterns in the Results Section

Reporting Only the p Value

A results section that reports only the p value of a multivariate test is not informative. The reader needs to know the effect size, the direction of the effect, and the confidence in the result. The results should report the test statistic, the degrees of freedom, and the p value, along with a measure of the effect size.

Reporting Only the Score Plot

A results section that shows only the score plot of a principal component analysis is not complete. The reader needs to know the variance explained by each component, the loadings of the variables, and the results of any validation. The score plot should be accompanied by the loadings plot and the variance explained.

Reporting the Model Without the Validation

A results section that reports the model without the validation is not complete. The reader needs to know whether the model was validated and how the validation was performed. The validation results should be reported in the results section.

Reporting the Analysis Without the Data

A results section that reports the analysis without the data is not complete. The reader needs to know where the data can be found and how the analysis can be reproduced. The data availability statement should be included in the manuscript.

Limitations of Multivariate Methods and Their Reporting

The Results Depend on the Preprocessing

The results of a multivariate analysis depend on the preprocessing steps that were applied to the data. Different preprocessing steps can produce different results. The methods section should describe the preprocessing steps in enough detail that the reader can reproduce them.

The Results Depend on the Parameters

The results of a multivariate analysis depend on the parameters that were set by the analyst. The number of components, the distance measure, and the linkage method all affect the results. The methods section should state the parameters that were used.

The Results are Not Always Stable

The results of a multivariate analysis can be unstable, especially when the sample size is small or the number of variables is large. The results should be validated with a procedure such as cross-validation or permutation testing. The validation results should be reported in the methods section.

The Results are Not Always Interpretable

The results of a multivariate analysis are not always easy to interpret. The loadings of the variables and the scores of the samples can be difficult to understand. The results should be presented in a way that is clear and interpretable to the reader.

Safety and Regulatory Context for Data Reporting

Data Management and Sharing Policies

The National Institutes of Health Data Management and Sharing Policy requires that NIH funded research share the data and the analysis scripts that support the published results. The policy applies to research that is funded by NIH grants and cooperative agreements. The data management and sharing plan should describe how the data will be shared and how the analysis scripts will be made available.

Publication Ethics

The Committee on Publication Ethics Core Practices describes the responsibilities of authors, reviewers, and editors in the publication process. The core practices include the handling of data, the reporting of conflicts of interest, and the investigation of misconduct. The methods section should be written in a way that is transparent and honest.

Author Identity

The ORCID for Researchers describes the use of the ORCID identifier to distinguish the author and to link the author to their research outputs. The ORCID identifier should be included in the manuscript to ensure that the author is correctly identified.

Professional Escalation Criteria

When to Consult a Biostatistician

A biologist should consult a biostatistician when the analysis is complex, when the data are not well understood, or when the results are not clear. A biostatistician can help with the selection of the method, the interpretation of the results, and the reporting of the methods.

When to Consult a Data Manager

A biologist should consult a data manager when the data is large, when the data is complex, or when the data needs to be shared. A data manager can help with the organization of the data, the documentation of the data, and the sharing of the data.

When to Consult a Journal Editor

A biologist should consult a journal editor when the journal has specific requirements for the methods section, when the journal has a specific reporting guideline, or when the journal has a specific data sharing policy. The journal editor can help with the requirements of the journal.

A Decision Framework for Choosing What to Report in Multivariate Methods

The Core Problem: Deciding Which Details Matter

Biologists often struggle not with writing the methods section but with deciding which of the many possible details to include. A manuscript cannot list every keystroke, yet it must contain enough information for another laboratory to reproduce the analysis. The solution is a practical decision framework that classifies each analysis detail into one of three categories: essential, important, or optional. This framework helps authors allocate limited manuscript space to the details that most affect reproducibility and interpretation.

The framework is built on a simple principle: report any detail that, if changed, would alter the numerical results or the biological interpretation. A detail that does not affect either can be omitted. This principle is consistent with the transparency expectations described in the Committee on Publication Ethics Core Practices, which require authors to report their methods accurately and completely enough for others to evaluate the work.

The Three Tier Classification System

Tier One: Essential Details

These details must appear in the methods section for every multivariate analysis. Omitting any of them makes the analysis unreproducible in a meaningful sense. The essential tier includes the software package and version, the specific function or command used, the preprocessing steps in order, the number of components or factors retained, the distance measure or similarity metric, the validation procedure, and the random seed for any stochastic process.

The software version is essential because multivariate functions change between versions. A function that performed one way in version 2.6.4 may perform differently in version 2.7.0. The preprocessing steps are essential because they determine the actual data matrix that enters the analysis. The number of components is essential because it determines how much of the variance is interpreted. The distance measure is essential because it defines the similarity between samples. The validation procedure is essential because it determines whether the model can be trusted. The random seed is essential because it determines the exact sequence of random numbers used in permutations or cross-validation folds.

How to Apply the Essential Tier

For each analysis, ask one question: if this detail were changed, would the reported results change? If the answer is yes, the detail is essential. For example, changing the scaling method from Pareto scaling to autoscaling changes the relative contribution of each variable to the principal components. Changing the number of permutations from 999 to 9999 changes the precision of the p value. Changing the random seed changes the exact composition of the cross-validation folds. Each of these details must be reported.

The Second Tier: Important Contextual Details

The second tier includes details that do not change the numerical results but affect the interpretation or the generalizability of the findings. These details should be reported when space permits and when they are relevant to the biological question.

The second tier includes the criterion used to select the number of components, the rationale for choosing one method over an alternative, the biological question that motivated the analysis, and the relationship of the analysis to other analyses in the manuscript. These details do not change the numbers but they help the reader understand why the analysis was performed and how the results should be interpreted.

For example, stating that the number of components was selected by the broken stick model is a second tier detail. The number of components itself is first tier. The broken stick criterion explains the decision but does not change the result. Similarly, stating that principal component analysis was chosen because the variables were continuous and the goal was to visualize overall sample structure is a second tier detail. It does not change the numbers but it helps the reader understand the analytical logic.

How to Apply the Second Tier

The second tier details should be included when they are available and when they do not consume excessive space. A methods section that includes only first tier details is technically complete but may be difficult to follow. A methods section that includes first and second tier details is both complete and readable. The second tier details are especially important when the analysis is unusual, when the method was chosen over a more common alternative, or when the biological interpretation depends on the analytical context.

The Third Tier: Optional Details

The third tier includes details that do not affect the results or the interpretation and can be omitted without loss of reproducibility. These details include the operating system version, the computer hardware, the time the analysis was run, and the specific code used to generate a plot.

The operating system is a third tier detail because most multivariate functions produce identical results across operating systems. The computer hardware is third tier because the results do not depend on the processor or the memory. The time the analysis was run is third tier because it has no effect on the results. The specific plotting code is third tier because the plot is a visualization of the results, not the results themselves.

How to Apply the Third Tier

The third tier details should be omitted from the methods section. They add noise and distract from the essential information. The exception is when a third tier detail becomes relevant because of a specific circumstance. For example, if the analysis was run on a system with a known numerical precision issue, the operating system might become a first tier detail. But in the absence of such a circumstance, the third tier details are omitted.

A Practical Decision Table for the Framework

The following table summarizes the decision framework and provides a quick reference for the biologist writing a methods section.

Detail CategoryExamplesDecision RuleReporting Action
First tierSoftware version, preprocessing steps, number of components, distance measure, validation procedure, random seedDoes the detail change the numerical results if changed?Always report in the methods section
Second tierMethod selection rationale, component selection criterion, biological questionDoes the detail affect interpretation or context?Report when relevant and available
Third tierOperating system, hardware, analysis time, plotting codeDoes the detail affect results or interpretation?Omit unless a specific circumstance makes it relevant

How to Use the Framework During the Analysis

The framework is most useful when applied during the analysis, not after the analysis is complete. A biologist who records the first tier details at the time of the analysis will have the information needed to write the methods section later. A biologist who tries to reconstruct the details after the analysis is complete will often find that some details have been lost.

The practical workflow is to keep a running log of the analysis that records each first tier detail as it is set. The log can be a simple text file, a laboratory notebook, or a script with comments. The log should record the software version, the function used, the preprocessing steps, the parameters, and the random seed. The log should be updated each time the analysis is changed or rerun.

The National Library of Medicine Research Methods Resources provides access to authoritative references on research methods that can help the biologist understand which details are likely to matter for a particular method. The EQUATOR Network provides reporting guidelines that can be used to check whether the methods section includes all the required details.

A Worked Example of the Framework in Action

Consider a biologist who has performed a principal component analysis on a metabolomics dataset. The analysis was performed in R version 4.3.1 using the prcomp function. The data were log transformed and Pareto scaled. The first three components were retained. The analysis was performed on a Windows 11 computer with 16 gigabytes of memory.

The first tier details are the R version, the prcomp function, the log transformation, the Pareto scaling, and the three components. These details must be reported. The second tier details are the reason for choosing principal component analysis and the criterion for retaining three components. These details should be reported when available. The third tier details are the Windows operating system and the memory size. These details are omitted.

The methods section would state that principal component analysis was performed in R version 4.3.1 using the prcomp function. The data were log transformed and Pareto scaled before analysis. The first three components were retained, which explained 68 percent of the total variance. The analysis was performed to visualize the overall structure of the metabolomics data and to identify patterns of sample separation.

Common Mistakes in Applying the Framework

The most common mistake is treating a second tier detail as a first tier detail and reporting it in the methods section. This mistake does not harm the manuscript but it wastes space. The second most common mistake is treating a first tier detail as a second tier detail and omitting it. This mistake is more serious because it makes the analysis unreproducible.

The most frequently omitted first tier detail is the random seed. Many biologists do not record the random seed because they do not realize that the analysis uses random number generation. The random seed is essential for any analysis that uses cross-validation, permutation tests, or bootstrap procedures. The random seed should be recorded and reported.

The second most frequently omitted first tier detail is the preprocessing order. The biologist may report that the data were log transformed and Pareto scaled but not state the order in which these steps were applied. The order matters because log transforming before scaling produces a different result than scaling before log transforming. The methods section should state the order of the preprocessing steps.

How the Framework Connects to Journal Requirements

The framework is consistent with the reporting requirements of most biological journals. The EQUATOR Network maintains a collection of reporting guidelines that specify the items that should be reported for different study types. The framework provides a practical way to decide which details to include when the guideline does not specify the exact items.

The framework is also consistent with the NIH Data Management and Sharing Policy, which requires that the data and the analysis scripts be shared in a way that supports the verification of the published results. The framework helps the biologist identify the details that must be shared to support verification.

A Checklist for Applying the Framework

The following checklist can be used to apply the framework to each multivariate analysis in a manuscript.

First, list every multivariate analysis that appears in the results. For each analysis, identify the first tier details. The first tier details are the software version, the function used, the preprocessing steps, the number of components, the distance measure, the validation procedure, and the random seed. Record each detail in the methods section.

Second, identify the second tier details. These are the method selection rationale, the component selection criterion, and the biological question. Include these details when they are available and relevant.

Third, identify the third tier details. These are the operating system, the hardware, and the plotting code. Omit these details unless a specific circumstance makes them relevant.

Fourth, check the methods section against the reporting guideline for the study type. The EQUATOR Network provides a searchable database of reporting guidelines. The guideline will list the items that should be reported and can help identify any missing information.

Fifth, check the methods section against the data availability statement. The data availability statement should state where the data and the analysis scripts can be accessed. The NIH Data Management and Sharing Policy describes the expectations for data sharing for NIH funded research.

The Framework as a Teaching Tool

The framework is also useful as a teaching tool for biologists who are learning to report multivariate methods. The framework provides a clear and simple way to decide what to include in the methods section. The biologist can apply the framework to each analysis and then check the result against the reporting guideline.

The framework is not a substitute for the reporting guidelines. The guidelines provide the specific items that should be reported for a particular study type. The framework provides a general principle for deciding which details matter. The two tools are complementary and should be used together.

The Framework and the Reproducibility Crisis

The framework addresses the reproducibility crisis in biological research by providing a practical way to ensure that the methods section contains the details needed for reproduction. The framework is based on the principle that the methods section should contain every detail that affects the numerical results or the biological interpretation. This principle is simple to state and simple to apply.

The framework also addresses the problem of methods sections that are too long and too detailed. The framework provides a way to distinguish the details that matter from the details that do not. The biologist can omit the third tier details without worrying that the methods section is incomplete.

The Framework and the Peer Review Process

The framework can also be used by reviewers to evaluate the completeness of a methods section. A reviewer can apply the framework to each analysis in the manuscript and check whether the first tier details are present. If a first tier detail is missing, the reviewer can request that the author provide the missing information.

The framework is consistent with the Committee on Publication Ethics Core Practices, which require that the methods be reported accurately and completely. The framework provides a practical way to determine whether the methods section is complete.

The Framework and the Data Management Plan

The framework is also useful for the data management plan that is required by many funding agencies. The NIH Grants and Funding pages describe the requirements for the data management plan. The plan should describe how the data will be shared and how the analysis scripts will be made available. The framework can help the biologist identify the details that need to be shared to support the verification of the results.

The framework is a practical tool that can be applied to any multivariate analysis. It is simple to use and it addresses the core problem of deciding which details to report. The framework is not a replacement for the reporting guidelines but a complement to them. The framework provides a general principle for deciding which details matter and the guidelines provide the specific items for each study type.

Frequently Asked Questions

What is the minimum information needed to report a multivariate analysis?

The minimum information is the software name and version, the preprocessing steps, the model parameters, the validation procedure, and the data availability. This information allows a reader to reproduce the analysis and to understand the results.

How do I report the version of the software?

The version of the software should be reported in the methods section. The version can be obtained from the software documentation or from the session information. The version should be reported with the package name and the function used.

How do I report the preprocessing steps?

The preprocessing steps should be reported in the order they were applied. Each step should be described with the parameters that were used. The description should be detailed enough that a reader can reproduce the steps.

How do I report the number of components?

The number of components should be reported with the criterion that was used to choose the number. The criterion can be the variance explained, a scree plot, or a cross-validation result. The number of components should be reported in the methods section.

How do I report the validation procedure?

The validation procedure should be reported with the number of folds, the number of permutations, and the metric used to evaluate the model. The validation procedure should be described in enough detail that a reader can reproduce it.

How do I report the data availability?

The data availability should be reported in the data availability statement. The statement should state where the data can be found and how the analysis scripts can be reproduced. The statement should be included in the manuscript.

What should I do if the journal has a specific reporting guideline?

The journal may have a specific reporting guideline that should be followed. The EQUATOR Network provides a searchable collection of reporting guidelines. The guideline should be followed in the preparation of the manuscript.

What should I do if the analysis is not reproducible?

If the analysis is not reproducible, the methods section should be revised to include the missing information. The software version, the preprocessing steps, the parameters, and the validation procedure should be described in the methods section. The analysis should be repeated with the recorded settings to confirm the results.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.