Zubair Khalid

Virologist/Molecular Biologist | Veterinarian | Bioinformatician

Conventional & Molecular Virology • Vaccine Development • Computational Biology

Dr. Zubair Khalid is a veterinarian and virologist specializing in conventional and molecular virology, vaccine development, and computational biology. Dedicated to advancing animal health through innovative research and multi-omics approaches.

Dr. Zubair Khalid - Veterinarian, Virologist, and Vaccine Development Researcher specializing in Computational Biology, Multi-omics, Animal Health, and Infectious Disease Research

Category: Guides

How to Present Data in Research Papers: Methods, Tables, and Figures

Researchers in the life sciences and related fields face a common challenge when preparing manuscripts: deciding whether a data set belongs in the text, in a table, or in a figure. This article compares the main data presentation methods with practical guidance for choosing among them. The guidance applies to students writing theses, early-career researchers preparing their first journal submissions, and experienced professionals who want to improve the clarity of their work. The core principle is that the presentation format should serve the reader's ability to verify, interpret, and reuse the data. A decision table later in this article matches common data types to appropriate presentation formats, and worked examples show how the same data can appear in text, tables, and figures.

Why Data Presentation Quality Matters

The quality of data presentation directly affects whether readers can understand and trust the findings. Poor reporting is common and persists even when journals update their instructions to authors. An audit of papers published in the Journal of Neurophysiology found that 60% of papers erroneously reported standard errors of the mean, 23% included undefined measures of variability, and 40% failed to define a statistical threshold for their tests. When p-values between 0.05 and 0.1 appeared, 64% of papers misinterpreted them as statistical trends. These problems remained despite the journal updating its Information for Authors in 2016 and 2018. Required reporting practices were consistently followed by only 34 to 37% of papers, while strongly encouraged practices were followed by only 9 to 26% of papers. See the F1000Research audit of statistical reporting for the full account.

A similar pattern emerged after The Journal of Physiology and British Journal of Pharmacology published an editorial series in 2011 to improve statistical reporting. A cross-sectional analysis of papers before and after the editorial advice found no evidence of improvement. Across the two periods, 76 to 84% of papers with written measures summarizing data variability used standard errors of the mean, and 90 to 96% of papers did not report exact p-values for primary analyses and post-hoc tests. Of papers that reported p-values between 0.05 and 0.1, 56 to 63% interpreted these as trends or statistically significant. See the PLoS ONE analysis of editorial advice effectiveness for details.

These findings matter because readers use the presented data to judge whether the conclusions are supported. When variability measures are undefined or misinterpreted, the reader cannot assess the precision of the estimates. When exact p-values are absent, the reader cannot verify the significance claims. The practical implication is that researchers should treat data presentation as a core scientific skill, not as a formatting afterthought.

Choosing Between Text, Tables, and Figures

The first decision is whether a given data set belongs in the running text, in a table, or in a figure. Each format has strengths and limitations, and the choice should depend on the nature of the data and the message the author wants to convey.

When to Use Text

Text is appropriate for a small number of values that support a specific claim. For example, a sentence might state that the median overall survival was 81 months with a 5-year survival rate of 58.3% in a cohort of 103 patients. This level of detail works when the reader needs only the headline numbers to follow the argument.

Text is also appropriate for describing statistical methods, reporting test statistics in the narrative, and highlighting the most important results that appear in full in tables or figures. The text should not duplicate every value that appears in a table or figure. Duplication wastes space and forces the reader to compare two versions of the same information.

A common failure pattern is writing results paragraphs that read as a list of numbers without interpretation. The text should state the finding and its meaning, then direct the reader to the table or figure for the complete data. For example, instead of writing "The mean was 12.4, the standard deviation was 3.1, and the p-value was 0.02," write "The treatment group showed a higher mean score than the control group (table 2)."

When to Use Tables

Tables are appropriate when the reader needs to look up specific values, compare multiple groups across several variables, or access exact numbers that would be difficult to extract from a figure. Tables work well for demographic characteristics, baseline clinical features, model coefficients, and sensitivity analyses.

A table should have a clear structure with labeled rows and columns, defined units, and footnotes explaining abbreviations and statistical conventions. The title should describe the content without forcing the reader to search the text for context. Each column should contain one variable, and each row should contain one observation or one group.

Tables become difficult to read when they contain too many columns or when cells are packed with text. A table with more than ten columns often needs to be split into two tables or restructured so that variables are stacked in rows. Long decimal values should be rounded to a sensible precision, and the rounding rule should be stated in the table notes.

When to Use Figures

Figures are appropriate when the reader needs to see patterns, trends, distributions, or relationships that are difficult to perceive from a table of numbers. Figures work well for time series, survival curves, dose-response relationships, group comparisons with many categories, and data distributions.

The choice of figure type should match the data structure. A systematic review of 703 research articles in top physiology journals found that papers rarely included scatterplots, box plots, and histograms that allow readers to critically evaluate continuous data. Most papers presented continuous data in bar and line graphs. This is problematic because many different data distributions can lead to the same bar or line graph, and the full data may suggest different conclusions from the summary statistics. See the PLoS Biology review of data presentation paradigms for the analysis.

For small sample size studies, univariate scatterplots allow readers to see every data point and assess the distribution directly. Bar graphs with error bars hide the underlying distribution and can mislead when the data are skewed or contain outliers. The recommendation is to show the raw data points whenever the sample size is small enough to do so without clutter.

At a Glance: Matching Data Types to Presentation Formats

The following decision table matches common data types to appropriate presentation formats. The table is a starting point for planning a manuscript, and the final choice should account for the specific message the author wants to convey.

Data type Recommended format Example Notes
Small set of summary values (fewer than 5 numbers) Text "The median follow-up was 64 months (IQR 25 to 83 months)" Use text when the reader needs only the headline values
Multiple groups compared across several variables Table Baseline characteristics by treatment group Tables allow precise look-up and comparison
Time series or trend over time Figure with line graph or connected points Survival curves, growth curves, incidence trends Figures show the shape of the trend
Distribution of continuous data in small samples Figure with univariate scatterplot, box plot, or histogram Individual participant values by group Bar graphs hide the underlying distribution
Relationship between two continuous variables Figure with scatterplot Correlation between biomarker level and outcome Add a fitted line only when the model is justified
Large data set with many values Table in main text or supplementary material Full regression output, all study sites Use supplementary tables for very large data sets
Geographic or spatial patterns Figure with map Regional prevalence estimates Maps require clear legends and scale bars

The table above is adapted from the principle that the presentation format should preserve the reader's ability to evaluate the data. When the sample size is small, showing individual data points is preferable to showing only summary statistics. When the data set is large, tables allow precise look-up while figures allow pattern recognition.

Core Principles for Clear Data Presentation

Several principles apply across all presentation formats. These principles come from published guidance on statistical reporting and from observed failures in published papers.

Define Every Measure of Variability

Every measure of variability must be defined in the text, table notes, or figure legend. The reader needs to know whether the error bars or plus-minus values represent standard deviation, standard error of the mean, or a confidence interval. The audit of Journal of Neurophysiology papers found that 23% of papers included undefined measures of variability. See the F1000Research audit for the full findings.

Standard deviation describes the spread of the data. Standard error of the mean describes the precision of the sample mean as an estimate of the population mean. Confidence intervals describe the range of plausible values for the population parameter. These measures answer different questions, and the choice should match the message. When the goal is to describe the data, use standard deviation. When the goal is to make inferences about the population, use confidence intervals.

Report Exact P-Values

Exact p-values allow the reader to assess the strength of evidence against the null hypothesis. Reporting only "p < 0.05" or "p < 0.01" loses information and prevents the reader from applying their own significance thresholds. The analysis of papers in The Journal of Physiology and British Journal of Pharmacology found that 90 to 96% of papers did not report exact p-values for primary analyses and post-hoc tests. See the PLoS ONE analysis for details.

When p-values are very small, reporting "p < 0.001" is acceptable because the exact value is not meaningful beyond that threshold. For p-values between 0.001 and 0.10, report the exact value to two or three decimal places.

Do Not Misinterpret Borderline P-Values

P-values between 0.05 and 0.1 are frequently misinterpreted as trends or as evidence of a meaningful effect. The audits found that 56 to 64% of papers with p-values in this range interpreted them as trends or statistically significant. See the PLoS ONE analysis and the F1000Research audit for the findings.

A p-value above the chosen significance threshold means that the data are not strong enough to reject the null hypothesis at that threshold. It does not mean that a trend exists or that the effect is real but underpowered. If the author wants to discuss a borderline result, the discussion should acknowledge the uncertainty and report the confidence interval for the effect size.

Distinguish Association from Causation

The language used to describe results should match the study design. Observational studies can identify associations but cannot establish causation. The recommendations for writing about health inequality emphasize the importance of clarity when reporting association and causation. See the International Journal for Equity in Health recommendations for guidance on accurate presentation of evidence.

A retrospective cohort study can show that a treatment is associated with improved survival, but the association may be confounded by patient selection. The text should use language such as "was associated with" instead of "caused" unless the study design supports a causal claim.

Practical Workflow for Preparing Data Presentations

The following workflow guides researchers from raw data to finished tables and figures. The workflow assumes that the analysis is complete and the researcher is preparing the manuscript.

Step 1: Inventory the Data

List every data set that will appear in the manuscript. For each data set, record the number of observations, the number of variables, the type of each variable (continuous, categorical, ordinal), and the main message the data support. This inventory becomes the basis for deciding which format to use.

Step 2: Assign Each Data Set to a Format

Use the decision table in the At a Glance section to assign each data set to text, a table, or a figure. For data sets that could work in more than one format, choose the format that best serves the reader. When in doubt, prefer the format that shows more detail instead of less.

Step 3: Draft the Tables

Create each table with a clear title, labeled columns and rows, defined units, and footnotes for abbreviations and statistical conventions. Round values to a sensible precision and state the rounding rule. Check that every column has a header and that the table can be understood without reference to the text.

Step 4: Draft the Figures

Create each figure with a clear legend that describes the data, the sample size, the measure of variability, and the statistical test if applicable. Choose the figure type that matches the data structure. For small samples, use scatterplots or box plots instead of bar graphs. Label axes with the variable name and units.

Step 5: Write the Results Text

Write the results text to highlight the main findings and direct the reader to the tables and figures for complete data. The text should not duplicate every value in the tables and figures. State the finding, give the key numbers, and interpret the result in one or two sentences.

Step 6: Cross-Check Consistency

Check that every table and figure is mentioned in the text, that the numbers in the text match the numbers in the tables and figures, and that the statistical conventions are consistent throughout. Inconsistencies between text and tables are a common reason for reviewer comments.

Tables: Structure and Common Failure Patterns

Tables are the workhorse of scientific data presentation, but they fail when the structure does not match the data or when the reader cannot extract the needed values.

Table Structure

A well-structured table has a clear title, labeled columns and rows, defined units, and footnotes. The title should describe the content without requiring the reader to search the text for context. For example, "Baseline characteristics of 103 patients with limited-stage small cell lung cancer by treatment group" is more informative than "Patient characteristics."

Each column should contain one variable, and each row should contain one observation or one group. The first column typically contains the variable names, and the remaining columns contain the values for each group. The table should be sorted in a logical order, such as by importance or by the order the variables appear in the text.

Footnotes should explain abbreviations, define statistical conventions, and state the rounding rule. For example, a footnote might state "Values are mean (SD) unless otherwise indicated" or "Abbreviations: CI, confidence interval, HR, hazard ratio."

Common Failure Patterns in Tables

The audit of reporting practices identified several recurring problems. Undefined measures of variability appeared in 23% of papers. See the F1000Research audit for the full list of failures.

Other common failures include tables that are too wide to read, tables that duplicate information in the text, tables with inconsistent decimal places, and tables that omit the sample size for each group. A table without sample sizes prevents the reader from assessing the precision of the estimates.

A table with too many columns often needs to be split. For example, a table with 15 columns of biomarker values could be split into two tables organized by biomarker class, or restructured so that biomarkers are stacked in rows with a column for the group.

Figures: Choosing the Right Type

The choice of figure type should match the data structure and the message. The systematic review of physiology journals found that bar and line graphs dominated even when the data were continuous and the sample sizes were small. See the PLoS Biology review for the analysis of figure types across 703 papers.

Scatterplots for Small Samples

For continuous data with small sample sizes, univariate scatterplots show every data point. The reader can see the distribution, identify outliers, and assess whether the summary statistics are representative. A bar graph with error bars shows only the mean and a measure of variability, which can hide a bimodal distribution or a cluster of outliers.

The recommendation from the PLoS Biology review is to use scatterplots, box plots, and histograms for continuous data in small sample size studies. The authors provide Excel templates for making univariate scatterplots quickly. See the PLoS Biology review for the templates and guidance.

Line Graphs for Time Series

Line graphs are appropriate for data collected over time, such as growth curves, survival curves, or repeated measurements. The connected points show the shape of the trend and allow the reader to see when changes occur.

Survival curves are a specific type of line graph used in clinical research. A Kaplan-Meier curve shows the proportion of patients surviving over time, with steps at each event. The curve should include the number of patients at risk at each time point, usually shown below the x-axis. The example of limited-stage small cell lung cancer used Kaplan-Meier analysis to assess disease-free survival and overall survival. See the Frontiers in Oncology study for an example of survival curve presentation.

Box Plots for Distributions

Box plots show the median, quartiles, and outliers in a compact format. They are useful for comparing distributions across multiple groups. The box shows the interquartile range, the line inside the box shows the median, and the whiskers show the range or a defined multiple of the interquartile range.

Box plots are preferable to bar graphs when the goal is to show the distribution instead of the mean. However, box plots still hide the individual data points, so for very small samples, a scatterplot overlaid on the box plot is preferable.

Bar Graphs for Categorical Data

Bar graphs are appropriate for categorical data where the height of the bar represents a count, a proportion, or a summary statistic. Bar graphs are not appropriate for continuous data with small sample sizes because they hide the distribution.

When bar graphs are used, the error bars must be defined. The audit found that 76 to 84% of papers with plotted measures summarizing data variability used standard errors of the mean, and only 2 to 4% of papers plotted the raw data used to calculate variability. See the PLoS ONE analysis for the findings.

Annotated Examples of Data Presentation

The following examples show how the same data can appear in text, a table, and a figure. The examples use a hypothetical cohort of calves monitored for bovine respiratory disease, based on the design of a study that followed preweaned Holstein heifer calves for the first 12 weeks of life. See the PLoS ONE study on bovine respiratory disease for the study design.

Text Example

"Of the 121 calves enrolled across two years, 23 were classified as healthy, 31 had onset lobar consolidation, 17 had chronic lobar consolidation, and 18 had resolved lobar consolidation based on thoracic ultrasonography. Differential expression analysis identified 163 genes differentially expressed at disease onset compared to healthy calves, 27 genes in chronic disease, and no genes in resolved disease."

This text gives the reader the headline numbers and the main finding. The full data would appear in a table or figure.

Table Example

BRD stage Calves (n) Differentially expressed genes Top enriched pathways
Healthy 23 Reference Reference
Onset lobar consolidation 31 163 Immune response, resource allocation shift
Chronic lobar consolidation 17 27 Immune response
Resolved lobar consolidation 18 0 None

This table allows the reader to compare the stages and see the pattern of gene expression changes. The table would include footnotes defining the statistical threshold (FDR < 0.05 and |logFC| > 1) and the abbreviations.

Figure Example

A figure could show the number of differentially expressed genes by disease stage as a bar graph, or a scatterplot of the log fold change for each gene at onset. The scatterplot would allow the reader to see the distribution of effect sizes and identify the genes with the largest changes.

Records and Measurements for Data Presentation Quality

Researchers can assess the quality of their own data presentation by keeping records and checking against a standard checklist. The following measures provide a basis for self-assessment.

Statistical Reporting Checklist

A short instrument for assessing the quality of data analysis reporting was developed and tested on 160 original medical research articles. The instrument consists of nine questions that assess the quality of health research from a reader's perspective, with a total score ranging from 0 to 10. A high score indicated that an article had a good presentation of findings in tables and figures and that the description of analysis methods was helpful to readers. See the Applied Sciences instrument for the full checklist.

The checklist covers whether the analysis methods are described clearly, whether the results are presented in a way that allows verification, and whether the figures and tables are self-contained. Researchers can apply this checklist to their own manuscripts before submission.

Records to Keep

For each table and figure, keep a record of the following:

  • The data source and the analysis that produced the values
  • The sample size for each group or time point
  • The measure of variability and whether it is defined
  • The statistical test and the exact p-value or confidence interval
  • The software and version used for the analysis

These records allow the researcher to respond to reviewer questions and to verify the numbers during the revision process.

Common Failure Patterns and How to Avoid Them

The audits of published papers identified several recurring failure patterns. Knowing these patterns helps researchers avoid them in their own manuscripts.

Undefined Measures of Variability

The failure to define whether error bars or plus-minus values represent standard deviation, standard error, or confidence intervals appeared in 23% of audited papers. See the F1000Research audit for the finding.

The fix is to state the measure of variability in the table notes or figure legend and to use the same convention throughout the manuscript. If the text says "mean (SD)," the tables and figures should use the same format.

Misinterpretation of Borderline P-Values

P-values between 0.05 and 0.1 were misinterpreted as trends or significant results in 56 to 64% of audited papers. See the PLoS ONE analysis and the F1000Research audit for the findings.

The fix is to report the exact p-value and the confidence interval, and to describe the result as not statistically significant at the chosen threshold. If the author believes the result is clinically meaningful, the discussion should acknowledge the uncertainty and call for larger studies.

Duplication of Data in Text and Tables

Repeating every value from a table in the text wastes space and forces the reader to compare two versions of the same information. The text should highlight the main findings and direct the reader to the table for complete data.

The fix is to write the results text first, then check whether every sentence adds information beyond what appears in the tables and figures. If a sentence merely restates a table value, remove it or replace it with interpretation.

Bar Graphs for Continuous Data

Bar graphs hide the distribution of continuous data and can mislead when the data are skewed or contain outliers. The systematic review found that most papers presented continuous data in bar and line graphs even when the sample sizes were small. See the PLoS Biology review for the analysis.

The fix is to use scatterplots, box plots, or histograms for continuous data, especially when the sample size is small enough to show individual points.

Limitations of Data Presentation Methods

Each presentation format has limitations that researchers should acknowledge.

Text Limitations

Text can only convey a small number of values before becoming unreadable. Long lists of numbers in the text force the reader to parse the numbers from the prose, which is error-prone. Text is also poorly suited for showing patterns or distributions.

Table Limitations

Tables become difficult to read when they contain too many columns or rows. A table with more than ten columns often needs to be split. Tables are also poorly suited for showing trends over time or relationships between variables, which are better seen in figures.

Figure Limitations

Figures can distort data when the scale is misleading or when the figure type does not match the data structure. Bar graphs with error bars can hide the distribution of the data. Figures also require more effort to create and to verify than tables.

Data Extraction Limitations

Readers and meta-researchers increasingly use automated tools to extract data from published papers. A benchmark study of large language models for data extraction from scientific papers found that extraction accuracy ranged from 79.6% to 91.3% across models. Extraction performance was higher for explicit, verbatim information and worse for variables that required more complicated inference. See the Behavior Research Methods benchmark for the analysis.

This finding has implications for data presentation. When data are presented clearly and consistently, automated extraction is more likely to succeed. When data are buried in prose or presented inconsistently, automated extraction is more likely to fail. Researchers should present data in structured formats that are easy for both humans and machines to parse.

Reporting Guidelines and Standards

Several resources provide guidance for reporting research findings. These resources help researchers decide what to report and how to present it.

EQUATOR Network

The EQUATOR Network is an international initiative that provides reporting guidelines for health research. The network maintains a library of reporting guidelines for different study types, including randomized trials, observational studies, and systematic reviews. See the EQUATOR Network for the guideline library.

Authors should identify the appropriate reporting guideline for their study type and follow it during manuscript preparation. The guideline will specify which data to report and how to present them.

Research Data Framework

The National Institute of Standards and Technology maintains a Research Data Framework that describes the components of research data management, including data presentation and sharing. See the NIST Research Data Framework for the framework description.

The framework emphasizes that research data should be findable, accessible, interoperable, and reusable. These principles apply to the presentation of data in papers as well as to the sharing of raw data.

Experimental Design Assistant

The NC3Rs provides an Experimental Design Assistant that helps researchers plan experiments and avoid common design flaws. See the NC3Rs Experimental Design Assistant for the tool.

Good experimental design is a prerequisite for clear data presentation. When the design is flawed, no amount of presentation skill can fix the underlying problems.

FAIR Data Principles

The FAIR principles state that data should be findable, accessible, interoperable, and reusable. A special collection of data papers on COVID-19 research described datasets shared according to these principles. See the Journal of Open Psychology Data collection for examples of FAIR data sharing.

Researchers should consider how their data will be reused by others. Presenting data in tables and figures that are self-contained and well-defined supports reuse.

Welfare and Safety Context for Data Presentation

In animal research and clinical studies, data presentation has direct implications for welfare and safety. Poorly presented data can lead to incorrect conclusions that affect treatment decisions or animal management.

Clinical Decision-Making

In clinical research, the presentation of survival data and treatment outcomes directly affects clinical decisions. The study of limited-stage small cell lung cancer found that patients with stage I-IIA disease had significantly better survival than those with stage IIB-IIIB, with 5-year overall survival of 69.2% versus 47.1%. See the Frontiers in Oncology study for the survival analysis.

If these survival data were presented poorly, clinicians might draw incorrect conclusions about which patients benefit from surgery. Clear presentation of survival curves with the number at risk at each time point allows clinicians to assess the evidence.

Animal Health Monitoring

In animal research, the presentation of disease incidence and treatment response data affects herd management decisions. The study of bovine respiratory disease in dairy calves used thoracic ultrasonography to classify calves into disease stages and identified transcriptomic changes associated with disease onset. See the PLoS ONE study on bovine respiratory disease for the study design.

Clear presentation of disease classification data allows producers and veterinarians to identify calves that need treatment and to evaluate the effectiveness of prevention programs.

Escalation Criteria

When data presentation reveals unexpected patterns, researchers should escalate to appropriate professionals. For example, if a survival analysis shows an unexpected treatment effect, the researcher should consult a biostatistician before interpreting the result. If an animal study shows unexpected disease patterns, the researcher should consult a veterinarian.

The escalation criteria depend on the context. In clinical research, unexpected safety signals should be reported to the institutional review board or data safety monitoring board. In animal research, unexpected welfare problems should be reported to the institutional animal care and use committee.

Frequently Asked Questions

What is the difference between a table and a figure for data presentation?

A table presents data in rows and columns, allowing the reader to look up specific values and compare groups across variables. A figure presents data visually, allowing the reader to see patterns, trends, and distributions. Tables are better for precise values and many variables. Figures are better for showing the shape of the data and relationships between variables.

When should I use a bar graph versus a scatterplot?

Use a bar graph for categorical data where the height of the bar represents a count, proportion, or summary statistic. Use a scatterplot for continuous data where the reader needs to see the distribution of individual values. For small sample sizes, scatterplots are preferable because bar graphs hide the underlying distribution. See the PLoS Biology review for the analysis of figure types.

How do I decide whether to report standard deviation or standard error of the mean?

Standard deviation describes the spread of the data and is appropriate when the goal is to describe the sample. Standard error of the mean describes the precision of the sample mean as an estimate of the population mean and is appropriate when the goal is to make inferences. Whichever measure you choose, define it in the table notes or figure legend.

What should I do with p-values between 0.05 and 0.1?

Report the exact p-value and the confidence interval for the effect size. Describe the result as not statistically significant at the 0.05 threshold. Do not describe the result as a trend or as marginally significant unless you define what you mean by those terms. The audits found that 56 to 64% of papers misinterpreted p-values in this range. See the PLoS ONE analysis for the findings.

How many decimal places should I use in tables?

Round values to a sensible precision that matches the measurement accuracy. For most biological data, two or three decimal places are sufficient. State the rounding rule in the table notes. Excessive decimal places imply a precision that the measurement does not support.

Should I put all my data in the main text or in supplementary material?

Put the data that support the main findings in the main text. Put large data sets, full regression outputs, and detailed sensitivity analyses in supplementary material. The main text should contain enough data for the reader to evaluate the primary claims without searching the supplementary material.

How do I make my tables and figures accessible to readers with visual impairments?

Use high-contrast colors, label axes directly instead of relying on color alone, and provide descriptive captions that convey the main finding. Avoid using color as the only way to distinguish groups. Consider providing a text description of the figure content in the caption.

What reporting guidelines should I follow for my study type?

Identify the appropriate reporting guideline from the EQUATOR Network library. The guideline will specify which data to report and how to present them. Common guidelines include CONSORT for randomized trials, STROBE for observational studies, and PRISMA for systematic reviews.

Related Articles

References and Further Reading

This article is educational and does not replace institutional policy, professional advice, or applicable safety and regulatory requirements.