# Bayesian analysis mistakes biology

## Quick Answer

- Bayesian analysis in biology fails most often through priors chosen for convenience instead of biological evidence, producing results that reflect analyst assumptions more than data.
- Credible intervals are frequently misread as frequentist confidence intervals, leading researchers to draw conclusions about fixed parameters that the Bayesian framework does not support.
- The most important limitation is that posterior inferences depend on both the likelihood and the prior, so transparent reporting of both is required for any result to be reproducible or interpretable.

## At a Glance

| Common Bayesian Error | Typical Consequence | Practical Correction |
| --- | --- | --- |
| Using a diffuse prior without biological justification | Posterior estimates pulled toward implausible parameter values, especially with small sample sizes | Elicit priors from published measurements, pilot data, or domain knowledge and document the rationale |
| Interpreting credible intervals as confidence intervals | Overstated certainty about a fixed population parameter | Report the full posterior distribution and describe intervals as statements about the posterior probability of the parameter |
| Ignoring posterior predictive checks | Model inadequacy goes undetected and results are presented with false precision | Simulate replicate datasets from the fitted model and compare them to observed data |
| Failing to report convergence diagnostics | Markov chain Monte Carlo results may be unreliable without verification | Run multiple chains, inspect trace plots, and report Gelman-Rubin statistics or equivalent convergence checks |
| Selecting a model after seeing the data without accounting for the selection | Uncertainty is underestimated because model selection is not incorporated into inference | Use a validation scheme or explicitly acknowledge the selection process in the interpretation |

## Why Biologists Adopt Bayesian Methods

Biologists increasingly use Bayesian analysis because it provides a coherent framework for combining prior knowledge with observed data. The approach is especially attractive in fields where experiments are expensive, sample sizes are small, or measurements are noisy. In genomics, ecology, epidemiology, and evolutionary biology, Bayesian methods allow researchers to incorporate information from previous studies, account for multiple sources of uncertainty, and fit complex hierarchical models that are difficult to handle with classical statistics.

The core idea is straightforward. A prior distribution expresses what is known about a parameter before data are collected. The likelihood describes how the data are generated given the parameter. Bayes theorem combines these two components to produce a posterior distribution, which represents the updated state of knowledge after observing the data. This posterior distribution is the complete result of a Bayesian analysis, and all subsequent inference should be based on it.

The appeal for biological research is that the framework matches how scientists actually think. Researchers often have prior knowledge from earlier experiments, comparative studies, or mechanistic understanding of the system. Bayesian analysis provides a formal way to bring that knowledge into the statistical model. It also produces probability statements that are intuitive, such as the probability that a parameter falls within a given range, which can be easier to communicate than a frequentist p-value.

However, the flexibility of the Bayesian framework is also its main source of error. The prior is a modeling choice, not a fixed feature of the data. The likelihood is a modeling choice as well. Both choices affect the posterior, and both must be justified on biological and statistical grounds. When researchers treat these choices as defaults or fail to document them, the resulting analysis can be misleading.

The National Library of Medicine provides access to authoritative biomedical books and research-method references that describe the principles of statistical modeling and the importance of transparent reporting in research. These references emphasize that the validity of any statistical analysis depends on the assumptions being stated clearly and checked against the data.

## Core Principles of Bayesian Analysis

### The Prior Distribution

The prior distribution encodes what is known about a parameter before the current data are observed. In biological research, priors can come from several sources. They can be based on measurements from previous experiments, on theoretical constraints, on physical or biological limits, or on expert judgment. The prior should reflect genuine knowledge, not convenience.

A common error is to use a diffuse prior, such as a normal distribution with a very large variance, in the belief that this makes the analysis objective. In practice, a diffuse prior is still a modeling choice. It can place substantial probability on biologically implausible parameter values, and with small samples the posterior can be pulled toward those implausible values. A prior that is diffuse on the natural scale may be informative on the transformed scale, and vice versa.

Another common error is to use a prior that is too narrow, reflecting overconfidence in a particular value. This can happen when a researcher uses a prior from a single previous study without accounting for the uncertainty in that study. The result is a posterior that is too narrow and that understates the true uncertainty.

The choice of prior should be documented and justified. The justification should include the source of the prior information, the reasoning for the chosen distribution, and the sensitivity of the results to the prior. A sensitivity analysis, in which the prior is varied and the posterior is re-examined, is a standard way to assess how much the conclusions depend on the prior.

### The Likelihood Function

The likelihood function describes how the data are generated given the parameters. It is the statistical model that connects the parameters to the observations. The choice of likelihood is as important as the choice of prior, and it is often the more consequential choice.

In biological research, the likelihood must reflect the actual data-generating process. For count data, a Poisson or negative binomial likelihood may be appropriate. For continuous measurements, a normal or log-normal likelihood may be appropriate. For binary outcomes, a Bernoulli or binomial likelihood is used. The choice of likelihood should be based on the nature of the data and the biological process that produced it.

A common mistake is to choose a likelihood for computational convenience instead of biological realism. For example, a normal likelihood may be used for data that are counts, or a likelihood may be used when the data are overdispersed. These choices can lead to biased estimates and underestimated uncertainty.

The likelihood should be checked against the data. Posterior predictive checks, in which replicate datasets are simulated from the fitted model and compared to the observed data, are a standard way to assess whether the likelihood is adequate. If the observed data are extreme relative to the replicate datasets, the model is not adequate.

### The Posterior Distribution

The posterior distribution is the result of the analysis. It represents the full state of knowledge about the parameters after the data have been observed. All inferences should be based on the posterior distribution.

The posterior distribution can be summarized in various ways. The posterior mean or median is a point estimate. The posterior standard deviation is a measure of uncertainty. A credible interval is an interval that contains a specified probability of the posterior distribution. For example, a 95% credible interval is an interval that contains 95% of the posterior probability.

A common mistake is to interpret a credible interval as a confidence interval. The two are not the same. A confidence interval is a frequentist concept that describes the long-run frequency with which the interval contains the true parameter. A credible interval is a Bayesian concept that describes the posterior probability that the parameter lies in the interval. The interpretation of a credible interval is only valid within the Bayesian framework, and it depends on the prior.

## Common Mistakes in Bayesian Analysis

### Using Inappropriate Priors

The prior is the most distinctive component of a Bayesian analysis, and it is the component that is most often misused. The most common error is to use a prior that is not justified by biological evidence. This can take several forms.

A diffuse prior is often used to avoid the appearance of subjectivity. However, a diffuse prior is still a choice, and it can be informative in ways that are not obvious. For example, a uniform prior on a variance parameter is not the same as a uniform prior on the standard deviation. A prior that is diffuse on one scale can be informative on another scale.

An informative prior can be used when there is strong prior knowledge. However, the prior should be based on evidence, not on the researchers expectations. A prior that is too narrow can dominate the data and produce a posterior that is too certain. A prior that is too wide can allow the data to dominate, but it can also allow the posterior to be pulled toward implausible values.

The solution is to elicit priors from biological evidence. This can be done by reviewing previous studies, by using pilot data, or by consulting domain experts. The prior should be documented, and the sensitivity of the results to the prior should be assessed.

### Misinterpreting Credible Intervals

A credible interval is a summary of the posterior distribution. It is an interval that contains a specified amount of posterior probability. The interpretation of a credible interval is that the parameter has a specified probability of lying in the interval, given the data and the prior.

A common mistake is to interpret a credible interval as a confidence interval. A confidence interval is a frequentist concept that describes the long-run frequency of the interval containing the true parameter. The two concepts are different, and the interpretation of a credible interval depends on the prior.

Another common mistake is to interpret a credible interval as a statement about the probability of a fixed effect. In the Bayesian framework, the parameter is treated as a random variable, and the posterior distribution describes the uncertainty about the parameter. The credible interval is a statement about the posterior distribution, not about the true value of the parameter.

The solution is to report the full posterior distribution and to describe the credible interval in terms of the posterior probability. The interpretation should be clear that the interval is a statement about the posterior distribution, not about the true value of the parameter.

### Ignoring Model Checking

A Bayesian analysis is not complete when the posterior is computed. The model must be checked against the data. This is done with posterior predictive checks, in which replicate datasets are simulated from the fitted model and compared to the observed data.

A common mistake is to skip the model checking step and to present the posterior as the final result. This can lead to a model that is inadequate but is presented as valid. The model may be inadequate because the likelihood is wrong, because the prior is wrong, or because the model structure is wrong.

The solution is to run posterior predictive checks and to report the results. The checks should be described in the methods, and the results should be reported in the results. If the model is inadequate, the model should be revised.

### Failing to Report Convergence

Bayesian analyses are often computed with Markov chain Monte Carlo (MCMC) methods. These methods produce a sample from the posterior distribution, and the sample is used to estimate the posterior summaries. The sample must be large enough and the chain must be converged for the estimates to be reliable.

A common mistake is to run a single chain and to report the results without checking convergence. The chain may not have converged, and the results may be unreliable. The chain may be too short, and the estimates may be noisy.

The solution is to run multiple chains, to check convergence with diagnostics, and to report the diagnostics. The effective sample size should be reported, and the trace plots should be examined. If the chains have not converged, the model should be adjusted or the number of iterations should be increased.

### Overlooking the Effect of the Prior on the Posterior

The posterior is a combination of the prior and the likelihood. The relative influence of the prior and the likelihood depends on the sample size and the strength of the prior. With a small sample, the prior can dominate the posterior. With a large sample, the likelihood can dominate the posterior.

A common mistake is to ignore the effect of the prior on the posterior. The prior can be informative, and the posterior can be a reflection of the prior more than the data. This is especially a problem with small samples.

The solution is to report the prior and the likelihood separately, and to describe the influence of the prior on the posterior. A sensitivity analysis, in which the prior is varied, can be used to assess the influence of the prior.

## Practical Workflow for a Valid Bayesian Analysis

### Step 1: Define the Biological Question

The first step is to define the biological question in terms of the parameters of interest. The question should be specific and should be answerable with the data. The parameters of interest should be identified, and the model should be designed to estimate those parameters.

### Step 2: Specify the Likelihood

The likelihood should be specified based on the nature of the data and the biological process that generated it. The likelihood should be checked against the data, and the model should be adjusted if the likelihood is not adequate.

### Step 3: Elicit the Prior

The prior should be elicited from biological evidence. The prior should be documented, and the rationale should be stated. The prior should be checked for sensitivity, and the results should be reported for a range of priors.

### Step 4: Compute the Posterior

The posterior should be computed with a reliable method. The method should be documented, and the convergence should be checked. The effective sample number should be reported, and the diagnostics should be reported.

### Step 5: Check the Model

The model should be checked with posterior predictive checks. The observed data should be compared to the replicate datasets, and the model should be adjusted if the observed data is extreme.

### Step 6: Report the Results

The results should be reported in a transparent way. The prior, the likelihood, the posterior, and the model checks should be reported. The results should be interpreted in terms of the posterior distribution, and the limitations should be stated.

## Records and Measurements

### Documenting the Prior

The prior should be documented in the methods section of the report. The documentation should include the source of the prior information, the distribution of the prior, and the parameters of the prior. The rationale for the prior should be stated, and the sensitivity of the results to the prior should be reported.

### Documenting the Likelihood

The likelihood should be documented in the methods section. The documentation should include the distribution of the likelihood, the parameters of the likelihood, and the justification for the choice. The likelihood should be checked against the data, and the results of the check should be reported.

### Documenting the Posterior

The posterior should be documented in the results section. The documentation should include the point estimate, the uncertainty, and the credible interval. The posterior distribution should be described, and the interpretation should be stated.

### Documenting the Model Checks

The model checks should be documented in the results section. The checks should include the posterior predictive checks and the convergence diagnostics. The results of the checks should be reported, and the implications for the model should be stated.

## Common Failure Patterns

### The Prior Dominates the Posterior

A common failure pattern is that the prior dominates the posterior. This happens when the prior is too strong or when the sample is too small. The posterior is a reflection of the prior, and the data has little influence. The result is an analysis that is not informative about the data.

The solution is to use a prior that is based on biological evidence and to check the sensitivity of the results to the prior. If the prior is too strong, the prior should be weakened or the sample should be increased.

### The Likelihood Is Inadequate

A common failure pattern is that the likelihood is inadequate for the data. The likelihood may be too simple, or the likelihood may be the wrong distribution. The model does not fit the data, and the posterior is not a valid summary of the data.

The solution is to check the likelihood against the data and to adjust the likelihood if it is not adequate. The model should be checked with posterior predictive checks, and the model should be adjusted if the observed data is extreme.

### The Model Is Not Identifiable

A common failure pattern is that the model is not identifiable. The parameters cannot be estimated from the data, and the posterior is not well-defined. The model may have too many parameters, or the parameters may be correlated.

The solution is to simplify the model or to add more data. The model should be checked for identifiability, and the parameters should be estimated with a sufficient sample.

### The Results Are Not Reproducible

A common failure pattern is that the results are not reproducible. The analysis is not documented, and the results cannot be reproduced by another researcher. The prior, the likelihood, and the posterior are not reported, and the analysis is not transparent.

The solution is to document the analysis and to report the results in a transparent way. The prior, the likelihood, and the posterior should be reported, and the analysis should be reproducible.

## Limitations and Interpretation

### The Prior Is a Choice

The prior is a choice, and the choice is not unique. Different priors can lead to different posteriors, and the results are not unique. The prior should be justified, and the sensitivity of the results to the prior should be reported.

### The Likelihood Is a Choice

The likelihood is a choice, and the choice is not unique. Different likelihoods can lead to different posteriors, and the results are not unique. The likelihood should be justified, and the sensitivity of the results to the likelihood should be reported.

### The Posterior Is a Combination

The posterior is a combination of the prior and the likelihood. The posterior is not a property of the data alone. The posterior is a property of the data, the prior, and the likelihood. The posterior should be interpreted in the context of the prior and the likelihood.

### The Credible Interval Is Not a Confidence Interval

The credible interval is not a confidence interval. The credible interval is a statement about the posterior distribution, and the interpretation is different from the interpretation of a confidence interval. The credible interval should be interpreted in the context of the posterior distribution.

## Reporting Guidelines and Transparency

### Reporting the Prior

The prior should be reported in the methods section. The prior should be described in terms of the distribution and the parameters. The rationale for the prior should be stated, and the source of the prior information should be stated.

### Reporting the Likelihood

The likelihood should be reported in the methods section. The likelihood should be described in terms of the distribution and the parameters. The rationale for the likelihood should be stated, and the justification for the likelihood should be stated.

### Reporting the Posterior

The posterior should be reported in the results section. The posterior should be described in terms of the point estimate, the uncertainty, and the credible interval. The interpretation of the posterior should be stated.

### Reporting the Model Checks

The model checks should be reported in the results section. The checks should include the posterior predictive checks and the convergence results. The results of the checks should be reported, and the implications for the model should be stated.

### Transparency and Reproducibility

The analysis should be transparent and reproducible. The prior, the likelihood, and the posterior should be documented, and the analysis should be reproducible. The data should be available, and the code should be available. The analysis should be reported in a way that allows another researcher to reproduce the results.

The [EQUATOR Network](https://www.equator-network.org/) provides reporting guidelines for research studies, and the guidelines can be used to improve the transparency and reproducibility of the analysis. The guidelines should be used to ensure that the analysis is reported in a complete and transparent way.

The [Committee on Publication Ethics](https://publicationethics.org/core-practices) provides guidance on publication ethics, and the guidance should be used to ensure that the analysis is reported in an ethical way. The guidance covers authorship, peer review, data, conflicts, and misconduct.

## Data Management and Sharing

### Data Management

The data should be managed in a way that is consistent with the analysis. The data should be documented, and the data should be available for the analysis. The data should be stored in a way that is secure and accessible.

### Data Sharing

The data should be shared in a way that is consistent with the analysis. The data should be available for the analysis, and the data should be shared in a way that is consistent with the analysis. The data should be shared in a way that is consistent with the analysis.

The [NIH Data Management and Sharing Policy](https://sharing.nih.gov/data-management-and-sharing-policy) provides guidance on data management and sharing. The policy should be used to ensure that the data is managed and shared in a way that is consistent with the analysis.

### Data Availability

The data should be available for the analysis. The data should be available in a way that is consistent with the analysis. The data should be available in a way that is consistent with the analysis.

## Professional Escalation Criteria

### When to Seek Help

A researcher should seek help when the analysis is not working. The analysis may not be working when the model does not converge, when the posterior is not sensible, or when the model checks fail. The researcher should seek help from a statistician or a bioinformatician.

### When to Revise the Model

The model should be revised when the model checks fail. The model should be revised when the posterior is not sensible, or when the model is not identifiable. The model should be revised when the analysis is not reproducible.

### When to Report the Limitations

The limitations should be reported when the analysis is not complete. The limitations should be reported when the prior is not justified, when the likelihood is not adequate, or when the model checks fail. The limitations should be reported in the discussion.

## A Decision Framework for Prior Elicitation and Sensitivity Auditing

The most consequential choice in any Bayesian analysis is the prior, yet biologists often select priors through habit, convenience, or a vague sense of what seems reasonable. A diffuse prior is not automatically safe, and an informative prior is not automatically biased. What matters is whether the prior can be defended with evidence and whether the conclusions survive reasonable variation in that prior. This section provides a practical decision framework for prior elicitation, a structured sensitivity auditing protocol, and a record system that makes the reasoning behind prior choices explicit and reviewable.

### The Prior Elicitation Decision Tree

Before writing a single line of model code, work through a structured decision process that forces the prior to be connected to biological evidence instead of statistical convention. The decision tree has four branches, each leading to a different prior construction strategy.

**Branch 1: Direct measurement exists.** If previous experiments in your system have measured the parameter of interest directly, use those measurements to construct the prior. Extract the mean and standard error from published studies, or use raw data if available. Construct a normal prior centered on the pooled estimate with variance derived from the standard error. Document the specific studies used and the date of the literature search. This branch is appropriate when the measurement conditions in the prior studies match your current experimental conditions.

**Branch 2: Mechanistic constraint exists.** If the parameter is bounded by biological or physical principles, construct a prior that respects those bounds. For example, a survival probability must lie between zero and one, a growth rate cannot exceed the maximum observed for the species, or a concentration cannot be negative. Use a beta distribution for probabilities, a log-normal distribution for positive quantities, or a truncated distribution when bounds are known. The key is that the bounds come from biological reasoning, not from computational convenience.

**Branch 3: Related system measurement exists.** If direct measurements are unavailable but measurements exist in a related species, a related tissue, or a related experimental condition, use those measurements with an expanded variance to reflect the additional uncertainty from extrapolation. The expansion factor should be justified. A common approach is to double the variance from the related study, but the justification must be stated explicitly. This branch requires careful documentation of why the related system is informative and how much uncertainty the extrapolation adds.

**Branch 4: No direct evidence exists.** If no direct or related measurements exist, you face the most difficult situation. The honest response is not to default to a diffuse prior but to construct a weakly informative prior that respects the scale of the data and the plausible range of the parameter. A weakly informative prior is one that rules out absurd values but does not strongly favor any particular value within the plausible range. For a regression coefficient, this might be a normal prior centered at zero with a scale chosen from the range of the predictor and outcome variables. The scale should be justified by the measurement units, not by a default setting in software.

After selecting a branch, document the reasoning in a prior specification table. The table should include the parameter name, the chosen distribution, the distribution parameters, the evidence source, the date of the evidence, and the reasoning for any variance expansion. This table becomes part of the analysis record and should be included in the supplementary materials of any report.

### Sensitivity Auditing Protocol

A prior sensitivity analysis is not optional in biological research. It is the primary defense against the criticism that results are driven by assumptions instead of data. The auditing protocol has three levels, each answering a different question about the robustness of the conclusions.

**Level 1: Local sensitivity.** Vary each prior parameter by a small amount, typically 10 to 20 percent, and re-run the analysis. Record how much the posterior mean and credible interval change. This level detects whether the posterior is stable under minor perturbations of the prior. If the posterior mean shifts by more than the posterior standard deviation, the analysis is sensitive to the prior at the local level.

**Level 2: Structural sensitivity.** Replace the prior distribution with a different distribution that has the same mean and variance. For example, replace a normal prior with a Student t prior or a logistic prior. This level detects whether the conclusions depend on the specific distributional form of the prior. If the posterior changes substantially, the conclusions are sensitive to the distributional assumption.

**Level 3: Global sensitivity.** Replace the prior with a weakly informative prior and with a diffuse prior, and compare the posteriors. This level detects whether the conclusions depend on the strength of the prior information. If the conclusions change direction or lose statistical support under a weakly informative prior, the analysis is not robust.

For each level, record the prior specification, the posterior mean, the posterior standard deviation, and the 95 percent credible interval. Present the results in a sensitivity table that allows a reader to see at a glance how the conclusions change across prior specifications. The table should be included in the results section or supplementary materials, and the interpretation should state clearly which conclusions are robust and which depend on the prior.

### The Prior Influence Ratio

A practical metric for communicating prior influence is the prior influence ratio, which compares the posterior precision to the prior precision. Precision is the inverse of variance. The ratio is calculated as the posterior precision divided by the prior precision. A ratio near one indicates that the prior dominates the posterior. A ratio much larger than one indicates that the data dominate the prior.

This ratio is useful because it provides a single number that summarizes the balance of influence between prior and data. It is not a formal test and has no threshold for acceptability, but it helps researchers and readers understand whether the analysis is data-driven or prior-driven. Report the ratio for each parameter of interest and interpret it in the context of the sample size and the strength of the prior evidence.

### A Record System for Prior Justification

The record system for prior justification should be maintained from the start of the analysis, not reconstructed at the end. Create a prior log with one entry per parameter. Each entry should contain the following fields.

**Parameter identifier.** The name of the parameter as it appears in the model code and in the report.

**Biological definition.** A plain language description of what the parameter measures in the biological system.

**Elicitation branch.** The branch of the decision tree that was followed, from direct measurement to no evidence.

**Evidence source.** The specific citation, dataset, or expert consultation that informed the prior. Include the date the evidence was gathered.

**Distribution and parameters.** The exact distribution and parameter values used in the model.

**Variance expansion justification.** If the variance was expanded from a related study, the reasoning and the expansion factor.

**Sensitivity audit results.** The results of the local, structural, and global sensitivity analyses.

**Date of last update.** The date the prior was last revised, with a note on what changed.

This log serves multiple purposes. It forces the researcher to articulate the reasoning behind each prior choice. It provides the material needed for transparent reporting in the methods section. It allows reviewers to assess whether the priors are justified. And it enables reproducibility, because another researcher can see exactly what was assumed and why.

### Common Failure Patterns in Prior Elicitation

Several failure patterns recur when biologists attempt to justify their priors. Recognizing these patterns helps avoid them.

**The literature cherry-pick.** The researcher selects a single study that supports their expected result and uses it as the prior, ignoring studies with conflicting estimates. The solution is to conduct a systematic search and to use a pooled estimate from all relevant studies, or to use the range of estimates to set the prior variance.

**The false precision trap.** The researcher uses a prior from a large study without accounting for the fact that the prior study conditions differ from the current conditions. The prior is too narrow because the variance does not include the additional uncertainty from extrapolation. The solution is to expand the variance when applying priors across conditions.

**The diffuse prior fallacy.** The researcher uses a diffuse prior to avoid criticism of subjectivity, but the diffuse prior is still a choice and can be informative on the transformed scale. The solution is to use a weakly informative prior that respects the scale of the data and to document the reasoning.

**The expert overconfidence problem.** The researcher consults a domain expert who provides a narrow range for the parameter, but the expert is overconfident and the range is too narrow. The solution is to consult multiple experts and to use the range of their estimates to set the prior variance.

**The sensitivity analysis avoidance.** The researcher skips the sensitivity analysis because it is time-consuming or because the results are not robust. The solution is to run the sensitivity analysis regardless and to report the results honestly, even if they show that the conclusions depend on the prior.

### Integration with Reporting Guidelines

The prior elicitation and sensitivity auditing process should be reported in the methods section of any manuscript. The [EQUATOR Network](https://www.equator-network.org/) provides reporting guidelines that can help structure this reporting. The guidelines emphasize transparency about all modeling choices, including priors. The prior log and sensitivity table should be included as supplementary materials, and the methods section should summarize the elicitation branch for each parameter and the results of the sensitivity audit.

The [Committee on Publication Ethics](https://publicationethics.org/core-practices) core practices address the importance of accurate reporting and transparency in research. A prior that is not justified or a sensitivity analysis that is not reported is a form of incomplete reporting that undermines the validity of the conclusions. The record system described here provides the documentation needed to meet these ethical standards.

### Practical Implementation Steps

Implement the decision framework in the following order.

**Step 1: Inventory the parameters.** List every parameter in the model and classify each one as a parameter of interest or a nuisance parameter. The prior elicitation effort should focus on the parameters of interest, but nuisance parameters also need priors that are at least weakly informative.

**Step 2: Gather evidence.** For each parameter of interest, search the literature for direct measurements, mechanistic constraints, or related system measurements. Record the evidence in the prior log.

**Step 3: Construct the prior.** Follow the decision tree to select the prior distribution and parameters. Document the reasoning in the prior log.

**Step 4: Run the analysis.** Fit the model and record the posterior summaries.

**Step 5: Run the sensitivity audit.** Perform the local, structural, and global sensitivity analyses. Record the results in the sensitivity table.

**Step 6: Interpret and report.** Calculate the prior influence ratio for each parameter. Interpret the conclusions in light of the sensitivity results. Report the prior log and sensitivity table in the supplementary materials.

### When to Escalate to a Statistician

The decision framework is designed to be used by biologists, but there are situations where professional statistical help is needed. Escalate to a statistician when the sensitivity audit shows that conclusions change direction under reasonable prior variation, when the prior influence ratio is near one for parameters of interest, when the model is not identifiable and the prior is the only thing constraining the parameters, or when expert elicitation produces conflicting priors that cannot be reconciled. A statistician can help design a more sophisticated elicitation protocol, implement a hierarchical prior that accounts for between-study variation, or restructure the model to reduce dependence on the prior.

The [National Library of Medicine](https://www.ncbi.nlm.nih.gov/books) provides access to authoritative references on statistical modeling and research methods that can support the prior elicitation process. These references describe the principles of Bayesian inference and the importance of transparent modeling choices. The [NIH Grants and Funding](https://grants.nih.gov/) pages describe the expectations for rigorous statistical methods in funded research, and the [ORCID for Researchers](https://info.orcid.org/researchers) pages describe the importance of maintaining a complete record of research outputs and methods. These resources support the broader goal of conducting Bayesian analyses that are defensible, transparent, and reproducible.

## Frequently Asked Questions

### What is the difference between a prior and a likelihood?

The prior is the distribution of the parameter before the data is observed. The likelihood is the distribution of the data given the parameter. The prior and the likelihood are combined to produce the posterior distribution.

### How do I choose a prior for my biological data?

The prior should be chosen based on biological evidence. The prior should be based on previous studies, pilot data, or expert judgment. The prior should be documented, and the rationale should be stated.

### What is a credible interval and how do I interpret it?

A credible interval is an interval that contains a specified amount of posterior probability. The credible interval is a statement about the posterior distribution, and the interpretation is based on the posterior distribution.

### How do I check if my Bayesian model is adequate?

The model should be checked with posterior predictive checks. The observed data should be compared to the replicate datasets, and the model should be adjusted if the observed data is extreme.

### How do I check if my Markov chain has converged?

The convergence should be checked with diagnostics. The diagnostics should be reported, and the effective sample number should be reported. If the chain has not converged, the model should be adjusted or the number of iterations should be increased.

### What should I report in a Bayesian analysis?

The prior, the likelihood, the posterior, and the model checks should be reported. The analysis should be transparent and reproducible, and the results should be interpreted in the context of the posterior distribution.

### How do I handle the sensitivity of the results to the prior?

The sensitivity of the results to the prior should be assessed with a sensitivity analysis. The prior should be varied, and the posterior should be examined. The results should be reported for a range of priors.

### When should I seek help from a statistician?

A statistician should be consulted when the analysis is not working. The analysis may not be working when the model does not converge, when the posterior is not sensible, or when the model checks fail.

## Using the Evidence

| Source | Best use in this topic | Important limitation |
|---|---|---|
| [Research Methods Resources](https://www.ncbi.nlm.nih.gov/books) | official guidance | Check the linked page for current local requirements |
| [EQUATOR Network](https://www.equator-network.org/) | official guidance | Check the linked page for current local requirements |
| [Core Practices](https://publicationethics.org/core-practices) | official guidance | Check the linked page for current local requirements |

## Related Bioinformatics Guides

- [Metabolomics Data Analysis Workflow: From Raw Data to Biological Insight](/knowledge/bioinformatics/metabolomics-data-analysis-workflow-from-raw-data-to-biological-insight)
- [Metagenomics Data Analysis: From Raw Reads to Biological Insights](/knowledge/bioinformatics/metagenomics-data-analysis-from-raw-reads-to-biological-insights)
- [Proteomics Data Analysis Workflow: From Raw Spectra to Biological Insights](/knowledge/bioinformatics/proteomics-data-analysis-workflow-from-raw-spectra-to-biological-insights)
- [Spatial Omics Data Analysis: From Image Processing to Biological Interpretation](/knowledge/bioinformatics/spatial-omics-data-analysis-from-image-processing-to-biological-interpretation)
- [Single-Cell Sequencing Analysis Pipeline: From Raw Data to Biological Insights](/knowledge/bioinformatics/single-cell-sequencing-analysis-pipeline-from-raw-data-to-biological-insights)

## References and Further Reading

- [Research Methods Resources](https://www.ncbi.nlm.nih.gov/books). National Library of Medicine.
- [EQUATOR Network](https://www.equator-network.org/). EQUATOR Network.
- [Core Practices](https://publicationethics.org/core-practices). Committee on Publication Ethics.
- [NIH Grants and Funding](https://grants.nih.gov/). National Institutes of Health.
- [ORCID for Researchers](https://info.orcid.org/researchers). ORCID.
- [Data Management and Sharing Policy](https://sharing.nih.gov/data-management-and-sharing-policy). National Institutes of Health.
- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Quantifying chromosomal instability from intratumoral karyotype diversity using agent-based modeling and Bayesian inference.](https://pubmed.ncbi.nlm.nih.gov/35380536). eLife, 2022.
- [Phylogenies and Diversification Rates: Variance Cannot Be Ignored.](https://pubmed.ncbi.nlm.nih.gov/30481343). Systematic biology, 2019.
- [Automated Classification of Circulating Tumor Cells and the Impact of Interobsever Variability on Classifier Training and Performance.](https://pubmed.ncbi.nlm.nih.gov/26504857). Journal of immunology research, 2015.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.