# Bayes' Theorem in Biology: Worked Examples
## Key Takeaways

- Bayes' theorem quantifies how new evidence, such as a positive RT-PCR result for a rare pathogen with 5% prevalence, updates an initial belief (prior probability) to a revised belief (posterior probability). Even with 95% sensitivity and 90% specificity, a positive result in this low-prevalence scenario yields a posterior probability of only 33.3%, highlighting the impact of the base rate on diagnostic interpretation.
- In environmental DNA (eDNA) studies for species detection, a positive eDNA signal from a rare invasive species with 10% site occupancy, even with 90% sensitivity and 95% specificity, results in a posterior probability of 66.7%. This necessitates confirmation with additional sampling or a different detection method to mitigate the risk of false positives.
- The prior probability, representing the initial prevalence or base rate (e.g., disease prevalence, species occupancy), is a critical determinant of the posterior probability. A sensitivity analysis demonstrating how the posterior probability changes with varying priors is essential for assessing the robustness of conclusions, particularly in low-prevalence settings.
- Sequential testing, where the posterior probability from one test becomes the prior for the next, is a powerful workflow. For instance, a second positive RT-PCR test (98% sensitivity, 95% specificity) following an initial positive result with a 33.3% posterior probability increases the confidence of a true positive to 90.8%.
- Decision thresholds for action (e.g., initiating treatment, declaring species presence) should be pre-defined based on the consequences of false positives and false negatives, rather than being set after data collection. This pre-specification is crucial for maintaining research integrity and avoiding bias in interpreting diagnostic or detection outcomes.

---

## Quick Answer

- Bayes' theorem updates a prior probability using new evidence, producing a posterior probability that reflects both the initial belief and the observed data.
- Apply Bayes' theorem in biology by defining the prior, the sensitivity and specificity of your test or detection method, and the prevalence or base rate in your study population.
- The posterior probability depends heavily on the prior and the test characteristics, so a positive result from a low-specificity test in a low-prevalence setting can still yield a low posterior probability.

## At a Glance

| Application | Prior Probability | Evidence (Test Result) | Posterior Probability | Interpretation |
| --- | --- | --- | --- | --- |
| Diagnostic test for a rare disease | 1% prevalence | Positive result, 95% sensitivity, 95% specificity | 16.1% | Most positive results are false positives when prevalence is low |
| Diagnostic test for a common disease | 30% prevalence | Positive result, 95% sensitivity, 95% specificity | 89.0% | Positive result is more informative when prevalence is higher |
| Species detection using environmental DNA | 10% site occupancy | Positive detection, 90% sensitivity, 95% specificity | 66.7% | A positive eDNA result requires confirmation when occupancy is low |

## The Structure of Bayes' Theorem in Biological Research

Bayes' theorem provides a formal rule for revising a probability estimate when new information becomes available. In biological research, the theorem is used to move from a prior probability, which is the initial degree of belief in a hypothesis before collecting data, to a posterior probability, which is the updated belief after observing evidence. The theorem is expressed as:

P(H|E) = [P(E|H) × P(H)] / P(E)

In this expression, P(H) is the prior probability of the hypothesis, P(E|H) is the likelihood of observing the evidence if the hypothesis is true, and P(E) is the total probability of the evidence under all possible hypotheses. The result, P(H|E), is the posterior probability of the hypothesis given the evidence.

The denominator P(E) is often the source of confusion in biological applications. It is not simply the probability of the evidence in isolation. It is the weighted sum of the probability of the evidence under the hypothesis and the probability of the evidence under the alternative hypothesis, with weights determined by the prior probabilities. In diagnostic testing, this denominator accounts for both true positives and false positives, which is why the prevalence of a condition in the population matters so much.

For a biologist, the practical value of Bayes' theorem is that it forces explicit consideration of the base rate. Many biological tests, whether they are molecular assays, imaging methods, or field detection techniques, are imperfect. A test with high sensitivity and high specificity can still produce a large number of false positives when the condition being tested is rare. Bayes' theorem quantifies this effect and provides a defensible number for the probability that a positive result is a true positive.

## Worked Example 1: Diagnostic Test Accuracy for a Pathogen

Consider a scenario in which a laboratory is evaluating a new PCR assay for detecting a bacterial pathogen in tissue samples. The assay has a sensitivity of 95 percent, meaning it correctly identifies 95 percent of samples that truly contain the pathogen. The assay has a specificity of 90 percent, meaning it correctly returns a negative result for 90 percent of samples that do not contain the pathogen. The prevalence of the pathogen in the submitted sample population is 5 percent.

The prior probability is the prevalence, which is 0.05. The likelihood of a positive test result if the pathogen is present is the sensitivity, which is 0.95. The total probability of a positive test result is the sum of the true positive rate and the false positive rate. The true positive rate is 0.05 multiplied by 0.95, which equals 0.0475. The false positive rate is 0.95 multiplied by 0.10, which equals 0.095. The total probability of a positive result is 0.1425.

The posterior probability is 0.0475 divided by 0.1425, which equals 0.333. This means that a positive PCR result in this population has only a 33.3 percent chance of being a true positive. The other 66.7 percent of positive results are false positives.

This result is counterintuitive to many researchers because the test appears to be quite accurate. The issue is the low prevalence. When the condition is rare, even a small false positive rate produces many more false positives than true positives. The posterior probability of 33.3 percent is the correct interpretation of a positive result in this context.

The practical implication for a laboratory professional is that a positive result from a low-prevalence population should be confirmed with a second test that has a different biological basis, or the result should be interpreted with caution. The posterior probability from the first test becomes the prior probability for the second test. If the second test has a sensitivity of 98 percent and a specificity of 95 percent, the posterior probability after the second positive result is calculated using the prior of 0.333. The true positive rate is 0.333 multiplied by 0.98, which equals 0.326. The false positive rate is 0.667 multiplied by 0.05, which equals 0.033. The total probability of a positive result is 0.359. The posterior probability is 0.326 divided by 0.359, which equals 0.908. After two positive tests, the probability that the sample truly contains the pathogen is 90.8 percent.

This sequential updating is a direct application of Bayes' theorem and is the reason the method is so valuable in diagnostic workflows. Each new piece of evidence refines the probability, and the prior for each step is the posterior from the previous step.

## Worked Example 2: Species Detection Using Environmental DNA

Environmental DNA, or eDNA, is a method used to detect the presence of a species by analyzing DNA shed into the environment. This method is used in conservation biology to detect rare or invasive species in water samples. The method is sensitive but not perfect, and the interpretation of a positive eDNA result requires Bayesian reasoning.

Consider a study designed to detect the presence of an invasive fish species in a network of lakes. The prior probability that a randomly selected lake contains the species is 10 percent, based on previous surveys and the known dispersal rate of the species. The eDNA assay has a sensitivity of 90 percent, meaning it detects the species in 90 percent of lakes where the species is present. The assay has a specificity of 95 percent, meaning it returns a negative result in 95 percent of lakes where the species is absent.

The prior probability is 0.10. The true positive rate is 0.10 multiplied by 0.90, which equals 0.09. The false positive rate is 0.90 multiplied by 0.05, which equals 0.045. The total probability of a positive result is 0.135. The posterior probability is 0.09 divided by 0.135, which equals 0.667. A positive eDNA result in this study has a 66.7 percent chance of being a true positive.

The interpretation of this result is that a positive eDNA detection is suggestive but not conclusive. The researcher should collect additional samples from the same lake, or use a different detection method, before declaring the species present. The posterior probability of 66.7 percent becomes the prior for the next sample. If a second independent eDNA sample from the same lake is positive, the posterior probability increases substantially.

The Bayesian approach also helps in the design of the study. If the researcher wants to achieve a posterior probability of at least 95 percent before declaring the species present, the number of replicate samples can be calculated. Each positive sample updates the posterior, and the researcher can determine how many positive samples are needed given the sensitivity, specificity, and prior. This is a direct application of the theorem to study design and decision thresholds.

## Worked Example 3: Gene Expression Classification

In gene expression studies, researchers often classify samples into categories such as diseased or healthy based on the expression levels of a panel of genes. A classifier is trained on a set of known samples and then applied to new samples. The output of the classifier is a probability that the new sample belongs to a particular class. Bayes' theorem provides the framework for interpreting this probability.

Suppose a classifier has been developed to distinguish between two types of tumor tissue based on gene expression profiles. The classifier has a sensitivity of 85 percent and a specificity of 92 percent. The prevalence of the more aggressive tumor type in the patient population is 20 percent.

The prior probability is 0.20. The true positive rate is 0.20 multiplied by 0.85, which equals 0.17. The false positive rate is 0.80 multiplied by 0.08, which equals 0.064. The total probability of a positive classification is 0.234. The posterior probability is 0.17 divided by 0.234, which equals 0.726. A sample classified as the aggressive type has a 72.6 percent chance of being correctly classified.

The classifier is not perfect, and the posterior probability of 72.6 percent is the correct interpretation of the classifier output. The researcher should not treat the classifier output as a definitive diagnosis. The posterior probability should be combined with other clinical or pathological information before a final decision is made.

This example illustrates that Bayes' theorem is not limited to diagnostic tests. It applies to any classification or prediction method that produces a probability. The prior is the prevalence of the condition in the population, and the likelihood is the sensitivity and specificity of the classifier.

## The Role of the Prior Probability

The prior probability is the starting point for any Bayesian analysis. In biological research, the prior can come from several sources. It can be the prevalence of a disease in a population, the proportion of sites occupied by a species, or the probability of a gene being differentially expressed based on previous experiments. The choice of the prior is a critical decision because it directly influences the posterior probability.

A common criticism of Bayesian methods is that the prior is subjective. In practice, the prior should be based on the best available evidence. If a researcher has no information about the prevalence of a condition, a non-informative prior can be used, such as a uniform distribution. However, the use of a non-informative prior does not eliminate the influence of the prior. It simply assigns equal probability to all possible values of the parameter.

The sensitivity of the posterior to the prior is an important consideration. In the diagnostic test example, the posterior probability of 33.3 percent was based on a prior of 5 percent. If the prior were 1 percent, the posterior would be 8.8 percent. If the prior were 20 percent, the posterior would be 70.4 percent. The posterior is highly sensitive to the prior when the test is imperfect and the condition is rare.

A researcher should always report the prior used in the analysis and should consider how the posterior changes with different priors. This is called a sensitivity analysis. A sensitivity analysis is a standard part of a Bayesian analysis and should be reported in the methods section of a paper. The reporting guidelines for research methods emphasize the importance of transparency in the assumptions and the analysis. The EQUATOR Network provides a collection of reporting guidelines that help researchers report their methods clearly and completely.

## Practical Workflow for Applying Bayes' Theorem

The following workflow provides a step-by-step approach to applying Bayes' theorem in a biological research setting.

### Step 1: Define the Hypothesis and the Evidence

The first step is to clearly define the hypothesis and the evidence. The hypothesis is the statement that the researcher wants to evaluate, such as the presence of a pathogen in a sample. The evidence is the observation that is used to update the probability, such as a positive test result.

### Step 2: Determine the Prior Probability

The prior probability is the probability of the hypothesis before the evidence is observed. The prior should be based on the best available evidence, such as the prevalence of the condition in the population. If no prior data are available, the researcher should state the assumption and justify it.

### Step 3: Determine the Likelihood and the False Positive Rate

The likelihood is the probability of observing the evidence if the hypothesis is true. This is the sensitivity of the test. The false positive rate is the probability of observing the evidence if the hypothesis is false. This is one minus the specificity of the test.

### Step 4: Calculate the Posterior Probability

The posterior probability is calculated using the formula. The researcher should perform the calculation and record the result. The calculation can be done by hand, with a spreadsheet, or with statistical software.

### Step 5: Interpret the Posterior Probability

The posterior probability is the updated probability of the hypothesis after the evidence. The researcher should interpret the posterior in the context of the research question. A posterior probability close to 1 indicates strong support for the hypothesis, while a posterior close to 0 indicates strong support against the hypothesis.

### Step 6: Conduct a Sensitivity Analysis

The researcher should vary the prior probability and the test characteristics to see how the posterior changes. This analysis shows the robustness of the conclusion to the assumptions.

### Step 7: Report the Results

The researcher should report the prior, the likelihood, the false positive rate, and the posterior in the methods and results sections. The reporting should be transparent and should follow the reporting guidelines for the study type.

## Records and Measurements

The application of Bayes' theorem requires accurate records of the test characteristics and the prior probability. The following records should be maintained for each analysis:

- The prior probability and its source
- The sensitivity of the test or classifier
- The specificity of the test or classifier
- The number of positive and negative results
- The calculated posterior probability
- The sensitivity analysis results

These records are essential for reproducibility. A researcher should be able to reproduce the posterior probability from the recorded data. The records also allow other researchers to evaluate the assumptions and the analysis.

The National Institutes of Health Data Management and Sharing Policy requires that research data be managed and shared in a way that is consistent with the principles of reproducibility. The policy emphasizes the importance of documenting the methods and the data so that other researchers can verify the results. The records of a Bayesian analysis are part of the data that should be managed and shared.

## Common Failure Patterns in Bayesian Applications

Several common errors occur when researchers apply Bayes' theorem in biological research.

### Ignoring the Prior

The most common error is to ignore the prior probability and interpret the test result as if it were the posterior. This error leads to an overestimation of the probability of the hypothesis when the condition is rare. The correct interpretation is the posterior probability, which accounts for the prior.

### Using the Wrong Prior

The prior probability must be appropriate for the population being studied. A prior based on a different population can produce a misleading posterior. For example, the prevalence of a disease in a referral population is higher than the prevalence in the general population. Using the referral population prevalence as the prior for a general population sample would overestimate the posterior.

### Confusing Sensitivity and Specificity

Sensitivity is the probability of a positive result when the condition is present. Specificity is the probability of a negative result when the condition is absent. Confusing these two values leads to an incorrect posterior. The likelihood is the sensitivity, and the false positive rate is one minus the specificity.

### Ignoring the False Positive Rate

The false positive rate is the probability of a positive result when the condition is absent. This value is often small, but it has a large effect on the posterior when the condition is rare. Ignoring the false positive rate leads to an overestimation of the posterior.

### Not Conducting a Sensitivity Analysis

A posterior probability is only as good as the assumptions. Without a sensitivity analysis, the researcher does not know how the posterior changes with the prior. The sensitivity analysis is a necessary part of the analysis.

## Limitations of Bayesian Inference in Biology

Bayesian inference is a powerful tool, but it has limitations. The posterior probability is only as good as the prior and the likelihood. If the prior is wrong, the posterior is wrong. If the sensitivity and specificity are not known accurately, the posterior is uncertain.

The posterior probability is a probability, not a certainty. A posterior of 0.95 does not mean the hypothesis is true. It means that, under the assumptions of the model, there is a 95 percent chance that the hypothesis is true. The remaining 5 percent is the chance of error.

The Bayesian framework does not eliminate the need for judgment. The choice of the prior is a judgment call, and the interpretation of the posterior is a judgment call. The framework makes the judgment explicit, but it does not remove it.

## Professional Escalation Criteria

A researcher should seek additional guidance or escalate the analysis when the following conditions are met:

- The posterior probability is highly sensitive to the prior, and the prior is uncertain
- The sensitivity or specificity of the test is not known with confidence
- The posterior probability is close to a decision threshold, and the decision is consequential
- The analysis is part of a regulatory submission or a clinical decision

In these cases, the researcher should consult a biostatistician or a methodologist with expertise in Bayesian analysis. The biostatistician can help with the choice of the prior, the sensitivity analysis, and the interpretation of the posterior.

## Reporting and Reproducibility

The reporting of a Bayesian analysis should be transparent and complete. The methods section should include the prior, the likelihood, the false positive rate, and the posterior. The results should include the sensitivity analysis. The data and the code used for the analysis should be shared in accordance with the data management and sharing policy.

The Committee on Publication Ethics core practices emphasize the importance of transparency and the integrity of the research record. The researcher should report the analysis honestly and should not select the prior or the analysis that produces the desired result. The researcher should report the analysis as it was conducted, and should disclose any assumptions.

The ORCID for Researchers provides a system for researchers to maintain a record of their work. The researcher should link the analysis and the data to their ORCID record to ensure that the work is attributed correctly and is available for review.

## Decision Thresholds and Sequential Testing in Field Biology

The worked examples in the preceding sections show how Bayes' theorem updates a single probability after one piece of evidence. In practice, biologists rarely make decisions from a single test result. Field surveys, diagnostic workflows, and monitoring programs generate multiple observations over time, and each observation should update the probability of the hypothesis. The challenge is knowing when to stop collecting data and act on the accumulated evidence. This section provides a practical decision framework for setting thresholds, planning sequential sampling, and recording the evidence trail that supports a defensible biological conclusion.

### Defining the Decision Threshold Before Data Collection

A decision threshold is the posterior probability at which you commit to a conclusion or action. In conservation biology, the threshold might be the probability of species presence that justifies an eradication program. In diagnostic pathology, the threshold might be the probability of infection that triggers treatment or quarantine. In gene expression studies, the threshold might be the probability of a disease classification that determines the next experimental step.

The threshold should be set before data collection begins. Setting the threshold after observing results invites bias, because the researcher may unconsciously choose a threshold that supports the desired conclusion. The Committee on Publication Ethics core practices emphasize the integrity of the research record, and a pre-specified threshold is part of that integrity. The threshold should be recorded in the study protocol or analysis plan, along with the justification for the chosen value.

The choice of threshold depends on the consequences of being wrong. If a false positive leads to an expensive or harmful action, such as culling a population or treating a healthy patient, the threshold should be high. If a false negative leads to a missed opportunity, such as failing to detect an invasive species before it spreads, the threshold may be lower. The threshold is a management decision, not a statistical one. The Bayesian calculation provides the probability, and the researcher or manager decides what probability is sufficient for action.

A common threshold in ecological detection is 95 percent posterior probability of presence before declaring a species established at a site. A common threshold in diagnostic medicine is 90 percent posterior probability of disease before initiating treatment. These values are not universal. They are examples of thresholds that reflect the consequences of error in those specific contexts. The threshold should be justified in the methods section of any report.

### Sequential Sampling Plans

Sequential sampling is a method of collecting data in stages and updating the posterior probability after each stage. The process continues until the posterior crosses a pre-specified threshold or until a maximum number of samples is reached. This approach is more efficient than fixed sampling because it stops early when the evidence is strong and continues when the evidence is ambiguous.

The prior for the first sample is the initial estimate of the probability. After the first sample, the posterior becomes the prior for the second sample. This process repeats for each sample. The key advantage is that the researcher does not need to decide the sample size in advance. The sample size is determined by the data and the threshold.

Consider a field survey for an invasive plant species. The prior probability of occupancy at a randomly selected site is 5 percent, based on the known distribution of the species. The detection method has a sensitivity of 80 percent and a specificity of 97 percent. The management threshold for declaring the species present is 95 percent posterior probability.

The first survey at a site returns a positive detection. The true positive rate is 0.05 multiplied by 0.80, which equals 0.04. The false positive rate is 0.95 multiplied by 0.03, which equals 0.0285. The total probability of a positive result is 0.0685. The posterior probability is 0.04 divided by 0.0685, which equals 0.584. This is below the 90 percent threshold, so a second sample is needed.

The second sample is also positive. The prior is now 0.584. The true positive rate is 0.584 multiplied by 0.80, which equals 0.467. The false positive rate is 0.416 multiplied by 0.03, which equals 0.0125. The total probability of a positive result is 0.4795. The posterior probability is 0.467 divided by 0.4795, which equals 0.974. This exceeds the 90 percent threshold, so the site is declared occupied by the species.

The sequential process required only two positive samples to reach the threshold. If the second sample had been negative, the posterior would have decreased, and the researcher would have needed additional samples or would have concluded that the species is absent. The sequential approach avoids the cost of collecting a fixed number of samples when the evidence is already strong.

### A Practical Decision Framework for Bayesian Thresholds

The following framework provides a structured approach to applying Bayesian thresholds in biological research and management. The framework is designed for field biologists, laboratory professionals, and researchers who need to make decisions from imperfect tests.

#### Step 1: Define the Hypothesis and the Decision

State the hypothesis clearly. The hypothesis is the statement being evaluated, such as the presence of a pathogen, the occupancy of a site by a species, or the classification of a sample. Define the decision that will be made based on the posterior probability. The decision should be specific, such as "initiate treatment" or "declare the species present."

#### Step 2: Set the Posterior Threshold

Choose the posterior probability that justifies the decision. The threshold should be based on the consequences of a false positive and a false negative. Record the threshold and the rationale in the analysis plan.

#### Step 3: Determine the Test Characteristics

Identify the sensitivity and specificity of the detection method. These values should come from validation studies, not from assumptions. If the test characteristics are uncertain, the analysis should include a range of values to show how the posterior changes.

#### Step 4: Collect the First Sample and Update the Posterior

Apply the Bayesian calculation to the first sample. The prior is the initial estimate, and the likelihood is the sensitivity and specificity of the test. Record the posterior.

#### Step 5: Compare the Posterior to the Threshold

If the posterior exceeds the threshold, stop and make the decision. If the posterior is below the threshold, collect another sample. If the posterior is below the threshold and the maximum number of samples has been reached, make the decision based on the final posterior and the consequences of error.

#### Step 6: Record the Evidence Trail

Record the prior, the test characteristics, each sample result, and the posterior after each sample. This record is essential for reproducibility and for defending the decision to stakeholders.

#### Step 7: Conduct a Sensitivity Analysis

Vary the prior and the test characteristics to see how the posterior changes. This analysis shows whether the decision is robust to the assumptions. If the decision changes with a reasonable change in the prior, the evidence is not strong enough for a confident decision.

### Records and Measurements for Sequential Bayesian Analysis

The records for a sequential Bayesian analysis should include the following items for each site or sample:

- The prior probability and its source
- The sensitivity and specificity of the detection method
- The posterior threshold and the rationale
- The result of each sample, positive or negative
- The posterior probability after each sample
- The final decision and the date of the decision
- The sensitivity analysis results

These records are essential for reproducibility. The National Institutes of Health Data Management and Sharing Policy requires that research data be managed and shared in a way that is consistent with the principles of reproducibility. The records of a sequential Bayesian analysis are part of the data that should be managed and shared. The records should be stored in a format that allows other researchers to verify the calculations and the decision.

The records also support the integrity of the research record. The Committee on Publication Ethics core practices emphasize the importance of transparency and the integrity of the research record. The records of the sequential analysis should be available for review by co-authors, reviewers, and stakeholders.

### Common Failure Patterns in Sequential Bayesian Decisions

Several common errors occur when researchers apply Bayesian thresholds and sequential sampling.

#### Setting the Threshold After the Data Are Collected

This is the most common error. The threshold is chosen after the posterior is calculated, which allows the researcher to select a threshold that supports the desired conclusion. The threshold should be set before data collection and recorded in the analysis plan.

#### Using the Wrong Prior for the Population

The prior must be appropriate for the population being studied. A prior based on a different population can produce a misleading posterior. For example, the prevalence of a disease in a referral population is higher than the prevalence in the general population. Using the referral population prevalence as the prior for a general population sample would overestimate the posterior.

#### Ignoring the False Positive Rate

The false positive rate is the probability of a positive result when the condition is absent. This value is often small, but it has a large effect on the posterior when the condition is rare. Ignoring the false positive rate leads to an overestimation of the posterior.

#### Confusing Sensitivity and Specificity

Sensitivity is the probability of a positive result when the condition is present. Specificity is the probability of a negative result when the condition is absent. Confusing these two values leads to an incorrect posterior. The likelihood is the sensitivity, and the false positive rate is one minus the specificity.

#### Not Conducting a Sensitivity Analysis

A posterior probability is only as good as the assumptions. Without a sensitivity analysis, the researcher does not know how the posterior changes with the prior. The sensitivity analysis is a necessary part of the analysis.

#### Stopping Too Early or Too Late

Stopping too early means the posterior is below the threshold and the decision is uncertain. Stopping too late means the posterior is above the threshold but the researcher continues to collect data. The sequential framework should include a maximum number of samples to prevent indefinite sampling.

### Troubleshooting a Sequential Bayesian Analysis

When a sequential Bayesian analysis produces an unexpected result, the following troubleshooting steps can help identify the cause.

#### Check the Prior

The prior is the starting point of the analysis. If the posterior is unexpectedly low after a positive result, the prior may be too low. If the posterior is unexpectedly high after a negative result, the prior may be too high. The prior should be based on the best available evidence and should be stated in the analysis plan.

#### Check the Test Characteristics

The sensitivity and specificity of the test should be based on validation studies. If the test characteristics are not known with confidence, the analysis should include a range of values. The posterior is sensitive to the test characteristics, especially when the condition is rare.

#### Check the Calculation

The Bayesian calculation is straightforward, but errors can occur. The true positive rate is the prior multiplied by the sensitivity. The false positive rate is one minus the prior multiplied by one minus the specificity. The total probability of a positive result is the sum of the true positive rate and the false positive rate. The posterior is the true positive rate divided by the total probability of a positive result.

#### Check the Threshold

The threshold should be set before data collection. If the threshold is too high, the analysis may require many samples. If the threshold is too low, the decision may be made with insufficient evidence. The threshold should reflect the consequences of error.

### Professional Escalation Criteria for Sequential Bayesian Decisions

A researcher should seek additional guidance or escalate the analysis when the following conditions are met:

- The posterior probability is highly sensitive to the prior, and the prior is uncertain
- The sensitivity or specificity of the test is not known with confidence
- The posterior probability is close to a decision threshold, and the decision is consequential
- The analysis is part of a regulatory submission or a clinical decision
- The sequential sampling reaches the maximum number of samples without reaching the threshold

In these cases, the researcher should consult a biostatistician or a methodologist with expertise in Bayesian analysis. The biostatistician can help with the choice of the prior, the sensitivity analysis, and the interpretation of the posterior. The biostatistician can also help with the design of the sequential sampling plan and the selection of the threshold.

### Reporting the Sequential Bayesian Analysis

The reporting of a sequential Bayesian analysis should be transparent and complete. The methods section should include the prior, the sensitivity and specificity, the threshold, and the maximum number of samples. The results should include the posterior after each sample and the final decision. The sensitivity analysis should be reported in the results or the supplementary materials.

The EQUATOR Network provides a collection of reporting guidelines that help researchers report their methods clearly and completely. The researcher should select the appropriate reporting guideline for the study type and follow it. The reporting guideline will specify the information that should be included in the methods and results sections.

The data and the code used for the analysis should be shared in accordance with the data management and sharing policy. The National Institutes of Health Data Management and Sharing Policy requires that research data be managed and shared in a way that is consistent with the principles of reproducibility. The data should include the prior, the test characteristics, the sample results, and the posterior after each sample.

The ORCID for Researchers provides a system for researchers to maintain a record of their work. The researcher should link the analysis and the data to their ORCID record to ensure that the work is attributed correctly and is available for review.

### Comparison of Fixed and Sequential Sampling

Fixed sampling collects a pre-determined number of samples and then calculates the posterior. Sequential sampling collects samples until the posterior reaches a threshold or a maximum number is reached. The choice between the two methods depends on the cost of sampling and the cost of the decision.

Fixed sampling is simpler and easier to plan. The sample size is determined in advance, and the analysis is straightforward. The disadvantage is that the sample size may be larger than necessary when the evidence is strong, or smaller than necessary when the evidence is weak.

Sequential sampling is more efficient because it stops when the evidence is strong. The disadvantage is that the sample size is not known in advance, and the analysis is more complex. The researcher must be prepared to collect the maximum number of samples if the evidence is weak.

The choice between the two methods should be made before data collection and should be based on the cost of sampling and the cost of the decision. If sampling is expensive, sequential sampling may be more efficient. If the decision is consequential, the threshold should be high and the maximum number of samples should be sufficient to reach the threshold.

### The Role of the Prior in Sequential Decisions

The prior plays a critical role in sequential decisions. A prior that is too low will require more samples to reach the threshold. A prior that is too high will require fewer samples but may lead to a false positive decision. The prior should be based on the best available evidence and should be stated in the analysis plan.

The sensitivity analysis should include a range of priors. The researcher should report the posterior for each prior and the decision that would be made for each prior. This analysis shows the robustness of the decision to the prior.

The prior can also be updated as new data are collected. This is the essence of Bayesian inference. The posterior after each sample becomes the prior for the next sample. This process is the foundation of sequential sampling and is the reason the method is so valuable in biological research.

### The Decision Threshold and the Cost of Error

The decision threshold should reflect the cost of error. If a false positive leads to an expensive or harmful action, the threshold should be high. If a false negative leads to a missed opportunity, the threshold should be lower. The threshold is a management decision, not a statistical one.

The researcher should document the rationale for the threshold in the analysis plan. The rationale should include the consequences of a false positive and a false negative. The threshold should be reviewed by the research team and, if necessary, by a biostatistician.

The threshold should also be reported in the methods section of any report. The reader should be able to understand why the threshold was chosen and how it affects the decision. The threshold is part of the research record and should be transparent.

### The Maximum Number of Samples

The maximum number of samples is the upper limit of the sequential sampling plan. The maximum should be set before data collection and should be based on the cost of sampling and the need for a decision. If the maximum is reached without reaching the threshold, the researcher should make the decision based on the final posterior and the consequences of error.

The maximum number of samples should be reported in the methods section. The reader should be able to understand the sampling plan and the decision rule. The maximum number of samples is part of the research record and should be transparent.

### The Decision Rule

The decision rule is the combination of the threshold and the maximum number of samples. The decision rule should be stated in the analysis plan before data collection. The decision rule should be specific enough to be applied consistently by any researcher.

The decision rule should include the following elements:

- The prior probability
- The sensitivity and specificity of the test
- The posterior threshold
- The maximum number of samples
- The decision to be made when the threshold is reached
- The decision to be made when the maximum is reached without reaching the threshold

The decision rule should be recorded in the analysis plan and reported in the methods section. The decision rule is part of the research record and should be transparent.

### The Role of the Biostatistician

A biostatistician can help with the design of a sequential Bayesian analysis. The biostatistician can help with the choice of the prior, the sensitivity analysis, and the interpretation of the posterior. The biostatistician can also help with the selection of the threshold and the maximum number of samples.

The researcher should consult a biostatistician when the analysis is complex or when the decision is consequential. The biostatistician can provide guidance on the design and the interpretation of the analysis. The biostatistician can also help with the reporting of the analysis.

The biostatistician should be acknowledged in the research report. The acknowledgment should be consistent with the authorship guidelines of the Committee on Publication Ethics. The biostatistician should be included as an author if they make a substantial contribution to the design, analysis, or interpretation of the study.

### The Decision in Practice

The decision framework described in this section is a practical tool for applying Bayes' theorem to biological research. The framework is designed to be used by researchers who need to make decisions based on imperfect tests. The framework is not a substitute for statistical expertise, but it provides a structured approach to the decision.

The framework is based on the principles of Bayesian inference and the reporting guidelines of the EQUATOR Network. The framework is consistent with the data management and sharing policy of the National Institutes of Health and the core practices of the Committee on Publication Ethics.

The framework should be used in conjunction with the other sections of this article. The worked examples in the earlier sections show how to calculate the posterior probability. The decision framework in this section shows how to use the posterior probability to make a decision. The records and measurements section shows how to record the analysis. The common failure patterns section shows how to avoid errors. The professional escalation criteria section shows when to seek additional guidance.

The decision framework is a practical tool for the researcher. It is not a substitute for judgment, but it provides a structured approach to the decision. The framework makes the decision explicit and transparent, which is essential for the integrity of the research record.

## Frequently Asked Questions

### What is the difference between the prior and the posterior probability?

The prior probability is the probability of the hypothesis before the evidence is observed. The posterior probability is the probability of the hypothesis after the evidence is observed. The posterior is calculated from the prior and the likelihood of the evidence.

### How do I choose the prior probability in a biological study?

The prior should be based on the best available evidence, such as the prevalence of a condition in a population or the proportion of a species in a habitat. If no data are available, a non-informative prior can be used, but the assumption should be stated.

### What is the sensitivity and specificity of a test?

Sensitivity is the probability that a test is positive when the condition is present. Specificity is the probability that a test is negative when the condition is absent. These two values are used in the Bayesian calculation.

### Why does a positive test result not always mean the condition is present?

A positive test result can be a false positive. The probability that a positive result is a true positive is the posterior probability, which depends on the prior and the sensitivity and specificity. When the condition is rare, the posterior can be low even with a sensitive and specific test.

### Can Bayes' theorem be used for any type of biological data?

Bayes' theorem can be used for any type of data where a hypothesis and evidence can be defined. It is used in diagnostic testing, species detection, gene expression classification, and many other areas of biology.

### What is a sensitivity analysis in Bayesian inference?

A sensitivity analysis is a process of varying the prior and the likelihood to see how the posterior changes. This analysis shows the robustness of the posterior to the assumptions.

### How do I report a Bayesian analysis in my paper?

The methods should include the prior, the sensitivity and specificity, and the posterior. The results should include the posterior and the sensitivity analysis. The reporting should be transparent and follow the guidelines for the field.

### What should I do if my posterior probability is close to a decision threshold?

If the posterior is close to a decision threshold, the decision is uncertain. The researcher should collect more data or consult a biostatistician to refine the analysis.

## Using the Evidence

| Source | Best use in this topic | Important limitation |
|---|---|---|
| [Research Methods Resources](https://www.ncbi.nlm.nih.gov/books) | official guidance | Check the linked page for current local requirements |
| [EQUATOR Network](https://www.equator-network.org/) | official guidance | Check the linked page for current local requirements |
| [Core Practices](https://publicationethics.org/core-practices) | official guidance | Check the linked page for current local requirements |

## Related Bioinformatics Guides

- [Metabolomics Data Analysis Workflow: From Raw Data to Biological Insight](/knowledge/bioinformatics/metabolomics-data-analysis-workflow-from-raw-data-to-biological-insight)
- [Genomic Data Integration: Combining Multi-Omics for Biological Insights](/knowledge/bioinformatics/genomic-data-integration-combining-multi-omics-for-biological-insights)
- [Metagenomics Data Analysis: From Raw Reads to Biological Insights](/knowledge/bioinformatics/metagenomics-data-analysis-from-raw-reads-to-biological-insights)
- [Proteomics Data Analysis Workflow: From Raw Spectra to Biological Insights](/knowledge/bioinformatics/proteomics-data-analysis-workflow-from-raw-spectra-to-biological-insights)
- [Spatial Omics Data Analysis: From Image Processing to Biological Interpretation](/knowledge/bioinformatics/spatial-omics-data-analysis-from-image-processing-to-biological-interpretation)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)


## References and Further Reading

- [Research Methods Resources](https://www.ncbi.nlm.nih.gov/books). National Library of Medicine.
- [EQUATOR Network](https://www.equator-network.org/). EQUATOR Network.
- [Core Practices](https://publicationethics.org/core-practices). Committee on Publication Ethics.
- [NIH Grants and Funding](https://grants.nih.gov/). National Institutes of Health.
- [ORCID for Researchers](https://info.orcid.org/researchers). ORCID.
- [Data Management and Sharing Policy](https://sharing.nih.gov/data-management-and-sharing-policy). National Institutes of Health.
- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Adjusting batch effects in microarray expression data using empirical Bayes methods.](https://pubmed.ncbi.nlm.nih.gov/16632515). Biostatistics (Oxford, England), 2007.
- [Gene-level alignment of single-cell trajectories.](https://pubmed.ncbi.nlm.nih.gov/39300283). Nature methods, 2025.
- [Exoplanet Biosignatures: Future Directions.](https://pubmed.ncbi.nlm.nih.gov/29938538). Astrobiology, 2018.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.