Logistic Regression in Clinical Research

By Dr. Zubair Khalid, DVM, MS, PhD ·

Logistic Regression in Clinical Research

Key Takeaways

  • Odds ratios (ORs) from logistic regression approximate relative risks (RRs) only when the outcome is rare (typically <10% prevalence); otherwise, ORs overestimate the true effect size, potentially misleading clinical decisions.
  • To ensure accurate interpretation, convert ORs to RRs using the baseline risk formula or report predicted probabilities alongside ORs, providing clinicians with directly interpretable measures of absolute risk.
  • Always report the confidence interval for the OR and the baseline risk in the reference group, as an OR without this context can lead to misjudgments regarding treatment efficacy or study conclusions.
  • The distinction between odds and risk is critical; odds are the ratio of event probability to non-event probability, while risk is the direct probability of an event, and this divergence becomes substantial as outcome prevalence increases beyond 10-20%.
  • For clinical decision-making, predicted probabilities offer the most direct and interpretable measure of absolute risk for specific patient profiles, surpassing the clinical utility of ORs when outcomes are common.

Quick Answer

  • Odds ratios from logistic regression approximate relative risk only when the outcome is rare, typically below 10 percent, and misinterpretation otherwise inflates the apparent effect size.
  • Convert odds ratios to risk ratios using the baseline risk formula, or report predicted probabilities alongside odds ratios to give clinicians an interpretable effect measure.
  • Always report the confidence interval and the baseline risk, because an odds ratio without context can mislead treatment decisions and study conclusions.

Understanding the Odds Ratio Problem in Clinical Research

Logistic regression is a standard tool in clinical research for modeling binary outcomes such as disease presence, treatment response, or mortality. The model produces coefficients that are exponentiated to yield odds ratios, which describe the change in the odds of the outcome associated with a one-unit change in a predictor variable. The central problem is that odds and risk are different quantities, and the distinction is frequently lost in clinical interpretation.

Odds represent the ratio of the probability of an event occurring to the probability of it not occurring. If the probability of an event is 0.2, the odds are 0.2 divided by 0.8, which equals 0.25. Risk, also called probability, is the proportion of individuals who experience the event. An odds ratio compares the odds in two groups, while a risk ratio compares the probabilities directly. When the event is rare, the odds and the probability are numerically close, and the odds ratio approximates the risk ratio. When the event is common, the odds ratio diverges from the risk ratio and overstates the association.

This distinction matters in clinical research because many outcomes are not rare. Hospital readmission, surgical complications, and adverse drug reactions can all occur in more than 10% of patients. In these settings, an odds ratio of 2.0 does not mean the risk is doubled. The actual risk increase depends on the baseline risk. A clinician who interprets the odds ratio as a risk ratio will overestimate the treatment effect and may recommend interventions that are not justified by the data.

The problem is compounded by the way results are reported in the literature. Many studies present odds ratios without the baseline risk or the predicted probabilities. The reader is left with a single number that has no intuitive clinical meaning. The solution is not to abandon logistic regression, which remains a powerful and flexible method, but to report results in a way that supports correct interpretation.

At a Glance

ConceptDefinitionClinical Interpretation
OddsRatio of the probability of an event to the probability of no eventLess intuitive for clinicians, ranges from 0 to infinity
RiskProbability of an event occurring in a defined populationDirectly interpretable as the chance of the outcome
Odds ratioRatio of odds between two groupsApproximates risk ratio only when the outcome is rare
Risk ratioRatio of probabilities between two groupsDirectly interpretable as the relative risk
Baseline riskRisk of the outcome in the reference groupRequired to convert an odds ratio to a risk ratio
Predicted probabilityModel-based estimate of the outcome probability for a given profileProvides patient-specific risk estimates

The table above summarizes the core distinction. The odds ratio is a valid measure of association, but its clinical meaning depends on the baseline risk. Researchers should report the baseline risk and, when possible, the predicted probabilities for clinically relevant profiles.

Core Principles of Logistic Regression

The Logistic Function and the Logit Link

Logistic regression models the log odds of the outcome as a linear combination of predictors. The logit link function transforms the probability, which is bounded between 0 and 1, into a value that can range from negative to positive infinity. The model equation is the log of the odds equals the intercept plus the sum of the product of each coefficient and its predictor.

The coefficients from the model are log odds. Exponentiating a coefficient gives the odds ratio for a one-unit change in the predictor, holding other variables constant. This is the standard output from statistical software and the basis for most reported results.

The model assumes a linear relationship between the log-odds and the continuous predictors. This assumption should be checked because a nonlinear relationship can produce biased coefficients. Common approaches include adding polynomial terms, using splines, or categorizing continuous variables, though categorization loses information and is generally discouraged.

The Outcome Variable

The outcome in logistic regression is binary, meaning it has two categories. The outcome must be coded consistently, and the reference category must be clearly defined. For example, if the outcome is mortality, the model can be coded as 1 for death and 0 for survival. The odds ratio then describes the odds of death for the exposed group relative to the reference group.

The outcome should be well-defined and measured without error. Misclassification of the outcome can bias the odds ratio toward the null, meaning the true association is underestimated. Researchers should validate the outcome definition and consider sensitivity analyses when misclassification is possible.

The Predictors

Predictors can be continuous, binary, or categorical. Continuous predictors are often centered or standardized to improve interpretability. Categorical predictors with more than two levels require dummy coding, and the reference category must be specified. The choice of reference category affects the interpretation of the coefficients but does not change the overall model fit.

Interaction terms can be included to assess whether the effect of one predictor depends on the level of another. Interactions are important in clinical research because treatment effects often vary by patient characteristics. However, interactions increase model complexity and require adequate sample size to estimate reliably.

Odds Ratio versus Risk Ratio

Why the Distinction Matters

The odds ratio is a measure of association that is mathematically convenient but clinically opaque. The risk ratio is more intuitive because it directly compares probabilities. The two measures diverge as the outcome becomes more common.

Consider a study where the risk of an event in the unexposed group is 0.5 and the risk in the exposed group is 0.75. The risk ratio is 0.75 divided by 0.5, which equals 1.5. The odds in the unexposed group are 0.5 divided by 0.5, which equals 1.0. The odds in the exposed group are 0.75 divided by 0.25, which equals 3.0. The odds ratio is 3.0 divided by 1.0, which equals 3.0. The odds ratio is twice the risk ratio, and a clinician who interprets the odds ratio as a risk ratio would conclude that the risk is tripled when it is actually increased by 50%.

This example illustrates the magnitude of the problem. The odds ratio is always further from 1 than the risk ratio when the outcome is common. The direction of the bias is always away from the null, meaning the odds ratio overstates the association.

When the Odds Ratio Approximates the Risk Ratio

When the outcome is rare, the odds and the probability are numerically close. If the probability of an event is 0.01, the odds are 0.01 divided by 0.99, which equals 0.0101. The odds are nearly identical to the probability. In this setting, the odds ratio is a good approximation of the risk ratio.

The threshold for "rare" is not fixed, but a common guideline is an outcome prevalence below 10%. At a prevalence of 10%, the odds are 0.10 divided by 0.90, which equals 0.111. The odds ratio will be modestly larger than the risk ratio. At a prevalence of 20%, the odds are 0.20 divided by 0.80, which equals 0.25, and the divergence becomes substantial.

Researchers should assess the outcome prevalence in the study population and report whether the odds ratio is a reasonable approximation of the risk ratio. If the outcome is common, the odds ratio should be converted to a risk ratio or the predicted probabilities should be reported.

Converting Odds Ratios to Risk Ratios

The Conversion Formula

The odds ratio can be converted to a risk ratio using the baseline risk in the reference group. The formula is:

Risk ratio = Odds ratio divided by (1 minus the baseline risk plus the baseline risk multiplied by the odds ratio)

This formula requires the baseline risk, which is the probability of the outcome in the reference group. The baseline risk can be estimated from the model intercept or from the observed data.

For example, if the baseline risk is 0.2 and the odds ratio is 2.0, the risk ratio is 2.0 divided by (1 minus 0.2 plus 0.2 multiplied by 2.0), which equals 2.0 divided by (0.8 plus 0.4), which equals 2.0 divided by 1.2, which equals 1.67. The risk ratio is 1.67, not 2.0. The odds ratio overstates the relative risk by about 20%.

The conversion is straightforward and can be done with a calculator or statistical software. The key requirement is the baseline risk, which should be reported in the study results.

Predicted Probabilities

An alternative approach is to report the predicted probabilities for clinically relevant groups. The logistic regression model can generate a predicted probability for any combination of predictor values. These probabilities are directly interpretable as the risk of the outcome for a patient with those characteristics.

For example, the model can predict the probability of post-operative infection for a patient with diabetes, a specific age, and a specific surgical procedure. The predicted probability is more informative than an odds ratio because it provides the absolute risk, which is what clinicians and patients need for decision-making.

Predicted probabilities can be presented in a table or a graph. A common approach is to show the predicted probability across the range of a continuous predictor, with the other predictors held at their mean or at clinically relevant values. This presentation is more interpretable than a single odds ratio.

Software Implementation

Most statistical software packages can compute predicted probabilities and risk ratios from a logistic regression model. In R, the margins package can compute the average marginal effects and the risk ratios. In Stata, the margins command provides similar functionality. The software can also compute confidence intervals for the risk ratios using the delta method or bootstrapping.

The conversion should be reported with the confidence interval to convey the uncertainty in the estimate. The confidence interval for the risk ratio is not the same as the confidence interval for the odds ratio, and it should be computed separately.

Reporting Guidelines for Odds Ratios

The STROBE Statement

The STROBE statement provides a checklist for reporting observational studies, including cohort, case-control, and cross-sectional studies. The checklist includes items on the statistical methods, the results, and the interpretation of the findings. The EQUATOR Network maintains a comprehensive collection of reporting guidelines, and the STROBE statement is one of the most widely used.

The STROBE checklist recommends that researchers describe the statistical methods in enough detail to allow replication. This includes the model specification, the handling of missing data, and the sensitivity analyses. The results should be reported with the precision of the estimates, including the confidence intervals.

The EQUATOR Network is the primary resource for identifying the appropriate reporting guideline for a study. Researchers should consult the network before writing the manuscript to ensure that all relevant items are addressed.

Reporting the Baseline Risk

The odds ratio alone is not sufficient for clinical interpretation. The baseline risk in the reference group should be reported so that the reader can convert the odds ratio to a risk ratio if desired. The baseline risk can be reported as the proportion of events in the reference group or as the predicted probability at the mean of the covariates.

The baseline risk is particularly important when the outcome is common. Without the baseline risk, the reader cannot assess the clinical significance of the odds ratio. The odds ratio may be statistically significant but clinically unimportant if the baseline risk is very low.

Reporting the Confidence Interval

The confidence interval conveys the precision of the estimate. A wide confidence interval indicates that the estimate is imprecise, and the true effect may be substantially different from the point estimate. The confidence interval should be reported for the odds ratio and for any converted risk ratio.

The confidence interval is also important for interpreting the clinical significance. A statistically significant odds ratio with a wide confidence interval may not be clinically meaningful. The reader should consider the lower and upper bounds of the confidence interval in the context of the clinical decision.

Reporting the Model Performance

The model performance should be reported to assess the predictive ability of the model. The area under the receiver operating characteristic curve, or the C-statistic, is a common measure of discrimination. The C-statistic ranges from 0.5, which indicates no discrimination, to 1.0, which indicates perfect discrimination.

The calibration of the model should also be assessed. Calibration refers to the agreement between the predicted probabilities and the observed outcomes. A well-calibrated model predicts probabilities that match the observed event rates. Calibration can be assessed with a calibration plot or the Hosmer-Lemeshow test, though the test has limitations and should be interpreted with caution.

Practical Workflow for Logistic Regression

Step 1: Define the Research Question

The research question should be specific and clinically relevant. The outcome should be clearly defined, and the predictors should be selected based on the clinical knowledge and the literature. The research question determines the model specification and the interpretation of the results.

Step 2: Prepare the Data

The data should be checked for completeness and accuracy. Missing data should be handled with an appropriate method, such as multiple imputation, and the method should be reported. Outliers and influential observations should be examined, and the data should be checked for errors.

Step 3: Check the Assumptions

The logistic regression model has several assumptions that should be checked. The outcome should be binary, the observations should be independent, and the log-odds should be linearly related to the continuous predictors. The linearity assumption can be checked with a plot of the log-odds against the predictor or with a statistical test.

Step 4: Fit the Model

The model should be fitted with the selected predictors. The model should be parsimonious, meaning that it includes only the predictors that are necessary for the research question. The model should not be overfitted, which occurs when too many predictors are included for the sample size.

Step 5: Evaluate the Model

The model should be evaluated for discrimination and calibration. The AUC should be reported, and the calibration should be assessed. The model should be validated internally, for example with bootstrapping, to assess the optimism in the performance estimates.

Step 6: Report the Results

The results should be reported according to the appropriate reporting guideline. The odds ratios should be reported with the confidence intervals, and the baseline risk should be reported. The predicted probabilities should be reported for clinically relevant groups.

Options and Tradeoffs in Model Specification

Continuous versus Categorical Predictors

Continuous predictors retain all the information in the data and are generally preferred. Categorizing a continuous predictor, such as age or blood pressure, loses information and can reduce the statistical power. The categorization also introduces arbitrary cut points that can affect the results.

However, continuous predictors may have a nonlinear relationship with the log-odds. In this case, the predictor can be modeled with a spline or a polynomial term. The choice of the functional form should be guided by the data and the clinical context.

Adjustment for Confounders

The logistic regression model can adjust for confounders, which are variables that are associated with both the exposure and the outcome. The adjustment is intended to estimate the effect of the exposure independent of the confounders. The selection of confounders should be based on the causal diagram and the literature, not on the statistical significance.

The adjustment for confounders can change the odds ratio substantially. The unadjusted odds ratio may be confounded, and the adjusted odds ratio is the preferred estimate. The adjusted model should be reported with the list of the covariates included.

Interaction Terms

Interaction terms allow the effect of one predictor to vary by the level of another predictor. For example, the effect of a treatment may differ between men and women. The interaction term is the product of the two predictors, and the model includes the main effects and the interaction.

The interaction should be specified based on the clinical hypothesis, not on the data mining. The interaction should be tested with a likelihood ratio test, and the results should be reported with the stratified estimates.

Sample Size Considerations

The sample size must be adequate for the number of predictors in the model. A common rule of thumb is that there should be at least 10 events per predictor variable, though this rule has been criticized and may be too simplistic. The sample size should be calculated based on the expected effect size and the desired precision.

The sample size affects the stability of the model. A small sample size can produce unstable estimates and wide confidence intervals. The model may also be overfitted, meaning that the model fits the observed data well but does not generalize to new data.

Observations and Measurements

Data Quality Checks

The data quality should be assessed before the analysis. The outcome variable should be checked for the correct coding and for the distribution of the events. The predictors should be checked for the range, the distribution, and the missing values. The data should be checked for duplicates and for inconsistencies.

The data quality checks should be documented in the analysis plan. The checks should be performed before the model is fitted, and the results should be reported in the study.

Model Diagnostics

The model should be examined for influential observations and outliers. The residuals should be examined for patterns that indicate a poor model fit. The model should be checked for multicollinearity, which occurs when the predictors are highly correlated.

The model diagnostics should be reported in the study. The diagnostics can identify problems with the model that are not apparent from the coefficients.

Sensitivity Analyses

The sensitivity analyses should be performed to assess the robustness of the results. The sensitivity analyses can include the analysis with the missing data handled differently, the analysis with the outliers excluded, and the analysis with the different model specifications. The sensitivity analyses should be reported in the study.

The sensitivity analyses are important because the results of the primary analysis may depend on the modeling choices. The sensitivity analyses can identify the choices that have a large impact on the results.

Records and Measurements

The Analysis Plan

The analysis plan should be written before the analysis is performed. The plan should describe the research question, the data, the model, the covariates, and the sensitivity analyses. The plan should be registered in a public registry, if possible, to prevent the selective reporting of the results.

The analysis plan is a record of the decisions that were made before the analysis. The plan can be compared with the final analysis to assess the consistency.

The Statistical Software

The statistical software and the version should be reported in the study. The software can produce different results due to the differences in the algorithms and the default settings. The software should be described in the methods section.

The code should be made available for the replication of the analysis. The code should be well-documented and should be run on the same data to reproduce the results.

The Data

The data should be described in the study, including the source, the collection methods, and the variables. The data should be made available for the replication, subject to the ethical and legal constraints. The data sharing is a standard practice in the clinical research, and the NIH has a data management and sharing policy that requires the data to be shared.

The data sharing policy is described in the NIH data management and sharing policy, which is available at the NIH sharing website. The policy requires the data to be shared in a repository that is appropriate for the data type.

Common Failure Patterns

Misinterpreting the Odds Ratio as a Risk Ratio

The most common failure is the misinterpretation of the odds ratio as a risk ratio. This failure occurs when the outcome is common and the odds ratio is larger than the risk ratio. The misinterpretation leads to an overestimation of the effect and can lead to the wrong clinical decisions.

The misinterpretation can be avoided by reporting the baseline risk and the predicted probabilities. The reader should be aware of the distinction between the odds and the risk.

Overfitting the Model

The overfitting occurs when the model is too complex for the sample size. The overfitted model has a good fit to the observed data but does not generalize to the new data. The overfitting can be detected by the internal validation, such as the bootstrapping.

The overfitting can be prevented by limiting the number of predictors and by using the shrinkage methods. The model should be validated in an external data set, if possible.

Ignoring the Linearity Assumption

The linearity assumption is often ignored. The relationship between the log-odds and the predictor may be nonlinear, and the model may be misspecified. The misspecification can produce the biased coefficients and the incorrect conclusions.

The linearity assumption should be checked with the graphical methods and the statistical tests. The nonlinearity can be modeled with the splines or the polynomial terms.

The Complete Separation

The complete separation occurs when the outcome is perfectly predicted by a predictor or a combination of the predictors. The model cannot be fitted, and the coefficients are infinite. The complete separation is more common in the small samples and the rare outcomes.

The complete separation can be detected by the large coefficients and the wide confidence intervals. The model can be fitted with the penalized methods, such as the Firth's method.

Limitations of Logistic Regression

The Binary Outcome

The logistic regression is limited to the binary outcomes. The outcome cannot have more than two categories. The multinomial logistic regression can be used for the outcomes with more than two categories, but the interpretation is more complex.

The binary outcome is a simplification of the clinical reality. The outcome may be a continuous measure that is dichotomized, and the dichotomization can lose information.

The Independence Assumption

The logistic regression assumes that the observations are independent. The observations may be correlated, such as the repeated measurements or the clustered data. The correlated data can be analyzed with the generalized estimating equations or the mixed-effects models.

The Linearity Assumption

The logistic regression assumes a linear relationship between the log-odds and the predictors. The linearity assumption is often violated, and the model may be misspecified. The misspecification can be addressed with the flexible modeling approaches.

The Causality

The logistic regression can be used to estimate the association between the predictors and the outcome, but it cannot establish the causality. The causality requires the study design and the causal inference methods. The observational studies are subject to the confounding and the bias.

Safety and Regulatory Context

The Reporting Guidelines

The reporting guidelines are the standards for the transparent reporting of the research. The EQUATOR Network is the resource for the reporting guidelines, and the researchers should consult the guidelines before the writing of the study. The guidelines are the STROBE for the observational studies, the CONSORT for the randomized trials, and the TRIPOD for the prediction models.

The reporting guidelines are the requirement for the publication in many journals. The guidelines are intended to improve the quality of the research and to reduce the waste in the research.

The Data Sharing

The data sharing is the requirement for the NIH-funded research. The NIH data management and sharing policy requires the data to be shared in a repository. The policy is described in the NIH sharing website.

The data sharing is the ethical responsibility of the researchers. The data sharing allows the replication of the analysis and the secondary analysis of the data.

The Publication Ethics

The publication ethics are the standards for the conduct of the research and the publication. The Committee on Publication Ethics (COPE) provides the core practices for the publication ethics. The practices include the authorship, the peer review, the data, the conflicts of interest, and the misconduct.

The researchers should be aware of the publication ethics and should follow the standards. The misconduct includes the fabrication, the falsification, and the plagiarism.

Professional Escalation Criteria

When to Consult a Biostatistician

The logistic regression analysis should be performed with the guidance of a biostatistician, especially when the model is complex. The biostatistician can help with the model specification, the model diagnostics, and the interpretation of the results.

The biostatistician should be consulted when the data are complex, such as the missing data, the correlated data, or the high-dimensional data. The biostatistician can also help with the sample size calculation and the study design.

When to Seek a Second Opinion

The results of the analysis should be reviewed by a second person, such as a colleague or a biostatistician. The second review can identify the errors in the analysis and the interpretation.

The second review is especially important when the results are surprising or when the results have important clinical implications. The second review can also be a part of the peer review process.

When to Report the Concerns

The concerns about the analysis should be reported to the appropriate person, such as the principal investigator or the institutional review board. The concerns may include the data errors, the analysis errors, or the ethical issues.

The concerns should be reported in a timely manner. The reporting of the concerns is the responsibility of the researcher.

A Practical Decision Framework for Choosing Between Odds Ratios and Risk Ratios

The Clinical Communication Problem

The choice between reporting odds ratios and risk ratios is also a statistical preference. It is a clinical communication decision that affects how practitioners interpret study findings and apply them to patient care. A researcher who reports an odds ratio of 3.0 for a common outcome may unintentionally lead clinicians to believe the risk is tripled when the actual risk increase is far smaller. The decision framework below helps researchers determine when to report which measure and how to present both when the clinical context demands it.

Step 1: Determine the Outcome Prevalence in the Reference Group

The first decision point is the observed prevalence of the outcome in the reference group. This is the proportion of participants in the unexposed or baseline category who experienced the event. The prevalence can be obtained directly from the study data or estimated from the model intercept.

If the prevalence is below 10 percent, the odds ratio is a reasonable approximation of the risk ratio, and the difference between the two measures is unlikely to change the clinical interpretation. If the prevalence is between 10 and 20 percent, the odds ratio will be modestly larger than the risk ratio, and the researcher should consider whether the difference matters for the intended audience. If the prevalence exceeds 20 percent, the odds ratio will be substantially larger than the risk ratio, and reporting only the odds ratio is likely to mislead.

The prevalence threshold is not a hard rule. The acceptable difference between the odds ratio and the risk ratio depends on the clinical context. A 10 percent relative difference may be acceptable for a screening test but not for a treatment decision with serious side effects.

Step 2: Assess the Clinical Purpose of the Study

The intended use of the study results determines the appropriate measure. If the study is designed to estimate the strength of an association for etiologic research, the odds ratio is a valid and useful measure. The odds ratio has desirable statistical properties, including its relationship to the coefficients in the logistic model and its use in case-control studies where the outcome prevalence is fixed by design.

If the study is designed to inform clinical decisions, such as whether to prescribe a treatment or recommend a screening test, the risk ratio and the absolute risk are more interpretable. Clinicians need to know the probability of the outcome with and without the intervention, not the ratio of odds. The predicted probabilities are the most direct way to communicate this information.

Step 3: Consider the Study Design

The study design constrains the choice of measure. In a case-control study, the odds ratio is the natural measure because the sampling is conditional on the outcome. The risk ratio cannot be estimated directly from a case-control study without additional information about the population at risk. In a cohort study or a randomized trial, the risk ratio can be estimated directly from the observed proportions.

The odds ratio from a case-control study can be converted to a risk ratio if the prevalence of the outcome in the population is known. This conversion is rarely possible because the prevalence is often unknown. The researcher should report the odds ratio and note that it approximates the risk ratio only if the outcome is rare.

Step 4: Select the Reporting Strategy

The reporting strategy should be selected based on the first three steps. The options are:

  • Report the odds ratio alone when the outcome is rare and the study is etiologic.
  • Report the odds ratio with the baseline risk when the outcome is common and the study is etiologic.
  • Report the risk ratio or the predicted probabilities when the study is intended to inform clinical decisions.
  • Report both the odds ratio and the risk ratio when the audience includes both researchers and clinicians.

The reporting strategy should be specified in the analysis plan before the results are known. This prevents the researcher from choosing the measure that makes the results look more favorable.

Step 5: Document the Decision

The decision to report a particular measure should be documented in the methods section of the manuscript. The documentation should include the prevalence of the outcome, the rationale for the choice of measure, and the conversion formula if a risk ratio is reported. This documentation allows the reader to assess whether the choice was appropriate.

The documentation also supports the transparency of the research. The EQUATOR Network provides reporting guidelines that include items on the statistical methods and the presentation of the results. The STROBE statement for observational studies and the CONSORT statement for randomized trials both require the statistical methods to be described in enough detail to allow replication.

A Worked Example of the Decision Framework

Consider a study of post-operative infection after a surgical procedure. The outcome is the occurrence of a surgical site infection within 30 days. The observed prevalence of infection in the reference group is 15 percent. The logistic regression model produces an odds ratio of 2.5 for the exposure of interest.

The prevalence is above 10 percent, so the odds ratio is not a good approximation of the risk ratio. The study is designed to inform clinical decisions about whether to use a particular surgical technique. The risk ratio is the more appropriate measure.

The risk ratio is calculated as 2.5 divided by (1 minus 0.15 plus 0.15 multiplied by 2.5), which equals 2.5 divided by (0.85 plus 0.375), which equals 2.5 divided by 1.225, which equals 2.04. The risk ratio is 2.04, not 2.5. The odds ratio overstates the relative risk by about 23 percent.

The researcher should report the risk ratio of 2.04 with its confidence interval and the baseline risk of 5 percent. The researcher should also report the predicted probabilities for the exposed and unexposed groups, which are 10.2 percent and 5 percent respectively. These probabilities are directly interpretable by the clinician.

A Record System for the Reporting Decision

The decision framework should be recorded in the study documentation. The record should include the following items:

  • The prevalence of the outcome in the reference group.
  • The intended use of the study results.
  • The study design.
  • The chosen reporting strategy.
  • The rationale for the choice.
  • The date of the decision.

The record can be a simple table in the analysis plan or a separate document. The record should be updated if the decision changes during the analysis. The record is a part of the audit trail that supports the reproducibility of the research.

The record also supports the peer review process. The reviewer can check whether the reporting decision was made before the results were known and whether the rationale is sound. The record is a part of the transparent reporting that is expected in clinical research.

Common Failure Patterns in the Reporting Decision

Choosing the Measure After Seeing the Results

The most common failure is to choose the reporting measure after the results are known. If the odds ratio is larger than the risk ratio, the researcher may be tempted to report the odds ratio to make the effect look larger. This is a form of selective reporting and is a violation of the publication ethics. The decision should be made before the analysis and documented in the analysis plan.

Ignoring the Baseline Risk

The second common failure is to report the odds ratio without the baseline risk. The reader cannot convert the odds ratio to a risk ratio without the baseline risk. The baseline risk should be reported in the results section, either as the proportion of events in the reference group or as the predicted probability at the mean of the covariates.

Reporting the Odds Ratio as a Risk Ratio

The third common failure is to report the odds ratio as a risk ratio in the abstract or the discussion. This is a misinterpretation that can mislead the reader. The abstract should report the measure that was actually estimated, and the discussion should interpret the measure correctly.

The Role of the Statistical Software

The statistical software can compute the risk ratio and the predicted probabilities from the logistic regression model. The researcher should use the software to compute the risk ratio and the confidence interval, not to report the odds ratio alone. The software can also produce the predicted probabilities for the clinically relevant groups.

The software output should be checked for the consistency with the decision framework. The researcher should verify that the prevalence of the outcome is correctly estimated and that the conversion formula is applied correctly. The software output should be documented in the study records.

The Professional Escalation

The decision framework should be applied by the researcher with the guidance of a biostatistician. The biostatistician can help with the calculation of the risk ratio and the confidence interval and with the interpretation of the results. The biostatistician should be consulted when the outcome is common and the odds ratio is likely to mislead.

The researcher should also consult the reporting guidelines from the EQUATOR Network before writing the manuscript. The guidelines provide the checklist of the items that should be reported and the standards for the transparent reporting. The guidelines are the STROBE for the observational studies and the CONSORT for the randomized trials.

The researcher should report the concerns about the reporting to the principal investigator or the institutional review board if the decision framework is not followed. The reporting of the concerns is the responsibility of the researcher and is a part of the publication ethics.

Frequently Asked Questions

What is the difference between an odds ratio and a risk ratio?

The odds ratio is the ratio of the odds of an event in two groups, while the risk ratio is the ratio of the probabilities of the event. The odds ratio is the output of the logistic regression, and the risk ratio is the ratio of the probabilities. The odds ratio approximates the risk ratio when the outcome is rare, but the odds ratio overstates the risk when the outcome is common.

When can I interpret an odds ratio as a risk ratio?

The odds ratio can be interpreted as a risk ratio when the outcome is rare, typically below 10%. When the outcome is more common, the odds ratio is larger than the risk ratio, and the interpretation as a risk ratio is incorrect. The baseline risk should be reported to allow the conversion.

How do I convert an odds ratio to a risk ratio?

The odds ratio can be converted to a risk ratio using the baseline risk in the reference group. The formula is the risk ratio equals the odds ratio divided by the quantity one minus the baseline risk plus the baseline risk multiplied by the odds ratio. The conversion requires the baseline risk, which should be reported in the study.

Why is the odds ratio larger than the risk ratio?

The odds ratio is larger than the risk ratio when the outcome is common because the odds are a nonlinear transformation of the probability. The odds are the probability divided by one minus the probability, and the odds are larger than the probability when the probability is greater than 0.5. The odds ratio is the ratio of the odds, and the ratio is larger than the ratio of the probabilities.

What is the baseline risk?

The baseline risk is the probability of the event in the reference group. The baseline risk is the risk in the group that is the reference for the comparison. The baseline risk is required to convert the odds ratio to a risk ratio and to interpret the clinical significance of the odds ratio.

What are the predicted probabilities?

The predicted probabilities are the model-based estimates of the probability of the outcome for a given set of predictor values. The predicted probabilities are directly interpretable as the risk of the outcome for a patient with those characteristics. The predicted probabilities can be reported in a table or a graph.

What is the STROBE statement?

The STROBE statement is a reporting guideline for the observational studies. The guideline provides a checklist of the items that should be reported in the study. The STROBE statement is one of the reporting guidelines that are available from the EQUATOR Network.

What is the role of the NIH data sharing policy?

The NIH data management and sharing policy requires the data from the NIH-funded research to be shared. The policy is described at the NIH sharing website. The data sharing is intended to improve the reproducibility and the transparency of the research.

Using the Evidence

SourceBest use in this topicImportant limitation
Research Methods Resourcesofficial guidanceCheck the linked page for current local requirements
EQUATOR Networkofficial guidanceCheck the linked page for current local requirements
Core Practicesofficial guidanceCheck the linked page for current local requirements

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.