Humane Endpoint Determination for Rodent Models of Sepsis

By Dr. Zubair Khalid, DVM, MS, PhD ·

Humane Endpoint Determination for Rodent Models of Sepsis

Key Takeaways

  • Sepsis models necessitate rigorous humane endpoint determination due to rapid clinical decline and overlapping signs with experimental insult, with a significant gap identified in published euthanasia criteria (only 9% in highly cited studies).
  • Objective clinical parameters, including a minimum 20-25% body weight loss from baseline or hypothermia below 34.0°C (or a >4.0°C drop), are critical for consistent endpoint assessment.
  • A composite scoring system should be complemented by single-parameter severe triggers (e.g., severe respiratory distress, unresponsiveness) to prevent masked decline and ensure immediate intervention.
  • Baseline assessments (weight, temperature, behavior) for at least 3 days prior to induction are crucial to account for individual variation, and scoring frequency should be every 4-6 hours during the acute phase.
  • Score sheets must differentiate pain-induced behaviors from sepsis signs, and analgesia administration must be documented to inform interpretation of clinical parameters.
  • Survival follow-up duration should be calibrated to the specific model's expected clinical time course of the infectious agent, rather than an arbitrary fixed period.

This article provides a structured framework for defining, implementing, and refining humane endpoints in rodent models of sepsis. It is written for veterinary researchers, laboratory animal veterinarians, and animal care staff who design or oversee sepsis studies and must balance scientific validity with welfare obligations. The procedural focus is on objective clinical assessment, scoring instrument design, and decision thresholds that can be applied consistently across investigators and institutions.

Sepsis models present particular challenges for endpoint determination because the clinical trajectory is rapid, the signs overlap with expected sequelae of the experimental insult, and the scientific objective often requires observing a defined period of morbidity. The Minimum Quality Threshold in Pre-Clinical Sepsis Studies (MQTiPSS) consensus, developed by 31 experts across 13 countries, identified that euthanasia criteria were defined in only 9% of the 260 most highly cited sepsis modeling articles published between 2003 and 2012, despite mice being used in 79% of those studies MQTiPSS study design and humane modeling recommendations. This gap between model use and welfare planning is the central problem this article addresses.

At a Glance

ParameterRecommended ApproachSource or Rationale
Baseline assessmentRecord weight, temperature, appearance, and behavior for at least 3 days before inductionIndividual variation exceeds strain norms
Scoring frequencyEvery 4 to 6 hours during acute phase, then at least twice dailySepsis trajectories change rapidly
Core parametersWeight, temperature, piloerection, posture, activity, respiration, hydration, ocular dischargeMQTiPSS humane modeling working group
Weight loss threshold20% to 25% from baseline, adjusted for model durationSystematic review of mouse endpoint definitions
Temperature thresholdHypothermia below 34.0°C or a drop exceeding 4.0°C from baselineMachine learning endpoint analysis
Single parameter triggersAny one parameter reaching a predefined severe score mandates euthanasiaPrevents composite score masking decline
Analgesia interactionScore sheets must distinguish pain from sepsis signsAnalgesia in clinically relevant rodent sepsis models
Survival follow-upDuration should reflect the clinical time course of the infectious agentMQTiPSS consensus recommendation

Why Sepsis Endpoints Differ from Other Disease Models

Sepsis induces a systemic inflammatory response that affects every organ system, producing clinical signs that are neither pathognomonic nor static. A rodent with peritonitis may appear nearly normal at one observation and moribund within hours. This rapid decompensation distinguishes sepsis from tumor models or chronic disease models, where weight loss and behavioral change accumulate over days or weeks and allow more leisurely endpoint decisions.

The scientific objective also complicates endpoint setting. Many sepsis studies measure survival as the primary outcome, which creates an inherent tension: the experiment requires observing death or near-death, while welfare standards require intervention before that point. The MQTiPSS consensus addresses this by recommending that survival follow-up reflect the clinical time course of the infectious agent used, instead of a fixed arbitrary duration MQTiPSS executive summary. In practice, this means a cecal ligation and puncture model with high-grade puncture may warrant a 24 to 48 hour observation window, while a low-grade model designed to study recovery may require 7 to 14 days. The endpoint criteria must be calibrated to the expected trajectory of the specific model.

Physiological Basis of Endpoint Parameters

Body Weight

Weight loss in sepsis reflects a combination of reduced food and water intake, catabolic metabolism, and third-space fluid losses. It is the most widely used endpoint parameter in mouse models, but the systematic review of endpoint definitions found large intra- and inter-model variance in how weight loss thresholds are applied machine learning-based humane endpoint definition. Some studies use a fixed percentage from baseline, others use a comparison to sham-operated controls, and still others combine weight with a clinical score. The choice matters because a 20% threshold applied to a model with a 5 day expected course is more permissive than the same threshold applied to a 48 hour acute model.

Temperature

Body temperature is a sensitive and early indicator of sepsis severity in rodents. Mice and rats are homeotherms with a narrow thermoneutral zone, and their thermoregulatory responses to systemic inflammation are rapid. Hypothermia in sepsis models correlates with poor outcome and is often the first objective parameter to deviate from baseline. Temperature measurement method matters: infrared thermometry of the tail or paw reflects peripheral perfusion instead of core temperature, while rectal probes provide core values but require handling that may itself induce stress hyperthermia. The machine learning analysis of sepsis model data identified temperature as one of the most informative parameters for endpoint prediction, particularly when combined with weight and sickness score machine learning-based humane endpoint definition.

Clinical Appearance and Behavior

Piloerection, ocular discharge, hunched posture, and reduced spontaneous activity form the core of most sepsis sickness scores. These signs reflect autonomic dysfunction, dehydration, and the sickness behavior syndrome mediated by proinflammatory cytokines acting on the central nervous system. The challenge is that these signs are also produced by pain, and the distinction matters for both welfare and scientific interpretation. The review of analgesia in rodent sepsis models notes that score sheets should be used to adapt analgesia or to terminate experiments using humane endpoints, while acknowledging that research is needed to differentiate behavioral changes caused by sepsis from those caused by pain or analgesia analgesia in rodent sepsis models. This overlap means that a well-designed score sheet must include parameters that are relatively specific to sepsis severity, such as respiratory rate and effort, instead of relying solely on general appearance.

The MQTiPSS Framework for Humane Modeling

The MQTiPSS consensus, published across multiple companion papers, provides the most current international expert guidance on humane endpoints in sepsis models. The humane modeling working group produced recommendations at two strength levels: "recommendation" for points with strong consensus and clear evidence, and "consideration" for points where the evidence base was thinner or expert opinion divided MQTiPSS humane modeling and study design. The full consensus statement and its detailed commentaries are published as a coordinated set MQTiPSS international expert consensus.

Key recommendations relevant to endpoint determination include: defining euthanasia criteria before study initiation, using a combination of objective parameters instead of a single sign, ensuring that the person assessing animals is blinded to treatment group, and reporting endpoint criteria in publications. The consensus also emphasizes that the choice of anesthetic, analgesic, and supportive care affects both the clinical signs observed and the validity of the model, so endpoint criteria cannot be separated from the broader experimental protocol.

Score Sheet Design Principles

A score sheet for sepsis models should be model-specific, pilot-tested, and revised if the observed trajectory does not match the predicted one. The parameters selected must be measurable with acceptable inter-observer reliability, sensitive to the expected severity range, and feasible to assess at the required frequency. A common structure assigns each parameter a score from 0 to 3 or 0 to 4, with a cumulative threshold for euthanasia and a separate rule that any single parameter reaching the maximum score triggers immediate euthanasia regardless of the total.

The composite score approach has a known failure mode: an animal can accumulate moderate scores across many parameters and reach the euthanasia threshold while appearing less distressed than an animal with a single severe abnormality. The single-parameter trigger rule addresses this by ensuring that severe hypothermia, marked respiratory distress, or unresponsiveness to handling always results in euthanasia even if the total score is below the threshold. Conversely, the composite score captures animals that deteriorate gradually across multiple systems without any single dramatic sign.

Pilot data are essential for calibrating thresholds. The systematic review of endpoint definitions found that published cut-off values vary widely, and the authors recommend that institutions generate their own baseline data for each model and strain machine learning-based humane endpoint definition. This is particularly important for newer models or genetically modified strains, where published reference values may not apply.

Clinical Assessment Sequence and Decision Points

Endpoint determination in sepsis models is a dynamic process, not a single observation. The assessment sequence should follow a fixed schedule with defined trigger points for escalation. Begin with a baseline assessment at least 24 hours before model induction to establish individual reference values for body weight, temperature, and behavior. The MQTiPSS consensus recommends that survival follow-up reflect the clinical time course of the infectious agent used, which means the monitoring frequency must match the expected disease trajectory Minimum Quality Threshold in Pre-Clinical Sepsis Studies.

The assessment sequence proceeds through three tiers. Tier one is the scheduled observation, performed at intervals determined by the model phase. Tier two is the triggered assessment, initiated when any single parameter crosses a predefined threshold. Tier three is the confirmatory evaluation, performed by a second observer or the attending veterinarian to decide between continued monitoring, intervention, or euthanasia.

Trigger Points That Change the Decision

A single abnormal parameter warrants increased monitoring frequency but does not mandate euthanasia. Two or more parameters in the moderate category, or any single parameter in the severe category, should trigger immediate veterinary assessment. The decision to euthanize rests on trajectory instead of absolute values. An animal losing weight at a steady rate of 5% per day over three days has a different prognosis than one losing 15% in a single day, even if both reach the same cumulative loss.

Pain-related behavior complicates this assessment. Analgesia can alter the clinical signs used for endpoint scoring, and behavioral changes caused by sepsis, pain, and analgesic drugs are difficult to differentiate Analgesia in clinically relevant rodent models of sepsis. Document analgesic administration and account for its effects when interpreting scores.

Clinical Signs and Endpoint Scoring Table

The following table provides a structured scoring framework for the most commonly used clinical parameters in rodent sepsis models. Scores are additive, with a cumulative score of 6 or more, or a score of 3 in any single category, indicating that euthanasia criteria have been met.

ParameterScore 0Score 1Score 2Score 3
Body weight change0 to 5% loss6 to 10% loss11 to 15% lossGreater than 15% loss or 20% loss over 48 hours
Body temperature36.5 to 38.0 C35.0 to 36.4 C33.0 to 34.9 CBelow 33.0 C or above 40.5 C
Spontaneous activityNormal locomotion, grooming, explorationReduced movement, intermittent groomingStationary when undisturbed, no groomingRecumbent, unresponsive to gentle handling
Response to handlingImmediate escape attemptSlowed response, brief resistanceMinimal response, no resistanceNo response, limp body tone
Respiratory effortNormal rate and patternMildly increased rateModerate dyspnea, abdominal breathingSevere dyspnea, gasping, cyanosis
Ocular dischargeNoneMild porphyrin stainingModerate crusting, partial eyelid closureComplete eyelid closure, matted fur
PostureNormal arched back when movingMildly hunched when stationaryHunched with piloerectionSeverely hunched, head lowered, unable to right
Food and water intakeNormal consumptionReduced consumptionMinimal consumption, dehydration evidentNo consumption for 24 hours, skin tenting

Temperature thresholds require adjustment for the measurement method. Infrared thermometry of the tail or paw reads lower than rectal temperature, and the difference varies with ambient temperature and vasomotor tone. Use the same method throughout the study and calibrate thresholds to the chosen technique.

Monitoring Parameters and What Each Detects

Body weight is the most reliable single parameter because it is objective, quantitative, and correlates with disease severity across models. Weight loss reflects a combination of reduced intake, increased metabolic demand, and fluid shifts. The systematic review by Mei and colleagues found large intra- and inter-model variance in endpoint determination, with body weight thresholds varying considerably across published studies Refining humane endpoints in mouse models of disease. This variance argues for model-specific validation of weight loss thresholds instead of reliance on generic cut-offs.

Temperature is a sensitive early indicator of systemic inflammation. Hypothermia in rodents corresponds to the decompensatory phase of sepsis and carries a poor prognosis. Hyperthermia may appear early in some models but is less consistent than in human sepsis. Temperature measurement requires restraint, which itself induces stress and transient hyperthermia. Time temperature measurements to avoid confounding by handling stress.

Clinical appearance parameters detect deterioration that precedes measurable weight or temperature changes. Porphyrin staining around the eyes and nares indicates stress in mice. Piloerection reflects autonomic dysregulation. Hunched posture and reduced grooming are nonspecific but reliable indicators of malaise. The MQTiPSS working group on humane modeling emphasized that euthanasia criteria were defined in only 9% of the 260 most highly cited sepsis studies reviewed, which underscores the need for explicit, published criteria in every protocol MQTiPSS Part I for study design and humane modeling endpoints.

Equipment and Technique Selection

Rectal temperature probes are the reference standard for rodents but carry risks of perforation and stress. Use a probe sized for the species, lubricated, and inserted to a consistent depth. Infrared thermometry is less stressful but less accurate. Implantable telemetry provides continuous temperature and activity data without handling stress, but requires surgical implantation and is not feasible for all studies.

Body weight measurement requires a balance accurate to 0.1 g for mice and 1.0 g for rats. Weigh animals at the same time each day to control for diurnal variation and gastrointestinal fill. Clinical scoring requires a standardized observation environment. Observe animals in their home cage first, then after gentle stimulation, then during handling. This sequence distinguishes spontaneous behavior from response to external stimuli.

Documentation and Communication

Record every observation on the score sheet at the time it is made. Do not rely on memory or batch recording. Each entry should include the date, time, animal identification, body weight, temperature, individual parameter scores, cumulative score, and the observer's initials. Note any analgesic or other treatment administered, as this affects interpretation of subsequent scores.

The score sheet serves multiple functions. It provides the objective basis for euthanasia decisions, documents welfare for regulatory review, and generates data on model progression. The Guide for the Care and Use of Laboratory Animals requires that institutional animal care and use programs ensure adequate veterinary care, which includes oversight of humane endpoint application. Score sheets should be reviewed by the attending veterinarian at scheduled intervals and whenever an animal reaches a trigger threshold.

Define in advance who has authority to euthanize an animal without further consultation. The attending veterinarian must have standing authority to euthanize any animal meeting endpoint criteria. Research staff should be trained to recognize the severe category signs and to contact the veterinary service immediately when these appear. When a second observer is unavailable, err on the side of earlier euthanasia.

Species-Specific Adjustments

Rats and mice differ in their physiological responses to sepsis and in the practical aspects of monitoring. Rats tolerate repeated handling and temperature measurement better than mice. Mice lose body heat more rapidly when ill and develop hypothermia faster. The scoring table applies to both species, but the thresholds for temperature and weight loss may require adjustment based on the specific model and the baseline values of the strain used.

Strain differences affect baseline temperature, activity levels, and susceptibility to sepsis. C57BL/6 mice differ from BALB/c mice in their inflammatory responses and clinical trajectories. Establish strain-specific baselines before the study begins and document them in the protocol. The NC3Rs resources on refinement provide practical guidance on species-specific welfare assessment and endpoint application.

The choice of monitoring equipment also depends on the species. Rectal probes for mice must be smaller and insertion must be shallower than for rats. Infrared thermometry is more commonly used in mice because of their small size. Telemetry implantation is feasible in both species but carries greater surgical risk in mice. Select equipment based on the species, the model, and the number of animals that require simultaneous monitoring.

Recognized Complications and Failure Modes

The most common failure in sepsis endpoint monitoring is not the absence of a score sheet but the absence of a decision hierarchy. A total score that triggers euthanasia only when several parameters accumulate will miss animals that deteriorate rapidly along a single axis. Conversely, a single-parameter trigger such as fixed weight loss can euthanise animals that are recovering. The MQTiPSS consensus recommends that survival follow-up reflect the clinical time course of the infectious agent used, which means endpoint criteria must be calibrated to the expected trajectory of the specific model instead of transferred wholesale between models MQTiPSS humane modeling recommendations.

Hypothermia is the most reliable early indicator of progression in murine sepsis, but it is also the parameter most often measured incorrectly. A single rectal temperature reading taken after handling-induced stress can be artefactually elevated by 1 to 2 degrees Celsius, masking true hypothermia. Infrared thermal imaging of the tail or plantar surface avoids handling stress but measures peripheral temperature, which can lag core changes. The discriminating check is to compare two modalities or to take repeated readings at fixed intervals after the animal has settled.

Piloerection and ocular discharge are scored as present or absent in most systems, yet they are among the first signs to appear and the last to resolve. An animal that has recovered systemically may retain a poor coat for days, and a score sheet that does not separate acute from chronic appearance will overestimate morbidity in the recovery phase. The corrective action is to weight acute parameters such as temperature and respiratory effort more heavily than appearance parameters in the first 48 hours, then reverse that weighting during the recovery window.

Common Errors and Corrective Actions

Less experienced observers tend to score behavior before physical parameters and to anchor on the most recent observation instead of the trend. A mouse that is lethargic at a single check may simply have been asleep, while a mouse that is active at every check except one may be in rapid decline. The corrective action is to record the trend across at least two consecutive observations and to require a confirmatory assessment before any euthanasia decision is executed.

A second recurring error is the failure to distinguish disease-related signs from analgesic side effects. Opioid analgesics can cause respiratory depression, reduced activity, and altered temperature regulation, all of which overlap with sepsis signs. Jeger and colleagues note that information on the efficacy of analgesia in sepsis models is scarce and that score sheets should be used to adapt analgesia or to terminate experiments using humane endpoints analgesia in rodent sepsis models. The practical corrective action is to record baseline behavior and physiological parameters after analgesic administration but before the septic insult, so that drug effects are separated from disease effects.

A third error is the use of a single observer without inter-rater reliability checks. Scoring systems are inherently subjective, and the same animal can receive different scores from different technicians. The corrective action is to have two observers score the same animals independently during the first week of a study and to reconcile discrepancies before the formal monitoring schedule begins.

Limitations of the Current Evidence

The evidence base for humane endpoints in rodent sepsis is heterogeneous. A systematic review of mouse studies using weight, temperature, and sickness scores found large intra- and inter-model variance in endpoint determination, driven by differing animal models, lack of standardized protocols, and heterogeneous performance metrics machine learning-based endpoint definition in mouse models. This means that published cut-off values cannot be treated as transferable constants.

Expert opinion still differs on several points. The MQTiPSS working groups reached consensus on 29 points, but only 20 at recommendation strength, with the remaining nine at consideration strength MQTiPSS executive summary. Areas of residual disagreement include whether weight loss should be expressed as a percentage of starting weight or as deviation from a control group mean, whether temperature should be measured at fixed clock times or at the same interval after the insult, and whether a single moribund state criterion should override all composite scores.

Machine learning approaches have been proposed as a way to define endpoints across models, but their output depends entirely on the quality and completeness of the input data. An algorithm trained on weight, temperature, and sickness scores cannot detect a parameter that was never recorded, and the published work in this area remains exploratory instead of prescriptive machine learning-based endpoint definition in mouse models.

Escalation and Referral

Most endpoint decisions are made by trained animal technicians or research staff, but escalation pathways must be defined in advance. A veterinarian should be consulted when an animal shows signs outside the expected model trajectory, when a new batch of animals shows a different response to the same insult, or when analgesic adjustments do not produce the expected improvement. The Guide for the Care and Use of Laboratory Animals requires that institutional animal care and use programs provide adequate veterinary care, which includes oversight of humane endpoint decisions.

Regulatory reporting is required when unexpected death occurs before the humane endpoint is reached, when an animal is found dead with no prior clinical signs, or when a protocol deviation results in unrelieved pain or distress. Institutional policies vary, and the attending veterinarian should determine whether an adverse event report is required. The AVMA practice resources provide general professional guidance on euthanasia methods and welfare assessment, while the NC3Rs resources offer practical refinement strategies that can reduce the frequency of endpoint triggers.

Troubleshooting Table

ObservationLikely CauseDiscriminating Check
Low temperature score but active behaviorHandling-induced stress before measurementRepeat after 10 minutes undisturbed, compare with infrared peripheral reading
Persistent poor coat after clinical recoveryChronic appearance parameter weighted too heavilyCheck temperature and activity trend, re-weight appearance in recovery phase
Reduced activity after analgesic administrationDrug effect instead of disease progressionCompare with baseline post-analgesia activity recorded before insult
Rapid decline between scheduled checksMonitoring interval too long for model trajectoryIncrease frequency during expected peak severity window
Different scores from different observersInter-rater variabilityIndependent dual scoring with reconciliation before study start
Weight loss below trigger but animal appears wellFixed weight threshold inappropriate for modelReview trend instead of absolute value, consult model-specific MQTiPSS guidance

Frequently Asked Questions

How Do I Set a Humane Endpoint When My Institution Lacks Telemetry or Infrared Thermal Imaging?

Manual assessment remains the standard when advanced monitoring is unavailable. A calibrated rectal or infrared thermometer, a reliable balance, and a validated clinical score sheet provide sufficient data for endpoint decisions. The MQTiPSS consensus on humane modeling endpoints recommends combining weight, temperature, and sickness scores instead of relying on any single parameter. Increase observation frequency to at least twice daily during the predicted peak of the disease course, and ensure the same observer scores consecutive time points to reduce inter-observer variability. Document the limitations of your monitoring approach in the study protocol and in any resulting publications, as this transparency supports accurate interpretation of your endpoint data.

What Should I Do When an Animal Deteriorates Between Scheduled Observation Points?

Unplanned deterioration requires an immediate response protocol. Staff must have clear instructions to interrupt other duties and assess any animal found in a moribund state within 15 minutes. The assessment should follow the same structured sequence used at scheduled time points, with particular attention to unresponsive behavior, labored breathing, and inability to maintain sternal recumbency. If the animal meets any single hard endpoint criterion, euthanasia should proceed without waiting for the next scheduled observation. The systematic review of endpoint definitions in mouse disease models found substantial heterogeneity in how quickly animals were assessed after deterioration, which can confound both welfare outcomes and scientific data. Record the time of discovery, the clinical findings, and the decision made.

How Do Endpoint Criteria Change When Using Older or Immunocompromised Animals?

Age and immune status shift both baseline values and the trajectory of decline. Aged rodents typically have higher baseline body weights but lose weight more slowly relative to their starting mass, so percentage weight loss thresholds may need adjustment. Immunocompromised animals may show less pronounced fever responses and more rapid progression to hypothermia. Establish strain-specific and age-specific baseline data before beginning the study, and pilot the model in a small cohort to characterize the expected clinical course. The MQTiPSS working group recommendations emphasize that survival follow-up should reflect the clinical time course of the infectious agent used, which varies with host factors. Consult your institutional veterinary staff when adapting published thresholds to a new population.

What Records Must I Keep for Endpoint Decisions and How Long Should I Retain Them?

Maintain a contemporaneous log for each animal that includes the date and time of every observation, the individual parameter values, the total score, the observer's identity, and any interventions performed. Photographs or video of representative clinical states can be valuable for training and for retrospective review of borderline decisions. The Guide for the Care and Use of Laboratory Animals describes the institutional oversight framework that governs these records. Retention periods are typically set by institutional policy and funding agency requirements, but five years after study completion is a common minimum. Ensure that records distinguish between scheduled observations and unscheduled assessments triggered by clinical concern.

How Do I Explain an Endpoint Decision to a Supervisor Who Wants to Maximize Data Collection?

Frame the discussion around scientific validity instead of welfare alone. An animal that has passed a validated endpoint threshold produces data that are difficult to interpret, because the physiological state at that point is confounded by agonal change. The MQTiPSS expert consensus explicitly links humane endpoints to translational quality, noting that poorly defined endpoints contribute to the poor reproducibility seen across preclinical sepsis studies. Present the score sheet data showing the animal's trajectory, and explain that the predefined endpoint was designed to protect the scientific objective by ensuring that all animals are assessed at comparable physiological states. Offer to review the endpoint criteria prospectively if the supervisor believes they are too conservative, using pilot data to justify any adjustment.

What Are the Cost Implications of Rigorous Endpoint Monitoring?

The primary costs are personnel time and training, not equipment. Twice-daily scoring of a moderate cohort requires approximately 30 to 60 minutes per day depending on group size. Training costs are incurred upfront but reduce over time as observers become proficient. The NC3Rs resources on refinement provide free score sheet templates and training materials that reduce the need for custom development. Automated monitoring systems such as telemetry or video tracking have significant purchase costs but can reduce personnel demands in long-term studies. Consider whether the increased monitoring frequency reduces total animal numbers by improving data quality, which can offset the personnel cost. Budget for these monitoring activities in the grant application instead of treating them as an afterthought.

Related Clinical & Scientific Guides

References and Further Reading

Related Articles

This article is educational professional reference material for veterinary audiences. It is not a substitute for veterinary diagnosis, individual clinical judgment, current product labeling, or applicable regulatory requirements.