# Scoring Severity of Procedures in Animal Research Protocols


## Key Takeaways

- Severity scoring integrates clinical examination, behavioral assessment, and procedure-specific knowledge to provide a graded, defensible metric for animal welfare impact, serving both prospective protocol review and retrospective refinement.
- A robust scoring framework incorporates multiple domains including pain, distress, functional impairment, and procedure-specific effects, utilizing validated instruments like grimace scales and objective measures such as body weight and temperature.
- Humane endpoints are the practical application of severity scoring, requiring pre-defined, procedure-specific criteria that combine multiple indicators to ensure timely intervention, minimizing animal suffering.
- Observer training and standardized documentation are critical for ensuring inter-rater reliability and reproducibility, with individual animal scoring and regular review of aggregate data driving protocol refinement.
- The distinction between predicted and actual severity is paramount, necessitating reassessment intervals and ad hoc evaluations when clinical signs deviate from the expected course, with the veterinarian's professional judgment guiding intervention.
- Species-specific considerations are essential, as validated scoring instruments and thresholds for parameters like body weight loss or temperature changes vary significantly between rodents, rabbits, and larger laboratory species.

---

Severity scoring translates the cumulative welfare impact of an experimental procedure into a graded, defensible metric. This article provides a framework for assigning severity scores to common experimental procedures in laboratory animals, with emphasis on practical application for veterinary researchers who design protocols, monitor animals, and advise animal care and use committees. The framework integrates clinical assessment, behavioral indicators, and procedure-specific knowledge to produce scores that are reproducible across observers and institutions.

The reader is assumed to be a qualified veterinarian or graduate-level researcher familiar with laboratory animal medicine. The content addresses a specific operational question: given a proposed procedure, what evidence should inform its severity classification, and how should that classification change when an animal's condition deviates from the expected course? The framework draws on established welfare science, published scoring instruments, and the refinement principles articulated by the [NC3Rs resources on replacement, reduction and refinement](https://www.nc3rs.org.uk/).

Severity classification serves two distinct functions. Prospectively, it informs protocol review, housing decisions, and staffing requirements. Retrospectively, it drives refinement by identifying procedures whose actual severity exceeds the predicted level. Both functions require a scoring system that is sensitive to subtle welfare changes, also to overt signs of distress. As demonstrated in studies of voluntary wheel running in mice, conventional clinical scoring can miss meaningful welfare compromise that behavioral measures detect [Häger et al., defining individual severity levels in mice](https://pubmed.ncbi.nlm.nih.gov/30335759/).

## At a Glance

| Parameter | Decision or Fact |
|---|---|
| Purpose of severity scoring | Prospective protocol review and retrospective refinement |
| Primary evidence sources | Clinical examination, behavioral assessment, physiological monitoring |
| Scoring domains | Pain, distress, functional impairment, procedure-specific effects |
| Time course | Score at baseline, during procedure, and at defined recovery intervals |
| Humane endpoints | Predefined, procedure-specific, and triggered by score thresholds |
| Observer training | Required to ensure inter-rater reliability |
| Documentation | Score each animal individually, not as a cohort average |
| Refinement trigger | Actual severity exceeding predicted severity |

## Conceptual Foundations of Severity Assessment

Severity is not a single physiological parameter but a composite of pain, suffering, distress, and lasting harm. The European Union Directive 2010/63/EU, which underpins much of the current severity classification framework, defines severity categories as non-recovery, mild, moderate, and severe. These categories are assigned prospectively based on the expected worst experience of an individual animal, then confirmed retrospectively against observed outcomes [Häger et al., defining individual severity levels in mice](https://pubmed.ncbi.nlm.nih.gov/30335759/).

The distinction between predicted and actual severity is critical. A procedure classified as moderate may prove mild in a well-managed cohort or severe in an individual animal with an unexpected complication. The scoring framework must therefore accommodate both planned reassessment intervals and ad hoc reassessment when clinical signs warrant. The [Guide for the Care and Use of Laboratory Animals](https://grants.nih.gov/grants/olaw/guide-for-the-care-and-use-of-laboratory-animals.pdf) emphasizes that the veterinarian's professional judgment governs when an animal's condition requires intervention, regardless of the protocol's predicted severity category.

### The Relationship Between Severity and Welfare

Welfare assessment and severity scoring are related but not identical exercises. Welfare assessment asks whether an animal's current state is acceptable. Severity scoring asks how much cumulative welfare compromise a procedure imposes. An animal can have acceptable welfare at a given observation point yet still accumulate substantial severity over the course of a study. Conversely, a brief but intense noxious stimulus may produce a high peak severity score despite rapid recovery. Both peak and cumulative severity matter for protocol review, but they inform different decisions.

### Observer Independence and Objectivity

Clinical scoring systems carry inherent subjectivity. The [Murine Sepsis Score and related instruments](https://pubmed.ncbi.nlm.nih.gov/30054760/) demonstrate that structured scoring improves prediction of outcomes compared with unstructured clinical impression, but they also show variability between instruments. The Mouse Clinical Assessment Score for Sepsis showed substantial inter-observer variability in the same study, underscoring the need for standardized training and clear operational definitions for each score level.

Behavioral measures offer a complementary, observer-independent channel. Voluntary wheel running in mice correlates with colitis severity and distinguishes severity levels that clinical scoring misses [Häger et al., defining individual severity levels in mice](https://pubmed.ncbi.nlm.nih.gov/30335759/). Automated home-cage monitoring, food and water intake, and body weight trajectories provide continuous data that can be analyzed with clustering algorithms to define severity boundaries empirically instead of arbitrarily.

## Physiological and Behavioral Domains of Severity

A defensible severity score integrates multiple domains because no single measure captures the full welfare impact of a procedure. The domains below apply across species, though the specific indicators and their weighting vary.

### Pain and Nociception

Pain assessment in laboratory animals relies on behavioral and physiological proxies. Grimace scales, which score orbital tightening, nose and cheek bulge, ear position, and whisker change, have been validated in mice and are increasingly applied to other rodents. These scales detect moderate pain but may miss low-grade or chronic pain. Clinical examination findings such as wound guarding, altered gait, and reduced grooming supplement grimace scoring.

### Distress and Anxiety

Distress manifests as behavioral and autonomic changes that may persist after the painful stimulus has resolved. Stereotypies, barbering, altered sleep-wake cycles, and exaggerated startle responses indicate chronic distress. Physiological correlates include sustained glucocorticoid elevation, though sampling itself can confound the measurement.

### Functional Impairment

A procedure's severity depends partly on how much it compromises normal function. A tumor that impairs ambulation scores higher than one of equal size at a non-weight-bearing site. Similarly, a surgical approach that limits feeding or urination imposes severity beyond the tissue trauma itself. Functional assessment should be species-appropriate and should consider the animal's natural behavior repertoire.

### Procedure-Specific Effects

Some procedures carry severity that is not captured by generic pain or distress scales. Seizure models impose neurological compromise that may not produce grimacing. Chemotherapy protocols cause gastrointestinal and bone marrow toxicity with characteriztic clinical courses. Scoring systems must therefore combine generic welfare indicators with model-specific endpoints. The [validated endoscopic scoring system for murine colitis and colorectal tumors](https://pubmed.ncbi.nlm.nih.gov/24193215/) exemplifies this principle, integrating inflammation extent, tumor burden, and procedural complications into a single numerical output.

## The Role of Humane Endpoints

Humane endpoints are the practical expression of severity scoring. A humane endpoint is the point at which an animal's pain, distress, or suffering is prevented, minimized, or terminated by euthanasia, by removing the cause, or by providing treatment. Severity scores provide the objective basis for defining these endpoints before the study begins.

Endpoint criteria should be procedure-specific and should combine multiple indicators instead of relying on a single sign. In the [cecal ligation and puncture sepsis model](https://pubmed.ncbi.nlm.nih.gov/30054760/), body temperature and composite clinical scores predicted death more reliably than any single parameter, supporting the use of multivariate endpoint criteria. The same study demonstrated that different scoring instruments have different predictive characteriztics, so the choice of instrument should be justified in the protocol.

Body temperature deserves particular attention as an endpoint criterion. Hypothermia is an early and reliable predictor of mortality in rodent sepsis models [Mai et al., body temperature and mouse scoring systems as surrogate markers of death](https://pubmed.ncbi.nlm.nih.gov/30054760/). Temperature monitoring is non-invasive, objective, and can be automated with implantable transponders or infrared thermometry. However, ambient temperature, anesthetic recovery, and handling stress can confound readings, so temperature thresholds should be validated for the specific model and housing conditions.

## Species-Specific Considerations

Severity scoring instruments developed for one species cannot be transferred without validation. The grimace scale, for example, has species-specific facial action units. Rats, rabbits, and non-human primates each require their own validated instruments. Body weight thresholds also differ: a 10 percent weight loss is a common humane endpoint in mice but may be less meaningful in larger species with greater fat reserves.

The [MSD Veterinary Manual](https://www.msdvetmanual.com/) provides species-specific reference values for physiological parameters and clinical signs that inform severity assessment. The [AVMA professional practice resources](https://www.avma.org/resources-tools) offer guidance on euthanasia criteria and anesthetic monitoring that supports endpoint decisions. Where a species-specific scoring instrument does not exist, the protocol should describe how generic indicators will be adapted and validated.

## Limitations of Current Scoring Systems

Published severity scoring systems have recognized gaps. Most were developed for acute procedures and perform poorly for chronic or progressive conditions. Many rely on subjective clinical judgment despite efforts to standardize definitions. The [systematic review of complication reporting in veterinary surgical research](https://pubmed.ncbi.nlm.nih.gov/31290167/) found that even in clinical veterinary medicine, definitions of complications and severity criteria are highly variable and often absent. Laboratory animal severity scoring faces the same challenge: instruments exist, but their use is inconsistent across institutions and studies.

The evidence base for severity scoring is also uneven. Some models, such as rodent sepsis and colitis, have multiple validated instruments. Others, particularly those involving non-human primates or agricultural species, have sparse published validation data. The [WOAH terrestrial animal health standards](https://www.woah.org/en/what-we-do/standards/codes-and-manuals/terrestrial-code-online-access/) address welfare in production and research contexts and provide a reference framework where species-specific laboratory data are lacking.

## A Practical Severity Scoring Matrix

A severity scoring matrix translates the conceptual domains discussed in Part 1 into a structured, defensible framework for protocol review and ongoing animal assessment. The matrix below provides a cross-species template. It is not a substitute for institution-specific scoring instruments, but it offers a common language for veterinary review of proposed procedures and for post-procedural monitoring.

| Severity Category | Expected Duration | Pain or Distress | Functional Impairment | Typical Procedures | Monitoring Frequency |
|---|---|---|---|---|---|
| Minor | Minutes to hours | Transient, mild, self-limiting | Minimal or absent | Single blood collection from a superficial vein, subcutaneous injection, brief manual restraint, imaging under light sedation | Once during recovery, then at next scheduled check |
| Moderate | Hours to days | Controlled with analgesics or anxiolytics, observable but not debilitating | Partial, temporary | Laparotomy with single-organ biopsy, repeated blood sampling over 24 to 72 hours, short-term single-housing for metabolic studies, chemically induced colitis with early humane endpoints | At least twice daily, with a validated clinical score |
| Severe | Days to weeks | Unrelieved or poorly responsive to analgesics, significant suffering | Marked, progressive, or irreversible | Major organ failure models, severe sepsis with fluid resuscitation withheld, tumor burden exceeding approved dimensions, prolonged restraint with sleep deprivation | At least three times daily, with a score that triggers predefined actions |

The boundaries between categories are not fixed. A procedure that is minor in an adult C57BL/6 mouse may be moderate in a neonatal rat or in a geriatric rabbit with concurrent disease. The matrix must be applied with reference to the specific animal, the model, and the institutional endpoint policy.

### Assigning Scores at Protocol Review

The initial severity assignment occurs during protocol review, before any animal is enrolled. The reviewer should work through the procedure step by step, from animal preparation through recovery, and assign a provisional score to each phase. The overall procedure score is the highest phase score, not the average. A procedure that involves brief restraint for an injection (minor) followed by 48 hours of single housing (moderate) receives a moderate score overall.

Several factors shift the assignment upward. General anesthesia itself does not reduce severity if recovery is prolonged or complicated. A procedure requiring neuromuscular blockade with mechanical ventilation is severe regardless of the analgesic plan, because the animal cannot exhibit pain-related behavior while paralysed. The [National Research Council guide for laboratory animal care](https://grants.nih.gov/grants/olaw/guide-for-the-care-and-use-of-laboratory-animals.pdf) emphasizes that the assessment must consider the cumulative effect of multiple procedures performed on the same animal, not each procedure in isolation. A study involving three minor procedures on consecutive days may warrant a moderate classification.

### Scoring During the In-Life Phase

Prospective scoring is necessary because predicted severity and actual severity diverge. The [Häger and colleagues study on defining individual severity levels in mice](https://pubmed.ncbi.nlm.nih.gov/30335759/) demonstrated that clinical scoring alone underestimated welfare compromise in a colitis model, while automated voluntary wheel running detected graded differences that correlated with histological damage. This finding supports the use of multiple, complementary assessment modalities instead of reliance on a single clinical observation sheet.

A practical in-life scoring system should include the following components, each with explicit criteria:

- **Body condition and weight change.** Weight loss of 10 to 15 percent from baseline in a mouse or rat warrants increased monitoring frequency. Weight loss exceeding 20 percent, or weight loss combined with hunched posture and piloerection, should trigger humane endpoint evaluation.
- **Behavioral repertoire.** Decreased voluntary locomotion, reduced grooming, social withdrawal, and altered nesting behavior are early indicators. Automated home-cage monitoring, where available, provides objective, continuous data that manual observation cannot match.
- **Clinical signs.** Respiratory rate and effort, ocular and nasal discharge, fecal consistency, and urine output are procedure-specific indicators. The [murine sepsis scoring comparison by Mai and colleagues](https://pubmed.ncbi.nlm.nih.gov/30054760/) found that body temperature decline was a reliable predictor of progression to endpoint in a caecal ligation and puncture model, and that a composite score outperformed any single parameter.
- **Pain-specific indicators.** Facial grimace scales, writhing, vocalisation, and guarding behavior are procedure-specific. Their absence does not confirm the absence of pain, particularly in prey species that mask signs.

### Monitoring Parameters and What Each Detects

| Parameter | What It Detects | Practical Notes |
|---|---|---|
| Body weight | Overall catabolic state, dehydration, inability or unwillingness to eat or drink | Daily weighing is the single most sensitive simple metric in rodents |
| Body temperature | Thermoregulatory failure, sepsis, shock, anesthetic recovery complications | A drop of more than 2 to 3 degrees Celsius from baseline in a mouse is an urgent finding |
| Voluntary wheel running | Spontaneous motivation, musculoskeletal integrity, overall wellbeing | Requires habituation before baseline collection, not suitable for all strains or models |
| Grimace score | Acute pain, moderate to severe | Requires training and standardized photography or observation, less useful for chronic low-grade pain |
| Respiratory rate and character | Pleural effusion, pneumonia, pulmonary metastasis, anesthetic complications | Count over a fixed interval, do not estimate |
| Posture and coat condition | Chronic pain, distress, systemic illness | Hunched posture with ruffled coat is a late sign in rodents |
| Fecal output and consistency | Gastrointestinal motility, dehydration, colitis severity | The [validated endoscopic scoring system for murine colitis by Kodani and colleagues](https://pubmed.ncbi.nlm.nih.gov/24193215/) provides a direct visual assessment that complements fecal scoring |

### Decision Points and Escalation Criteria

Every scoring system must have predefined action thresholds. These thresholds should be written into the protocol before the study begins and should specify three tiers of response:

1. **Increased monitoring.** Score reaches a predefined threshold but remains below the humane endpoint. Monitoring frequency increases, and the veterinary team is notified.
2. **Intervention.** The animal receives additional analgesia, fluid therapy, or other supportive care. The attending veterinarian may adjust the analgesic plan based on the current formulary and label references.
3. **Humane euthanasia.** The animal reaches the humane endpoint. The protocol must specify the endpoint criteria in objective, measurable terms.

The [AVMA professional practice resources](https://www.avma.org/resources-tools) provide guidance on euthanasia methodology and on the veterinarian's role in endpoint decisions. The veterinarian has the authority to euthanise an animal at any point, regardless of the protocol's predicted severity, when welfare is compromised.

### Documentation and Communication

Severity scores are only useful if they are recorded consistently and reviewed regularly. Each animal should have an individual record that includes the baseline score, daily scores, any interventions, and the final outcome. Aggregate data should be reviewed at least quarterly to identify procedures that consistently score higher than predicted. This review feeds directly into refinement: a procedure that repeatedly causes unexpected weight loss or early endpoint should be modified or replaced.

The [NC3Rs refinement resources](https://www.nc3rs.org.uk/) offer practical guidance on refining specific procedures and on implementing severity assessment in practice. The [systematic review of complication reporting in veterinary surgical research by Follette and colleagues](https://pubmed.ncbi.nlm.nih.gov/31290167/) found that definitions and classification criteria for complications were highly variable across studies, which limits the ability to compare outcomes. The same problem applies to severity scoring in research animals. Institutions should adopt a standardized scoring instrument, train all personnel in its use, and audit inter-observer reliability periodically.

### Species-Specific Adjustments

The correct scoring approach varies with species. In mice and rats, body weight and grimace scoring are practical and well validated. In rabbits, fecal output and food intake are more sensitive indicators than weight, because rabbits lose weight slowly relative to the severity of their illness. In larger species such as dogs and pigs, clinical examination findings, appetite, and interaction with handlers carry more weight, and grimace scales are less developed. The [MSD Veterinary Manual](https://www.msdvetmanual.com/) provides species-specific guidance on clinical signs of pain and distress that can inform scoring instruments for non-rodent species.

Production animals used in research present additional considerations. Pigs and sheep may be group-housed, and individual scoring requires either temporary separation or observation systems that can identify individuals. The [WOAH terrestrial animal health standards](https://www.woah.org/en/what-we-do/standards/codes-and-manuals/terrestrial-code-online-access/) address welfare in the context of animal health and production, and these standards inform the assessment of research animals in agricultural settings.

### Recognized Failure Modes and Early Detection

Severity scoring systems fail in predictable ways. The most common failure is observer drift, where the same clinician assigns different scores to identical presentations over time. This is detected by periodic re-scoring of archived video or photographic records without reference to prior scores. A second failure is the conflation of disease severity with procedure severity. A minor surgical procedure in an animal with advanced systemic disease produces a clinical picture dominated by the disease, not the procedure. The scoring record must separate these contributions, or refinement decisions will target the wrong variable.

A third failure is the reliance on a single parameter. Body temperature alone, for example, can mislead in sepsis models. In a comparison of surrogate endpoint markers, temperature declined in a severity-dependent manner, but the variability of composite clinical scores differed substantially between instruments, with one scoring system showing considerable inconsistency across repeated assessments. Early detection of this failure requires plotting multiple parameters over time and flagging discordance, such as a normal activity score with a falling temperature. A fourth failure is the application of a scoring system beyond its validated range. Endoscopic scoring systems validated for murine colitis, for instance, integrate inflammation and tumor burden in ways that do not transfer to other species or to non-inflammatory lesions.

### Common Errors and Corrective Action

Less experienced scorers tend to anchor on the most salient sign. An animal with severe piloerection but normal posture may be scored higher than one with mild piloerection and a hunched stance, even when the composite picture is worse in the second case. The corrective action is to score each domain independently and sum the result before assigning a category, never to assign a global score first.

A second common error is the under-recognition of behavioral suppression. Inactivity is frequently read as calm. Voluntary wheel running in mice has been shown to reflect colitis severity directly, while conventional clinical scoring indicated only marginal welfare compromise. The corrective action is to treat a decline in spontaneous behavior as a positive finding, not an absence of findings. A third error is the failure to define terms. In a systematic review of surgical complication reporting in dogs and cats, 92 percent of articles mentioned complications, but only 7.3 percent defined the term, and classification criteria were highly variable or absent. The corrective action is to adopt a published classification scheme and record the definition in the protocol before the study begins.

### Limitations of Current Evidence

The evidence base for severity scoring is uneven. Most validated instruments exist for mice, particularly for colitis, sepsis, and restraint stress models. Data for rats, rabbits, and larger species rely more heavily on extrapolation and expert opinion. The National Research Council guidance describes the obligation to assess pain and distress but does not prescribe a universal scoring instrument, leaving institutions to adapt published tools to local conditions. Expert opinion still differs on whether composite scores or single physiological parameters should drive endpoint decisions. The sepsis literature suggests that a combination outperforms either alone, but the optimal weighting is unresolved. There is also genuine uncertainty about the predictive validity of grimace scales across strains and sexes, and about whether behavioral suppression reflects pain, distress, or both.

### Escalation and Referral Criteria

Escalation is warranted when a score reaches a pre-defined threshold, when a score increases across two consecutive observations, or when any single parameter reaches a humane endpoint criterion. The threshold must be defined in the protocol and approved before the study begins. Referral to a veterinary specialist is appropriate when the cause of a rising score is unclear, when pain is suspected but not controlled, or when a procedure produces an unexpected complication. Laboratory involvement is indicated when clinical signs suggest a systemic inflammatory response, infection, or organ dysfunction that requires diagnostic confirmation. Regulatory reporting is required when an animal is found dead unexpectedly, when a procedure causes more severe effects than predicted at protocol review, or when a humane endpoint is exceeded. The NC3Rs provides practical guidance on refining procedures and improving welfare, and their resources should be consulted when a protocol produces recurrent unexpected severity. Institutional policies and applicable animal welfare standards, such as those published by the World Organization for Animal Health, govern the specific reporting pathway.

### Troubleshooting Guide

| Observation | Likely Cause | Discriminating Check |
| --- | --- | --- |
| Score rises but animal appears bright | Observer drift or over-weighting of one domain | Re-score from video, compare domain totals |
| Activity falls, clinical score normal | Behavioral suppression under-read | Measure voluntary activity or open-field behavior |
| Temperature falls, other scores stable | Impending decompensation | Increase monitoring frequency, check hydration and perfusion |
| Composite score varies between observers | Instrument ambiguity | Review definitions, retrain on anchor cases |
| Score plateaus above threshold | Inadequate analgesia or endpoint delay | Escalate to veterinary review, apply humane endpoint |
| Unexpected death with low prior scores | Scoring interval too long or wrong parameters | Shorten interval, add a more sensitive parameter |

## Frequently Asked Questions

### How Do I Assign a Severity Score When the Protocol Uses a Novel Procedure With No Published Precedent?

Base the initial score on the closest analogous procedure with published severity data, then add a margin for uncertainty. Review the physiological and behavioral domains most likely to be affected and score each conservatively. For example, a novel minimally invasive imaging technique might be scored against published endoscopic procedures in mice, which have validated scoring systems for inflammation and lesion burden [flexible colonoscopy scoring in mice](https://pubmed.ncbi.nlm.nih.gov/24193215/). State explicitly in the protocol that the score is provisional. Schedule more frequent monitoring during the first cohort and include a predefined trigger for reassessment if observed signs exceed the predicted level.

### What Should I Do When the Ideal Monitoring Equipment Is Not Available?

Use the most sensitive observer-based tool you can implement reliably. Body temperature measurement and structured clinical scoring outperform unstructured observation for predicting deterioration in rodent sepsis models [body temperature and mouse scoring systems as sepsis markers](https://pubmed.ncbi.nlm.nih.gov/30054760/). If telemetry or automated home-cage monitoring is unavailable, increase the frequency of manual assessments and use a published scoring instrument instead of ad hoc notes. Voluntary wheel running offers an observer-independent behavioral readout that requires only standard caging and can grade severity without disturbing the animal [wheel running as a severity grading method in mice](https://pubmed.ncbi.nlm.nih.gov/30335759/). Document the limitation and justify why the alternative remains adequate for welfare oversight.

### How Does Severity Scoring Differ Between Rodents and Larger Laboratory Species?

Rodent scoring relies heavily on composite clinical instruments, grimace scales, and behavioral readouts because individual examination is rapid and inexpensive. Larger species permit more granular physiological monitoring, serial blood sampling, and imaging, but each assessment carries greater handling stress and cost. The same procedure, for example laparotomy, may receive a lower score in a dog than a mouse because the larger patient tolerates the physiological insult better and postoperative analgesia is easier to titrate. Species-specific reference values for pain and distress are available in standard veterinary references [MSD Veterinary Manual professional edition](https://www.msdvetmanual.com/). Always calibrate thresholds for humane intervention to the species, not to a generic mammalian baseline.

### How Should I Handle Disagreement Between Observers Scoring the Same Animal?

Treat disagreement as data, not failure. Have both observers score the animal independently and record the discrepancy. If scores differ by more than one severity grade, the animal should be assessed by a third observer or the attending veterinarian. Persistent inter-observer disagreement often indicates that the scoring instrument lacks clear anchor descriptions. Revise the instrument with explicit behavioral definitions and test it again. Published instruments that include unambiguous criteria, such as the Murine Sepsis Score, reduce variability when observers are trained together [body temperature and mouse scoring systems as sepsis markers](https://pubmed.ncbi.nlm.nih.gov/30054760/). Document the resolution process in the animal record.

### What Records Must I Keep for Severity Scoring Beyond the Raw Scores?

Record the time and date of each assessment, the observer identity, the scoring instrument version, and any deviation from the monitoring schedule. Note the animal's body weight, body temperature if measured, and the specific findings that triggered each score. Record any intervention taken, the time to response, and whether the animal returned to its baseline. This level of detail supports retrospective review when a protocol amendment is needed. The [Guide for the Care and Use of Laboratory Animals](https://grants.nih.gov/grants/olaw/guide-for-the-care-and-use-of-laboratory-animals.pdf) emphasizes that veterinary care records must document animal welfare assessments and clinical decisions. Structured records also allow you to compare severity across studies and refine future protocols.

### How Do I Explain a Severity Score to a Supervisor or IACUC Member Who Questions It?

Present the score as the output of a structured assessment, not a subjective judgment. Show the component scores for each domain, the monitoring data that support them, and the published benchmark you used for comparison. If the score is higher than expected, explain which physiological or behavioral parameter drove the escalation and what intervention was triggered. Reference the refinement framework that underpins the assessment, such as the [NC3Rs guidance on the 3Rs](https://www.nc3rs.org.uk/), to frame the score as part of continuous improvement instead of a static label. Be prepared to justify why the score changed over time and what evidence would prompt a further change.

## Related Clinical & Scientific Guides

* [Refining IACUC Protocols to Minimize Animal Pain and Distress](/knowledge/veterinary-medicine/laboratory-animal-science/refining-iacuc-protocols-minimize-animal-pain-distress)
* [Health Monitoring Programs for Laboratory Animal Facilities](/knowledge/veterinary-medicine/laboratory-animal-science/health-monitoring-programs-for-laboratory-animal-facilities)
* [Anesthetic Risk Assessment in Laboratory Animals: Preoperative Evaluation](/knowledge/veterinary-medicine/laboratory-animal-science/anesthetic-risk-assessment-in-laboratory-animals-preoperative-evaluation)


## References and Further Reading

- [Running in the wheel: Defining individual severity levels in mice.](https://pubmed.ncbi.nlm.nih.gov/30335759/). 2018.
- [Flexible colonoscopy in mice to evaluate the severity of colitis and colorectal tumors using a validated endoscopic scoring system.](https://pubmed.ncbi.nlm.nih.gov/24193215/). 2013.
- [Prioritization of Zoonotic Diseases in Kenya, 2015.](https://pubmed.ncbi.nlm.nih.gov/27557120/). 2016.
- [Body temperature and mouse scoring systems as surrogate markers of death in cecal ligation and puncture sepsis.](https://pubmed.ncbi.nlm.nih.gov/30054760/). 2018.
- [A systematic review of criteria used to report complications in soft tissue and oncologic surgical clinical research studies in dogs and cats.](https://pubmed.ncbi.nlm.nih.gov/31290167/). 2020.
- [Comparison between COPD Assessment Test (CAT) and modified Medical Research Council (mMRC) dyspnea scores for evaluation of clinical symptoms, comorbidities and medical resources utilization in COPD patients.](https://pubmed.ncbi.nlm.nih.gov/30150099/). 2019.
- [Guide for the Care and Use of Laboratory Animals, 8th Edition](https://grants.nih.gov/grants/olaw/guide-for-the-care-and-use-of-laboratory-animals.pdf). National Academies Press, 2011.
- [NC3Rs Resources on Replacement, Reduction and Refinement](https://www.nc3rs.org.uk/). NC3Rs.
- [MSD Veterinary Manual, Professional Edition](https://www.msdvetmanual.com/). MSD Veterinary Manual.

## Related Articles

- [Anesthesia for Laboratory Rabbits: Protocols and Monitoring](/knowledge/veterinary-medicine/laboratory-animal-science/anesthesia-for-laboratory-rabbits-protocols-and-monitoring)
- [Animal Model Selection for Neurological Research](/knowledge/veterinary-medicine/laboratory-animal-science/animal-model-selection-for-neurological-research)
- [Selecting Animal Models for Neurological Research](/knowledge/veterinary-medicine/laboratory-animal-science/selecting-animal-models-neurological-research)
- [Selecting Appropriate Animal Models for Pain Research](/knowledge/veterinary-medicine/laboratory-animal-science/selecting-appropriate-animal-models-for-pain-research)
- [Applying the 3Rs in Veterinary Research: Practical Examples](/knowledge/veterinary-medicine/laboratory-animal-science/applying-the-3rs-in-veterinary-research-practical-examples)

> This article is educational professional reference material for veterinary audiences. It is not a substitute for veterinary diagnosis, individual clinical judgment, current product labeling, or applicable regulatory requirements.


<div data-calculator="fluid-rate"></div>