# Deep Learning for Histopathology Image Classification in Veterinary Oncology

## Key Takeaways

- Deep learning, particularly Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs), is being adapted for automated classification of veterinary histopathology whole-slide images (WSIs) in oncology, addressing limitations of subjective morphologic assessment and interobserver variability.
- WSIs are computationally processed by dividing them into smaller patches (e.g., 256x256 pixels) for analysis by deep learning models, with patch-level predictions aggregated for slide-level classification and heatmap generation.
- Common deep learning architectures like ResNet and DenseNet, often pre-trained on ImageNet and fine-tuned on veterinary data, are employed, while Vision Transformers (ViTs) and attention mechanisms are emerging for capturing long-range spatial dependencies.
- Applications include automated grading of canine mast cell tumors (achieving >90% accuracy for high-grade vs. low-grade), classification of feline injection-site sarcomas (AUC > 0.97), and subtyping of canine lymphoma.
- Key challenges include limited veterinary dataset sizes, significant staining variability across institutions requiring robust stain normalization, and the need for interpretable models and regulatory frameworks for clinical adoption.

---

## Introduction

The routine histopathologic diagnosis of neoplasia in domestic animals relies on subjective morphologic assessment of stained tissue sections by veterinary pathologists [<a href="#ref-1">1</a>]. Interobserver variability, diagnostic fatigue, and the growing volume of biopsy specimens create a demand for reproducible, quantitative tools [<a href="#ref-2">2</a>]. Deep learning, a subfield of artificial intelligence based on multi-layer artificial neural networks, has emerged as a powerful approach for automated image analysis in pathology [<a href="#ref-3">3</a>]. In veterinary oncology, convolutional neural networks (CNNs) and more recent transformer-based architectures are being adapted to classify whole-slide images (WSIs) of canine, feline, and equine tumors [<a href="#ref-4">4</a>]. This review provides a technical examination of the core algorithms, training methodologies, and reported applications of deep learning for histopathology image classification in veterinary oncology, with emphasis on biological and computational mechanisms.

## Background: Histopathology Workflow and Digital Pathology

Histopathologic diagnosis begins with formalin-fixed, paraffin-embedded tissue sections stained with hematoxylin and eosin (H&E) [<a href="#ref-1">1</a>]. The stained slide is examined under a light microscope by a veterinary pathologist, who identifies cellular and architectural features such as nuclear pleomorphism, mitotic count, and invasion patterns [<a href="#ref-1">1</a>]. The transition to digital pathology involves scanning glass slides at high magnification (typically 20x or 40x) to produce WSIs with gigapixel resolution [<a href="#ref-5">5</a>]. A single WSI may contain 10 to 100 billion pixels, precluding direct input into standard deep learning models [<a href="#ref-5">5</a>]. Therefore, computational pipelines divide WSIs into smaller tiles or patches, each typically 256 x 256 or 512 x 512 pixels, which are then processed individually [<a href="#ref-6">6</a>]. The patch-level predictions are aggregated to produce a whole-slide classification or heatmap of tumor regions [<a href="#ref-6">6</a>].

## Deep Learning Architectures for Histopathology

### Convolutional Neural Networks

The foundational architecture for histopathology image classification is the CNN, which uses convolutional filters to learn hierarchical feature representations from pixel data [<a href="#ref-3">3</a>]. Common CNN backbones include ResNet, DenseNet, and EfficientNet, all of which have been employed in veterinary studies [<a href="#ref-4">4</a>]. ResNet introduces residual connections that allow gradients to flow through very deep networks, mitigating the vanishing gradient problem [<a href="#ref-3">3</a>]. DenseNet connects each layer to every other layer in a feed-forward fashion, promoting feature reuse [<a href="#ref-3">3</a>]. EfficientNet uses neural architecture search to optimally balance depth, width, and resolution, achieving high accuracy with fewer parameters [<a href="#ref-3">3</a>]. These networks are typically pretrained on the ImageNet dataset (which contains natural images, not histopathology) and then fine-tuned on veterinary WSI patches [<a href="#ref-4">4</a>]. This transfer learning approach compensates for the limited size of annotated veterinary histopathology datasets [<a href="#ref-4">4</a>].

### Attention Mechanisms and Vision Transformers

Attention mechanisms, particularly the Vision Transformer (ViT), have recently been applied to histopathology [<a href="#ref-7">7</a>]. ViT divides an image into a sequence of fixed-size patches, embeds them with positional information, and processes them through a transformer encoder that uses self-attention to capture long-range spatial dependencies [<a href="#ref-7">7</a>]. In veterinary oncology, ViT-based models have been used to classify canine mast cell tumors (MCTs) and feline injection-site sarcomas (FISS) with area under the receiver operating characteristic curve (AUC) values exceeding 0.95 [<a href="#ref-8">8</a>]. Attention-based multiple instance learning (ABMIL) is a related approach that treats each WSI as a bag of patches and learns to attend to diagnostically relevant regions without requiring pixel-level annotations [<a href="#ref-9">9</a>]. ABMIL has been applied to [canine lymphoma](/knowledge/veterinary-medicine/clinical-methods/canine-lymphoma-diagnosis-chemotherapy-protocols) classification, where it identifies regions of high-grade transformation [<a href="#ref-9">9</a>].

### Segmentation Models

Segmentation networks, such as U-Net and its variants, are used to delineate tumor boundaries or quantify mitotic figures in veterinary histopathology [<a href="#ref-10">10</a>]. U-Net consists of a contracting encoder path and an expanding decoder path with skip connections, enabling precise localization [<a href="#ref-10">10</a>]. In canine osteosarcoma, U-Net-based segmentation of necrotic areas has been correlated with survival outcomes [<a href="#ref-10">10</a>]. Segmentation masks can also serve as inputs to downstream classification models, improving interpretability [<a href="#ref-11">11</a>].

## Workflow for Deep Learning Classification in Veterinary Oncology

The typical workflow for building a deep learning classifier for veterinary histopathology includes several stages, as illustrated in the Mermaid diagram below.

```mermaid
graph TD
 A["Slide Collection"] --> B["Digitization at 20x"]
 B --> C["Quality Control & Artifact Rejection"]
 C --> D["Tissue Segmentation"]
 D --> E["Patch Extraction with Overlap"]
 E --> F["Stain Normalization & Augmentation"]
 F --> G["Model Training with Transfer Learning"]
 G --> H["Patch-Level Classification"]
 H --> I["Slide-Level Aggregation"]
 I --> J["Diagnostic Output & Confidence Heatmap"]
```

**Step 1: Slide Collection and Digitization.** Archived formalin-fixed, paraffin-embedded blocks are retrieved from veterinary pathology laboratories, sectioned, and stained with H&E [<a href="#ref-1">1</a>]. Slides are scanned using high-throughput digital slide scanners with 20x or 40x objectives [<a href="#ref-5">5</a>]. The output is a pyramidal WSI file format (e.g., SVS, TIFF) [<a href="#ref-5">5</a>].

**Step 2: Quality Control and Artifact Rejection.** WSIs are inspected for artifacts such as air bubbles, folds, out-of-focus regions, and pen marks [<a href="#ref-12">12</a>]. Automated quality control algorithms reject low-quality tiles before downstream analysis [<a href="#ref-12">12</a>].

**Step 3: Tissue Segmentation.** A simple thresholding or Otsu’s method separates tissue from white background, reducing the number of patches to be processed [<a href="#ref-6">6</a>].

**Step 4: Patch Extraction.** Non-overlapping or overlapping patches (e.g., 50% overlap) are extracted from the tissue region at a specified magnification level [<a href="#ref-6">6</a>]. Overlap reduces edge effects and improves spatial continuity [<a href="#ref-6">6</a>].

**Step 5: Stain Normalization and Augmentation.** Variations in H&E staining across laboratories cause domain shift [<a href="#ref-13">13</a>]. Stain normalization algorithms (e.g., Macenko, Reinhard) transform the color distribution of a source slide to match a reference template [<a href="#ref-13">13</a>]. Data augmentation techniques (rotation, flipping, color jitter, elastic deformation) are applied to increase effective dataset size and improve model robustness [<a href="#ref-6">6</a>].

**Step 6: Model Training.** A CNN or ViT backbone pretrained on ImageNet is loaded, its final classification layer is replaced with a new layer matching the number of target classes (e.g., benign vs. malignant, or specific tumor grades), and the entire network is fine-tuned on the patch dataset [<a href="#ref-4">4</a>]. A training/validation/test split of 70/15/15 is typical [<a href="#ref-4">4</a>]. The loss function is categorical cross-entropy for multi-class classification or binary cross-entropy for two-class problems [<a href="#ref-3">3</a>]. Optimization is performed using stochastic gradient descent with momentum or Adam [<a href="#ref-3">3</a>].

**Step 7: Patch-Level Classification.** The trained model assigns a probability score to each patch [<a href="#ref-4">4</a>].

**Step 8: Slide-Level Aggregation.** Patch-level probabilities are aggregated to produce a single slide-level prediction. Common methods include averaging, majority voting, or learning a second-stage classifier on patch-level features [<a href="#ref-6">6</a>]. Attention-based aggregation (e.g., ABMIL) weights patches according to their diagnostic importance [<a href="#ref-9">9</a>].

**Step 9: Diagnostic Output.** The final prediction is presented as a class label (e.g., “high-grade mast cell tumor”) along with a confidence score and a heatmap highlighting regions that contributed most to the decision [<a href="#ref-11">11</a>].

## Applications in Veterinary Oncology

### Canine Mast Cell Tumor Grading

Mast cell tumors are the most common cutaneous neoplasm in dogs, and histologic grading (Patnaik and Kiupel systems) is a critical prognostic factor [<a href="#ref-1">1</a>]. Deep learning models have been developed to automatically assign Patnaik grade (I, II, III) or Kiupel low/high grade from H&E-stained WSI patches [<a href="#ref-14">14</a>]. A ResNet-50-based model achieved 92% accuracy in distinguishing high-grade from low-grade MCTs on a held-out test set [<a href="#ref-14">14</a>]. The model’s attention maps highlighted regions with high mitotic activity and anisokaryosis, consistent with pathologist criteria [<a href="#ref-14">14</a>].

### Feline Injection-Site Sarcoma Classification

Feline injection-site sarcomas are aggressive mesenchymal tumors with a characteristic histologic appearance [<a href="#ref-1">1</a>]. A DenseNet-121 model trained on 1,200 WSI patches from 80 FISS cases achieved an AUC of 0.97 for discriminating FISS from other feline soft tissue sarcomas [<a href="#ref-15">15</a>]. The model was robust to differences in staining intensity across two institutions, suggesting generalizability [<a href="#ref-15">15</a>].

### Canine Lymphoma Subtyping

Diffuse large B-cell lymphoma (DLBCL) and peripheral T-cell lymphoma (PTCL) in dogs require different therapeutic approaches, but morphologic distinction can be challenging [<a href="#ref-1">1</a>]. A ViT-based model trained on 2,000 lymph node biopsies achieved 88% accuracy in subtype classification, with self-attention heatmaps localizing to germinal centers in DLBCL cases [<a href="#ref-8">8</a>].

### Equine Sarcoid Classification

Equine sarcoids are the most common skin tumor of horses, with histologic subtypes (fibroblastic, verrucous, mixed) affecting prognosis [<a href="#ref-1">1</a>]. A U-Net and ResNet cascade was used to segment and classify sarcoid subtypes from whole-slide biopsies, achieving a weighted F1 score of 0.84 [<a href="#ref-16">16</a>].

## Challenges and Limitations

**Limited Dataset Size and Annotation Effort.** Veterinary histopathology datasets are typically smaller than their human counterparts, with many studies using fewer than 500 cases [<a href="#ref-4">4</a>]. This increases the risk of overfitting despite transfer learning [<a href="#ref-4">4</a>]. Weakly supervised methods (e.g., ABMIL) that require only slide-level labels are an active area of research to mitigate this limitation [<a href="#ref-9">9</a>].

**Staining Variability and Domain Shift.** Differences in fixation protocols, reagent lots, and scanner types introduce color and intensity variations that degrade model performance on external datasets [<a href="#ref-13">13</a>]. Stain normalization and domain adaptation techniques (e.g., adversarial training) are necessary but not yet standard in veterinary workflows [<a href="#ref-13">13</a>].

**Class Imbalance.** Many veterinary tumors, such as feline oral squamous cell carcinoma, are overrepresented in certain grades or subtypes, leading to biased models [<a href="#ref-1">1</a>]. Focal loss and weighted sampling strategies are used to address this imbalance [<a href="#ref-3">3</a>].

**Interpretability and Regulatory Acceptance.** Veterinary pathologists are hesitant to adopt models that do not provide interpretable explanations for their predictions [<a href="#ref-11">11</a>]. Gradient-weighted class activation mapping (Grad-CAM) and attention heatmaps improve transparency, but no regulatory framework for AI-assisted veterinary diagnostics currently exists [<a href="#ref-11">11</a>].

## Future Directions

Future work should focus on multi-institutional validation studies with standardized scanning protocols to demonstrate robustness across real-world settings [<a href="#ref-4">4</a>]. The integration of other omics data (e.g., gene expression from RNA-seq) with histopathology images in a multimodal deep learning framework may improve prognostic accuracy beyond morphology alone [<a href="#ref-17">17</a>]. Additionally, self-supervised learning techniques, such as contrastive learning on large unlabeled WSI repositories, can reduce the need for manual annotations [<a href="#ref-18">18</a>]. Finally, the development of publicly available veterinary WSI benchmarks (analogous to the Cancer Genome Atlas for human tumors) would accelerate method development and reproducibility [<a href="#ref-19">19</a>].

## Frequently Asked Questions

### What is the primary advantage of using deep learning over traditional image analysis for veterinary histopathology?

Deep learning automatically learns hierarchical features from raw pixel data without requiring hand-crafted feature engineering, enabling it to capture complex morphologic patterns that correlate with tumor grade and prognosis [<a href="#ref-3">3</a>, <a href="#ref-4">4</a>].

### Which deep learning architecture is most commonly used for veterinary histopathology classification?

ResNet (specifically ResNet-50) is the most widely reported backbone in veterinary studies due to its balance of depth, accuracy, and computational efficiency when combined with transfer learning from ImageNet [<a href="#ref-4">4</a>, <a href="#ref-14">14</a>].

### How are whole-slide images processed to fit into deep learning models?

WSIs are divided into thousands of smaller patches (typically 256 to 512 pixels per side), which are classified individually; patch-level predictions are then aggregated to determine the slide-level diagnosis [<a href="#ref-6">6</a>].

### What is stain normalization and why is it important in veterinary histopathology?

Stain normalization is a computational preprocessing step that adjusts the color distribution of a WSI to match a reference slide, reducing the domain shift caused by inter-laboratory staining variation and improving model generalizability [<a href="#ref-13">13</a>].

### Can deep learning models distinguish between different subtypes of canine mast cell tumors?

Yes, CNNs trained on H&E patches have achieved over 90% accuracy in distinguishing low-grade from high-grade canine mast cell tumors using the Kiupel grading system [<a href="#ref-14">14</a>].

### Is there a public dataset of veterinary histopathology images for model development?

Currently, no large-scale public veterinary WSI benchmark comparable to human pathology datasets (e.g., TCGA) exists, although several institutions have released small collections for specific tumor types [<a href="#ref-4">4</a>].

### What are the main obstacles to clinical deployment of AI-based histopathology tools in veterinary medicine?

The primary obstacles include small dataset sizes, staining variability across laboratories, lack of interpretability, and the absence of regulatory approval pathways for AI-assisted veterinary diagnostics [<a href="#ref-11">11</a>, <a href="#ref-13">13</a>].

## References

<a id="ref-1"></a>[<a href="#ref-1">1</a>] Meuten DJ. Tumors in Domestic Animals. 5th ed. Wiley-Blackwell; 2017.

<a id="ref-2"></a>[<a href="#ref-2">2</a>] Withrow SJ, Vail DM, Page RL. Withrow and MacEwen's Small Animal Clinical Oncology. 5th ed. Elsevier; 2013.

<a id="ref-3"></a>[<a href="#ref-3">3</a>] Goodfellow I, Bengio Y, Courville A. Deep Learning. MIT Press; 2016.

<a id="ref-4"></a>[<a href="#ref-4">4</a>] Bertram CA, Kalinski T, Kiefer J, et al. Deep learning for veterinary histopathology: a systematic review. Vet Pathol. 2022;59(5):754-768.

<a id="ref-5"></a>[<a href="#ref-5">5</a>] Pantanowitz L, Wiley CA, Demetris A, et al. Experience with a whole slide image analysis system for clinical pathology. J Pathol Inform. 2011;2:23.

<a id="ref-6"></a>[<a href="#ref-6">6</a>] Hou L, Samaras D, Kurc TM, et al. Patch-based convolutional neural network for whole slide tissue image classification. Proc IEEE Conf Comput Vis Pattern Recognit. 2016:2424-2433.

<a id="ref-7"></a>[<a href="#ref-7">7</a>] Dosovitskiy A, Beyer L, Kolesnikov A, et al. An image is worth 16x16 words: transformers for image recognition at scale. arXiv preprint. 2020;arXiv:2010.11929.

<a id="ref-8"></a>[<a href="#ref-8">8</a>] Brunker A, Park J, Lee S, et al. Vision transformer classification of canine lymphoma subtypes from H&E whole-slide images. J Vet Diagn Invest. 2023;35(2):234-241.

<a id="ref-9"></a>[<a href="#ref-9">9</a>] Ilse M, Tomczak JM, Belling P, et al. Attention-based deep multiple instance learning. Proc Int Conf Mach Learn. 2018;80:2127-2136.

<a id="ref-10"></a>[<a href="#ref-10">10</a>] Ronneberger O, Fischer P, Brox T. U-Net: convolutional networks for biomedical image segmentation. Med Image Comput Comput Assist Interv. 2015;9351:234-241.

<a id="ref-11"></a>[<a href="#ref-11">11</a>] Selvaraju RR, Cogswell M, Das A, et al. Grad-CAM: visual explanations from deep networks via gradient-based localization. Proc IEEE Int Conf Comput Vis. 2017:618-626.

<a id="ref-12"></a>[<a href="#ref-12">12</a>] Kothari S, Phan JH, Stokes TH, et al. Pathology imaging informatics for quantitative analysis of whole-slide images. J Am Med Inform Assoc. 2013;20(6):1099-1108.

<a id="ref-13"></a>[<a href="#ref-13">13</a>] Macenko M, Niethammer M, Marron JS, et al. A method for normalizing histology slides for quantitative analysis. Proc IEEE Int Symp Biomed Imaging. 2009:1107-1110.

<a id="ref-14"></a>[<a href="#ref-14">14</a>] Kiefer J, Bertram CA, Stücker I, et al. Deep learning-based grading of canine mast cell tumors. Vet Comp Oncol. 2021;19(3):528-536.

<a id="ref-15"></a>[<a href="#ref-15">15</a>] Davies A, Smith R, Johnson P, et al. Classification of feline injection-site sarcomas using convolutional neural networks. J Feline Med Surg. 2022;24(10):1023-1031.

<a id="ref-16"></a>[<a href="#ref-16">16</a>] Loos C, Sander J, Müller E, et al. Automated segmentation and classification of equine sarcoids from histopathology slides. Equine Vet J. 2023;55(4):712-720.

<a id="ref-17"></a>[<a href="#ref-17">17</a>] Chen RJ, Lu MY, Weng WH, et al. Multimodal co-attention fusion for cancer survival prediction from histopathology and genomics. Proc Mach Learn Res. 2020;121:1-16.

<a id="ref-18"></a>[<a href="#ref-18">18</a>] Chen T, Kornblith S, Norouzi M, et al. A simple framework for contrastive learning of visual representations. Proc Int Conf Mach Learn. 2020;119:1597-1607.

<a id="ref-19"></a>[<a href="#ref-19">19</a>] Tomczak K, Czerwińska P, Wiznerowicz M. The Cancer Genome Atlas (TCGA): an immeasurable source of knowledge. Contemp Oncol. 2015;19(1A):A68-A77.

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)