Zubair Khalid

Virologist/Molecular Biologist | Veterinarian | Bioinformatician

Conventional & Molecular Virology • Vaccine Development • Computational Biology

Dr. Zubair Khalid is a veterinarian and virologist specializing in conventional and molecular virology, vaccine development, and computational biology. Dedicated to advancing animal health through innovative research and multi-omics approaches.

Dr. Zubair Khalid - Veterinarian, Virologist, and Vaccine Development Researcher specializing in Computational Biology, Multi-omics, Animal Health, and Infectious Disease Research

Category: Blog

Auditory Communication: How Animals Use Sound to Convey Messages

Animals across nearly every major taxonomic group produce and interpret sound for social coordination, survival, and reproduction. Auditory communication involves three linked components: signal production, transmission through the environment, and reception by a listener that changes its behavior or physiology. This article examines how birds, mammals, amphibians, and insects use sound for mating, warning, and territorial defense, with attention to the neural mechanisms that support these behaviors and the practical tools now available for studying them.

At a Glance: Auditory Signals Across Species

The table below summarizes representative auditory signals, their primary functions, and the contexts in which they occur. These examples illustrate the diversity of acoustic communication strategies across the animal kingdom.

Species Signal Type Primary Purpose Notable Feature
Domestic cat (Felis catus) Meows, purrs, hisses, growls Social interaction, internal state expression Up to 21 distinct vocalizations described in the literature
Concave-eared torrent frog (Amolops tormotus) Bird-like melodic advertisement calls Mate attraction Produces and detects ultrasonic frequencies above 20 kHz
North Atlantic right whale (Eubalaena glacialis) Low-frequency upcalls Contact and individual identification Upcalls encode individual identity across ages and sexes
Zebra finch (Taeniopygia guttata) Learned song Mate attraction and sexual selection Song structure influenced by developmental environment
House mouse (Mus musculus) Ultrasonic vocalizations Social interaction Frequencies outside human hearing range
Great reed warbler (Acrocephalus arundinaceus) Complex song syllables Territorial defense and mate attraction Syllable classification improved by advanced signal processing
Male grasshopper (Glyptobothrus maritimus) Stridulation Courtship signaling Sound-producing organ density varies with traffic noise exposure

The Building Blocks of Auditory Signals

Sound signals vary along several physical dimensions that determine their function and range. Frequency, measured in hertz, describes the pitch of a signal. Amplitude, measured in decibels, describes loudness. Duration and temporal patterning carry additional information. Animals exploit these dimensions in species-specific ways.

Most birds, amphibians, and reptiles detect and produce sounds below approximately 12 kHz. Mammals including bats, cetaceans, and some rodents extend this range into ultrasound, defined as frequencies greater than 20 kHz. The concave-eared torrent frog from Huangshan Hot Springs in China represents a striking exception among amphibians. Males produce diverse bird-like melodic calls with pronounced frequency modulations that contain spectral energy in the ultrasonic range. Acoustic playback experiments in the frogs' natural habitat showed that both audible and ultrasonic call components evoke male vocal responses. Electrophysiological recordings from the auditory midbrain confirmed ultrasonic hearing capacity in this species and in a sympatric species facing similar environmental constraints. This upward extension of both call harmonics and hearing sensitivity likely co-evolved in response to intense, predominantly low-frequency ambient noise from fast-flowing streams (ultrasonic communication in frogs).

Signal structure often reflects the acoustic properties of the habitat. Species living near noisy water sources face masking of low-frequency signals. The torrent frog example demonstrates how selection can shift signal energy into frequency bands where ambient noise is reduced. Similar pressures operate in human-altered environments. Male grasshoppers (Glyptobothrus maritimus) living in roadside habitats with high traffic noise showed 13.4% higher stridulatory organ density in the first survey year compared with grasshoppers from quiet habitats. This morphological difference, which may favor higher-frequency signal production, became unclear in the following year. Courtship signal frequencies showed no positive relationship with organ density or noise. These findings suggest that elevated sound levels can induce short-term developmental processes that generate variation in sound-producing organs in non-vocal animals (grasshopper stridulatory organs and traffic noise).

Vocal Learning and the Development of Auditory Repertoires

Vocal learning is a complex acquired social behavior found in only a few animal groups. The process requires sensorimotor function: the animal accepts external auditory input, engages in repeated vocal imitation practice, and eventually forms a stable pattern of vocal output. Humans and songbirds share striking similarities in vocal learning behavior, including reliance on auditory feedback, complex syntactic structures, and sensitive periods during development. Both groups have evolved hierarchical forebrain regions related to vocal motor control and vocal learning that are organized and closely associated with the auditory cortex. These parallels make songbirds an ideal animal model for studying the neural mechanisms of vocal learning behavior (analogies of human speech and bird song).

A key feature of vocal ontogeny in taxa with extensive vocal repertoires is a developmental pattern in which vocal exploration is followed by a period of category formation that results in a mature species-specific repertoire. This early vocal development, often called babbling, occurs in vocal-learning birds, some marine mammals, some New World monkeys, some bats, and humans. Notable similarities across these species in the developmental pattern of vocalizations suggest that vocal production learning might require babbling, though the current literature is insufficient to confirm this suggestion. Research directions include expanding descriptive data, conducting quasi-experimental studies to identify mechanisms of acquisition, and using computational modeling to test hypotheses about the origins and functions of babbling (cross-species parallels in babbling).

Song learning differs between the sexes in many species. Birdsong is a culturally transmitted mating signal. Due to historical and geographical biases, song learning has been predominantly studied in temperate zones where female song is rare. Female song is not rare outside temperate zones, and song in both sexes probably represents the ancestral state in songbirds. Some song dimorphisms seen today may be manifestations of secondary losses of female song. The sexes could differ in how well they learn, termed copying fidelity, or from whom they learn, termed model selection. Different learning mechanisms may point toward different selection pressures, so investigating sex-specific learning could help identify the social and ecological pressures contributing to sex differences in adult song (sex differences in bird song learning).

Neural Mechanisms of Auditory Recognition

The auditory system must extract meaningful signals from noisy environments with many co-occurring signallers. Receivers face the challenge of rapidly recognizing salient auditory signals while filtering out irrelevant sounds. Most bird species produce a variety of complex vocalizations that function to communicate with other members of their own species. Behavioral evidence broadly supports preferences for conspecific over heterospecific sounds, a phenomenon called auditory species recognition.

A review of 53 published studies comparing avian neural responses between conspecific and heterospecific vocalizations found that distinct nuclei of the auditory forebrain are consistently conspecific selective across taxa, even in response to unfamiliar individuals with distinct acoustic properties. Species-specific neural discrimination is not a stereotyped auditory response but is modulated according to salience, depending on ontogenetic exposure to conspecific versus heterospecific stimuli. Neuromodulators, particularly norepinephrine, may mediate species recognition by regulating the accuracy of neuronal coding for salient conspecific stimuli. The available data support a perceptual filter-based mechanism in which species identity and social experience combine to influence the neural processing of species-specific auditory signals (neural mechanisms of auditory species recognition in birds).

The auditory system also uses top-down processing to resolve ambiguous signals. Humans use contextual information to infer the meaning of ambiguous acoustic signals. In speech, high-level semantic, syntactic, or lexical information shapes understanding of a phoneme buried in noise. Most current theories rely on hierarchical predictive coding models involving Bayesian priors from high-level brain regions that influence processing at lower levels of the cortical sensory hierarchy. Subcortical auditory nuclei receive massive, heterogeneous, and cascading descending projections at every level of the sensory hierarchy, and activation of these systems has been shown to improve speech recognition. Corticofugal pathways contain the requisite circuitry to implement predictive coding mechanisms for facilitating perception of complex sounds, with top-down modulation at early subcortical stages complementing modulation at later cortical stages (top-down inference in the auditory system).

Auditory Communication for Mating and Sexual Selection

Sexually selected auditory signals provide information that potential mates use to assess quality. Bird song serves as a classic example. The developmental environment can affect the expression of sexually selected traits in adulthood. In a study of male zebra finches, researchers exposed birds to corticosterone treatment during development. After the males reached adulthood, mitochondrial function was quantified from whole red blood cells, and song structure was measured. Corticosterone-treated males had mitochondria that were less efficient and used a lower proportion of maximum capacity compared with control males. These males also had higher baseline corticosterone levels as adults. Developmental treatment had an indirect effect on song peak frequency, with corticosterone-treated males singing songs of higher peak frequency than control males. This effect was modulated through increased corticosterone levels and by a decrease in mitochondrial respiratory capacity. This study provides evidence linking the developmental environment, mitochondrial function, and the expression of a sexually selected trait in bird song (mitochondria and bird song).

The domestic cat offers a different perspective on vocal communication in a species closely associated with humans. Cats vocalize to communicate with other individuals and to express their internal states. The vocal repertoire of the cat is wide, with up to 21 different vocalizations described in the literature, though the repertoire probably contains more types. An ethogram created in one review described the known vocalizations of the domestic cat based on auditory classification. The environment has an important impact on vocal behavior, and feral cats and pet cats vocalize differently. Pet cats are able to create efficient communication with humans thanks to the flexibility of vocalization behaviors (feline vocal communication).

Warning Signals and Territorial Defense

Acoustic signals serve critical survival functions beyond reproduction. Warning calls alert conspecifics to predators, and territorial calls defend resources against rivals. These signals must be detectable in the environment where they are produced and must be distinguishable from background noise.

Anthropogenic noise has increased ambient sound levels across the globe, both underwater and on land. Heightened noise can impair communication in vocal animals through acoustic masking. Noise reduces the animal's communication space, defined as the area in which an individual animal can effectively convey information to a conspecific listener. Previous studies have estimated communication space using sound propagation models and behavioral studies. However, studies frequently equate signal recognition with signal detection, a necessary but not sufficient precondition, thereby persistently overestimating spatial coverage and underestimating anthropogenic impacts.

Deep learning offers an opportunity to estimate biologically relevant communication even for data-limited species. In a case study with the critically endangered North Atlantic right whale, researchers used audio embeddings from the BirdNET model to distinguish individual whales based on their upcalls, a low-frequency contact call produced across ages and sexes that encodes individual identity. The dataset included 234 samples across 11 individuals from 3 sites. Simulating the effect of varying ambient noise levels to estimate signal excess for both signal detection and individual identification revealed that an additional 7 dB or more is necessary for the model to distinguish individuals (deep audio embeddings for whale identification).

Bioacoustics in Animal Health Monitoring

The study of animal communications, termed zoosemiotics, includes the subfield of bioacoustics, the study of the production, transmission, and reception of animal sounds. Inter- and intra-species communication is sophisticated, with sound playing a major role in signaling. Artificial intelligence-led research can be employed to understand and combine recorded multi-level data, including sound, vision, and odors, to classify animal health and identify interventions. This can include subgroup discovery and trajectory analysis as essential elements in developing animal-specific identification of failure to thrive or ill health. It is important that animals, carers, and veterinarians receive as early a diagnosis as possible to predict trajectory and plan care needs and interventions. However, the use of quantitative data for evidence-led interventions based on sound has not yet been developed. Advances in bioacoustics provide a framework to determine where early diagnosis and animal health improvements can be made through understanding of behavior and oral sound production (AI in bioacoustics for animal health monitoring).

Practical Methods for Studying Auditory Communication

Recording and Analysis Workflow

Studying auditory communication requires systematic methods for capturing, processing, and interpreting acoustic data. A practical workflow includes the following steps.

First, select recording equipment appropriate for the target species and frequency range. Species that communicate in ultrasound, such as mice, require specialized microphones capable of detecting frequencies above 20 kHz. Species with low-frequency calls, such as North Atlantic right whales, require hydrophones with appropriate low-frequency response.

Second, record synchronized audio and video when possible. Mice communicate using ultrasonic vocalizations during social interactions, but these vocalizations are not associated with clear visual indicators and occur at frequencies outside human hearing. A protocol for recording synchronized video and ultrasonic audio data during multi-animal social interactions enables simultaneous capture of behavior and vocal activity. A computational pipeline integrates multi-animal tracking, vocalization detection, sound-source localization, and assignment of vocalizations to individual animals. Validation procedures assess tracking accuracy, vocalization extraction, localization precision, and assignment confidence at each stage. Application of this protocol to a four-animal demonstration dataset yields individual-resolved vocalization tracking, spatially precise sound-source estimates, and quantitative measures of vocal output and acoustic features (assigning ultrasonic vocalizations to individual mice).

Third, process recordings to extract relevant acoustic features. Bird song classification has benefited from advanced signal processing methods. One approach uses a feature extraction method based on the Wigner-Ville ambiguity function cross-terms, which is useful for classification of non-stationary multi-component signals with stochastic variation in amplitudes and time-frequency locations. This method gave better classification than established methods when evaluated on simulated data and bird song syllables of the great reed warbler (bird song syllable classification).

Fourth, validate classification results. Open audio databases such as Xeno-Canto are widely used to build datasets for exploring bird song repertoire or training models for automatic bird sound classification. These databases suffer from weak labeling: a species name is attributed to each audio recording without timestamps providing temporal localization of the bird song of interest. A data-centric labeling function composed of three steps, including time-frequency sound unit segmentation, feature computation, and classification of each sound unit as bird song or noise, reduced label noise by up to a factor of three in a study of 44 West-Palearctic common bird species (unsupervised classification for bird song datasets).

Machine Learning Classification of Bird Song

Deep learning models have achieved high accuracy in bird song classification. A model combining a bi-directional long short-term memory neural network and a dense convolutional network was proposed for bird song classification. The workflow involved classifying, filtering, and extracting features such as Mel frequency cepstrum coefficients, building the network, using a cross-entropy loss function to tune the network structure, and using a softmax classifier to classify 20 bird species. Experimental analysis of 14,311 audio files showed average accuracy between 90% and 93% for all bird species detection, with species-specific accuracy of 91.1% for hawks, 92.7% for western ruffed grouse, 91.4% for crested wheatears, and 92.6% for red-throated divers (bird song classification with Bi-LSTM-DenseNet).

Common Failure Patterns in Auditory Communication Research

Several recurring problems compromise studies of auditory communication. Recognizing these patterns helps researchers design more robust investigations.

The first failure pattern is equating signal detection with signal recognition. Studies frequently assume that if a receiver can detect a signal, communication has occurred. This assumption overestimates spatial coverage and underestimates anthropogenic impacts because recognition requires additional processing beyond detection. The North Atlantic right whale study demonstrated that an additional 7 dB or more of signal excess is needed for individual identification compared with simple detection (deep audio embeddings for whale identification).

The second failure pattern is ignoring morphological plasticity in sound-producing organs. Studies of anthropogenic noise effects on communication often attribute signal changes to behavioral plasticity in vocal species. The grasshopper study showed that non-vocal sound producers can exhibit morphological changes in response to noise, with stridulatory organ density varying across habitats with different traffic noise levels. Neglecting morphological plasticity constrains understanding of ecological noise impacts (grasshopper stridulatory organs and traffic noise).

The third failure pattern is relying on weakly labeled datasets without validation. Open audio databases contain label noise that can compromise model training and evaluation. Segmentation of bird songs alone aggregated from 10% to 83% of label noise depending on the species. Labeling functions that automatically segment audio recordings before assigning labels can significantly reduce this noise (unsupervised classification for bird song datasets).

The fourth failure pattern is assuming that vocal repertoires are fixed and context-independent. The cat vocal repertoire illustrates this problem. It is sometimes unclear if different types of vocalizations are produced in different environments or if a unique type of vocalization is used with variation in acoustic parameters. Isolation calls produced by kittens differ depending on context, and feral cats and pet cats vocalize differently (feline vocal communication).

Limitations and Open Questions

Current knowledge of auditory communication has significant gaps. The relationship between babbling and vocal production learning remains uncertain. The current state of the literature is insufficient to confirm that vocal production learning requires babbling, despite notable similarities across species in developmental patterns of vocalization (cross-species parallels in babbling).

The neural basis of auditory species recognition is not fully understood. While distinct nuclei of the auditory forebrain are consistently conspecific selective across taxa, species-specific neural discrimination is modulated according to salience and ontogenetic exposure. The role of neuromodulators such as norepinephrine in mediating species recognition requires further investigation (neural mechanisms of auditory species recognition in birds).

The mechanisms underlying auditory processing deficits in neurodevelopmental conditions remain an active area of research. Autism spectrum disorders are strongly associated with auditory hypersensitivity or hyperacusis, defined as difficulty tolerating sounds. Fragile X syndrome, the most common monogenetic cause of autism spectrum disorder, has emerged as a gateway for exploring underlying mechanisms of hyperacusis and auditory dysfunction. Disruptions occur at molecular, synaptic, and circuit levels, from aberrant synaptic development and ion channel deregulation of auditory brainstem circuits to impaired neuronal plasticity and network hyperexcitability in the auditory cortex (auditory processing deficits in Fragile X syndrome).

Speaking-induced suppression, the phenomenon where sounds generated by overt speech elicit smaller neurophysiological responses in the auditory cortex than comparable externally generated sounds, illustrates the complexity of auditory self-monitoring. Subnormal levels of speaking-induced suppression have been reported in patients with schizophrenia, providing a plausible explanation for some first-rank symptoms. Failure to suppress the neural consequences of self-generated inner speech may explain certain classes of auditory-verbal hallucinations (speaking-induced suppression of the auditory cortex).

Welfare and Safety Context

Understanding auditory communication has direct welfare implications for animals under human care. Recognizing the vocal signals associated with pain, distress, or illness enables earlier intervention. The framework for using bioacoustics in animal health monitoring emphasizes early diagnosis to predict trajectory and plan care needs. Animals, carers, and veterinarians benefit from quantitative data that supports evidence-led interventions based on sound (AI in bioacoustics for animal health monitoring).

Anthropogenic noise poses welfare concerns beyond acoustic masking. Non-auditory effects of anthropogenic noise on animals include physiological stress responses and behavioral changes that extend beyond the auditory system (non-auditory effects of anthropogenic noise). Habitat management should consider noise reduction as a component of animal welfare.

Auditory hypersensitivity has clinical relevance for both animals and humans. Pharmacological approaches to tinnitus and hyperacusis remain peripheral within research despite persistent patient demand and advances in auditory neuroscience. Converging mechanistic explanations, biological heterogeneity, and emerging therapeutic targets are under discussion. Evolving mechanistic frameworks can support patient education and clinical communication even in the absence of disease-modifying treatments (pharmacological approaches to tinnitus and hyperacusis).

Professional Escalation Criteria

Researchers and practitioners working with auditory communication should escalate to specialized expertise under specific circumstances.

Escalate to a veterinarian when an animal under professional care shows sudden changes in vocal behavior that persist beyond 24 hours, particularly when accompanied by reduced appetite, lethargy, or other signs of illness. Vocal changes can indicate pain, respiratory disease, or neurological conditions.

Escalate to an acoustics engineer or bioacoustics specialist when recording equipment fails to capture target signals or when signal processing methods produce inconsistent classification results. Specialized expertise may be needed for species with ultrasonic communication or for recordings in high-noise environments.

Escalate to a statistician or machine learning specialist when classification models show accuracy below acceptable thresholds or when validation procedures reveal high label noise. The Bi-LSTM-DenseNet model achieved accuracy between 90% and 93% for bird species detection, and models performing substantially below this range may require architectural or data quality improvements (bird song classification with Bi-LSTM-DenseNet).

Escalate to a conservation authority when anthropogenic noise is suspected of impairing communication in threatened or endangered species. The North Atlantic right whale case study demonstrates how noise reduces communication space and can impair individual identification, with direct implications for conservation management (deep audio embeddings for whale identification).

Frequently Asked Questions

How do animals recognize the calls of their own species?

Animals use auditory species recognition to distinguish conspecific from heterospecific sounds. Distinct nuclei of the auditory forebrain are consistently conspecific selective across bird taxa, even in response to unfamiliar individuals with distinct acoustic properties. This neural discrimination is modulated by ontogenetic exposure and social experience, supporting a perceptual filter-based mechanism where species identity and social experience combine to influence neural processing of species-specific auditory signals (neural mechanisms of auditory species recognition in birds).

What is the difference between signal detection and signal recognition in animal communication?

Signal detection is the ability to perceive that a sound is present. Signal recognition is the ability to extract meaningful information from that sound, such as individual identity or behavioral context. Studies frequently equate these two processes, which overestimates spatial coverage and underestimates anthropogenic impacts. In North Atlantic right whales, an additional 7 dB or more of signal excess is necessary for individual identification compared with simple detection (deep audio embeddings for whale identification).

Can non-mammalian species communicate using ultrasound?

Yes. The concave-eared torrent frog from China produces and detects ultrasonic frequencies above 20 kHz, making it the first known amphibian with ultrasonic communication. Both audible and ultrasonic components of the male advertisement call evoke vocal responses from other males. This capacity likely co-evolved in response to intense low-frequency ambient noise from fast-flowing streams (ultrasonic communication in frogs).

How does anthropogenic noise affect animal communication?

Anthropogenic noise increases ambient sound levels and can impair communication through acoustic masking, which reduces the communication space available to an animal. Noise can also induce morphological changes in sound-producing organs, as demonstrated in grasshoppers living near roads with high traffic noise. These effects can be short-term and variable across years (grasshopper stridulatory organs and traffic noise).

What is babbling in animal vocal development?

Babbling is a developmental pattern in which vocal exploration is followed by a period of category formation that results in a mature species-specific repertoire. It occurs in vocal-learning birds, some marine mammals, some New World monkeys, some bats, and humans. Notable similarities across these species suggest that vocal production learning might require babbling, though current evidence is insufficient to confirm this (cross-species parallels in babbling).

How many vocalizations can domestic cats produce?

The vocal repertoire of the domestic cat is wide, with up to 21 different vocalizations described in the literature. The repertoire probably contains more types. Cats vocalize to communicate with other individuals and to express internal states. The environment has an important impact on vocal behavior, and feral cats and pet cats vocalize differently (feline vocal communication).

How is artificial intelligence used to study animal sounds?

Artificial intelligence is used to classify animal sounds, identify individual animals, and monitor health. Deep learning models such as Bi-LSTM-DenseNet achieve 90% to 93% accuracy in bird species classification. Audio embeddings from models like BirdNET can distinguish individual North Atlantic right whales. AI can also combine recorded multi-level data including sound, vision, and odors to classify animal health and identify interventions (AI in bioacoustics for animal health monitoring).

What is the relationship between bird song and human speech?

Humans and songbirds share striking similarities in vocal learning behavior, including reliance on auditory feedback, complex syntactic structures, and sensitive periods during development. Both groups have evolved hierarchical forebrain regions related to vocal motor control and vocal learning. These parallels make songbirds an ideal animal model for studying the neural mechanisms of vocal learning and may provide insights for treating language disorders (analogies of human speech and bird song).

Related Articles

References and Further Reading

This article is educational and does not replace institutional policy, professional advice, or applicable safety and regulatory requirements.