How We Create Meaning Through Words, Images, and Gestures

Published on
August 20, 2026

Department of Humanities and Social Sciences, Indian Institute of Technology Jammu, India

Areas of Expertise
Cognitive Science & Creativity, Multimodal Communication, Neurocognitive Research, Decision-Making

A year ago, my student Riya Jain and I analysed thirty audio recordings of people telling us the worst day of their lives. Half our participants had their hands taped loosely to their sides; the rest were free to move as they spoke. We weren’t listening to what they said. We were timing the silences.

Across 4,242 individual pauses, a pattern emerged. Long silences; those stretching past a second, tied to searching for a feeling rather than a fact: grew more frequent when people recounted painful memories. When we prevented people from gesturing, those long pauses grew more frequent still, in both the emotional and the neutral stories. Stop the hands, and the mind works harder to find words. That study, just published in Acta Psychologica, sits alongside a decade of work I’ve done on how meaning gets built out of language, image, gestures and translation, and all of it points to the same conclusion: no single channel carries meaning on its own. Words, pictures, and gestures are jointly constructing something none of them could produce alone.

We tend to imagine communication as parcel delivery: a thought wrapped in words, handed over, unwrapped intact. My own research keeps running into evidence against this. In an fMRI study with Prof. Bipin Indurkhya (Poland) and Prof. Elisabetta Gola (Italy), we showed participants uncaptioned pairs of images meant to be read metaphorically, with no language involved at all. If metaphor were a purely linguistic operation, dressed up in pictures, we shouldn’t have seen much happening in the brain’s language centres. Instead, interpreting the visual metaphors reliably activated the inferior frontal gyrus (BA47), a region central to language processing. Language switches on even when there isn’t a word in sight. That tells us something important: “verbal” and “visual” are not two separate meaning-making systems that occasionally borrow from each other. They are two entry points into one shared conceptual apparatus. This matters more than ever now that machines are doing a growing share of our reading and writing for us. In work with Dr. Barnali Chaudhary on philosophical translation in the age of large language models, we argued that AI translation tools produce fluent, plausible sentences without anything resembling the interpretive judgment a human translator brings to a conceptually densetext, the decision about which of several possible senses an author intended, given the argument three pages earlier and the tradition it belongs to. We called the resulting redistribution of responsibility “distributed translational agency”: the human translator shifts from primary author to epistemic supervisor, checking outputs whose fluency can mask an absence of genuine understanding. We tested that suspicion more directly in a follow-up study with Dr. Barnali Chaudhary, comparing how humans and large language models interpret metaphors. We scored interpretations along two dimensions: how embodied they were (grounded in bodily, sensory experience) and how culturally accurate, with a third measure of phenomenological depth as the outcome. Human interpretations were well predicted by both embodiment and cultural grounding. For the models, both predictorsexplained far less of the variation, and yet the models expressed more confidence in their answers than the humans did. What they produced looked like understanding and wasn’t grounded in much of anything: surface-coherent but phenomenologically thin, metaphor as statistical pattern-matching rather than something built from a body that has actually felt sharpness, warmth, or weight.

Pictures are not simply words in disguise. They have their own grammar, one I’ve spent years trying to map. In an early study, I paired a photograph of the Taj Mahal with one of wine bottles, conceptually unrelated, but the marble minarets and the slender bottle necks share a shape. Using eye-tracking, we found that people register this kind of low-level perceptual similarity: in shape, colour, texture, before they consciously register any conceptual connection between the two images. That subconscious pull is often what starts a metaphorical reading in the first place; perceptual similarity is not decoration on top of the metaphor, it’s frequently the mechanism that gets it started.

How the image is composed then changes how strongly the metaphor lands. In another study, I compared pictorial similes – two images simply placed side by side, implying “this is like that” -against hybrid pictorial metaphors, where the two concepts are fused into a single object: a man’s head merged with a chicken’s, rather than a man drawn beside one. The fused, hybrid versions were felt more strongly than the side-by-side similes, but only when viewers encountered them cold, without a caption spelling out the comparison. Juxtaposition merely suggests a search for similarity.

Fusion forces the resolution: it transforms how you see one thing by making you look at it through another, within a single depicted space.

Even the smallest pictorial marks turn out to carry systematic meaning. In research on the flourishes cartoonists draw around characters’ heads: what comics scholars call “pictorial runes” or “emanata”, we found that droplets and spikes read as generic emotional arousal, spirals read specifically as negative emotion, and twirls read as confusion or dizziness. No one teaches readers this vocabulary explicitly. It’s absorbed the way a child absorbs grammar, through exposure, and it does real interpretive work that no caption is doing for the reader.

Which returns me to the tape on my participants’ wrists. Gesture is often treated as decoration: hand-waving that accompanies the real communicative work happening in speech. Our data point the other way. When we physically restricted gesture, we weren’t suppressing a communicative flourish; we were removing a channel that appears to help speakers organize memory, regulate emotion, and search for words, especially under emotional load. The long pauses that grew when hands were tied down look like the audible residue of a cognitive system working without one of its usual tools.

What ties an fMRI scanner, an eye-tracker, and a stopwatch on silent pauses together is a single claim: meaning-making is distributed across modalities that we’ve studied in isolation largely for disciplinary reasons, not because the mind treats them separately. A word activates conceptual machinery that overlaps substantially with what a picture activates. A picture fuses two concepts into one depicted space and encodes emotion in the curl of a line, things a sentence structurally cannot do. A gesture carries part of the cognitive load of finding the words in the first place. None of these channels is a mere illustration of what’s “really” happening in language.

This has stakes well beyond the seminar room. As large language models increasingly mediate howwe write, translate, and converse, our own data suggest real caution about what fluency is tellingus. A model can produce a metaphor interpretation, or a translated sentence, that reads as confidentand coherent while remaining, by the measures we used, considerably thinner than a human one:ungrounded in the embodied and cultural experience that gives our own language its depth. If wewant machines, classrooms, or communication design to actually traffic in meaning rather thanmerely simulate its surface, we need to keep building the fuller picture: not just what words say,but what images do and what hands are doing while the words are still being found.

References

Jain R, Ojha A. Gesture inhibition and emotional content selectively modulate long silent pauses in autobiographical narratives. Acta Psychologica. 2026 Aug 1;268:107287.
Article DOI

Science Factors.

The molecular machinery that helps plants survive heat

0
When we think about the impact of climate change on agriculture, our minds often turn to parched soils and empty reservoirs. However, beneath the...

Threads in the Sand: What the Genomes of the Thar Desert Tell Us About India

0
While monuments crumble and written histories fade, human DNA preserves stories that stretch to millennia. India, home to nearly 1.48 billion people has one...

Electrochemical Divergent Synthesis of Bioactive Heterocycles by Regulating the Applied Electricity

0
N-Heterocycles and the Rise of Sustainable Synthesis Nitrogen-containing heterocycles (N-heterocycles) serve as the structural backbone for the majority of modern pharmaceuticals. Recent data indicate that...

Hidden Chemistry Above the Oceans: How Reactive Molecules Shape Climate and Aerosol Formation

0
Marine Aerosols and the Climate Connection Aerosols, the tiny solid particles or liquid droplets suspended in air, affect the Earth's climate in various ways and...

The Sperm Selection Puzzle: How Microfluidics Could Recreate Nature

0
From an extraordinary journey inside the female reproductive tract to tiny engineered channels, scientists are exploring whether the principles of natural sperm selection can...

The New Face of Water Pollution: Beyond Microplastics

0
The Changing Face of Water Pollution Twenty years ago, water pollution was something we could usually see. Oil slicks floating on rivers, untreated sewage entering...

Evolving Landscape and Clinical Potential of Robotic Telesurgery

0
Robotic Telesurgery: From Concept to Clinical Reality In the last 4 decades surgery has evolved from open techniques to minimally invasive robotic assisted approaches with...

Who Decides What Is Beautiful? The Changing Science and Culture of Female Beauty

0
‘Female beauty’is always seen to improve about the age of puberty; but, if we should attempt to define in what this beauty consists or...