A year ago, my student Riya Jain and I analysed thirty audio recordings of people telling us the worst day of their lives. Half our participants had their hands taped loosely to their sides; the rest were free to move as they spoke. We weren’t listening to what they said. We were timing the silences.
Across 4,242 individual pauses, a pattern emerged. Long silences; those stretching past a second, tied to searching for a feeling rather than a fact: grew more frequent when people recounted painful memories. When we prevented people from gesturing, those long pauses grew more frequent still, in both the emotional and the neutral stories. Stop the hands, and the mind works harder to find words. That study, just published in Acta Psychologica, sits alongside a decade of work I’ve done on how meaning gets built out of language, image, gestures and translation, and all of it points to the same conclusion: no single channel carries meaning on its own. Words, pictures, and gestures are jointly constructing something none of them could produce alone.
“The next time you watch someone search for the right word, watch their hands too. They mayalready be saying it.”
What words actually do
We tend to imagine communication as parcel delivery: a thought wrapped in words, handed over, unwrapped intact. My own research keeps running into evidence against this. In an fMRI study with Prof. Bipin Indurkhya (Poland) and Prof. Elisabetta Gola (Italy), we showed participants uncaptioned pairs of images meant to be read metaphorically, with no language involved at all. If metaphor were a purely linguistic operation, dressed up in pictures, we shouldn’t have seen much happening in the brain’s language centres. Instead, interpreting the visual metaphors reliably activated the inferior frontal gyrus (BA47), a region central to language processing. Language switches on even when there isn’t a word in sight. That tells us something important: “verbal” and “visual” are not two separate meaning-making systems that occasionally borrow from each other. They are two entry points into one shared conceptual apparatus. This matters more than ever now that machines are doing a growing share of our reading and writing for us. In work with Dr. Barnali Chaudhary on philosophical translation in the age of large language models, we argued that AI translation tools produce fluent, plausible sentences without anything resembling the interpretive judgment a human translator brings to a conceptually densetext, the decision about which of several possible senses an author intended, given the argument three pages earlier and the tradition it belongs to. We called the resulting redistribution of responsibility “distributed translational agency”: the human translator shifts from primary author to epistemic supervisor, checking outputs whose fluency can mask an absence of genuine understanding. We tested that suspicion more directly in a follow-up study with Dr. Barnali Chaudhary, comparing how humans and large language models interpret metaphors. We scored interpretations along two dimensions: how embodied they were (grounded in bodily, sensory experience) and how culturally accurate, with a third measure of phenomenological depth as the outcome. Human interpretations were well predicted by both embodiment and cultural grounding. For the models, both predictorsexplained far less of the variation, and yet the models expressed more confidence in their answers than the humans did. What they produced looked like understanding and wasn’t grounded in much of anything: surface-coherent but phenomenologically thin, metaphor as statistical pattern-matching rather than something built from a body that has actually felt sharpness, warmth, or weight.
What images do that sentences cannot
Pictures are not simply words in disguise. They have their own grammar, one I’ve spent years trying to map. In an early study, I paired a photograph of the Taj Mahal with one of wine bottles, conceptually unrelated, but the marble minarets and the slender bottle necks share a shape. Using eye-tracking, we found that people register this kind of low-level perceptual similarity: in shape, colour, texture, before they consciously register any conceptual connection between the two images. That subconscious pull is often what starts a metaphorical reading in the first place; perceptual similarity is not decoration on top of the metaphor, it’s frequently the mechanism that gets it started.
How the image is composed then changes how strongly the metaphor lands. In another study, I compared pictorial similes – two images simply placed side by side, implying “this is like that” -against hybrid pictorial metaphors, where the two concepts are fused into a single object: a man’s head merged with a chicken’s, rather than a man drawn beside one. The fused, hybrid versions were felt more strongly than the side-by-side similes, but only when viewers encountered them cold, without a caption spelling out the comparison. Juxtaposition merely suggests a search for similarity.
Fusion forces the resolution: it transforms how you see one thing by making you look at it through another, within a single depicted space.
Even the smallest pictorial marks turn out to carry systematic meaning. In research on the flourishes cartoonists draw around characters’ heads: what comics scholars call “pictorial runes” or “emanata”, we found that droplets and spikes read as generic emotional arousal, spirals read specifically as negative emotion, and twirls read as confusion or dizziness. No one teaches readers this vocabulary explicitly. It’s absorbed the way a child absorbs grammar, through exposure, and it does real interpretive work that no caption is doing for the reader.
The hands are thinking too
Which returns me to the tape on my participants’ wrists. Gesture is often treated as decoration: hand-waving that accompanies the real communicative work happening in speech. Our data point the other way. When we physically restricted gesture, we weren’t suppressing a communicative flourish; we were removing a channel that appears to help speakers organize memory, regulate emotion, and search for words, especially under emotional load. The long pauses that grew when hands were tied down look like the audible residue of a cognitive system working without one of its usual tools.
Meaning as a shared construction
What ties an fMRI scanner, an eye-tracker, and a stopwatch on silent pauses together is a single claim: meaning-making is distributed across modalities that we’ve studied in isolation largely for disciplinary reasons, not because the mind treats them separately. A word activates conceptual machinery that overlaps substantially with what a picture activates. A picture fuses two concepts into one depicted space and encodes emotion in the curl of a line, things a sentence structurally cannot do. A gesture carries part of the cognitive load of finding the words in the first place. None of these channels is a mere illustration of what’s “really” happening in language.
This has stakes well beyond the seminar room. As large language models increasingly mediate howwe write, translate, and converse, our own data suggest real caution about what fluency is tellingus. A model can produce a metaphor interpretation, or a translated sentence, that reads as confidentand coherent while remaining, by the measures we used, considerably thinner than a human one:ungrounded in the embodied and cultural experience that gives our own language its depth. If wewant machines, classrooms, or communication design to actually traffic in meaning rather thanmerely simulate its surface, we need to keep building the fuller picture: not just what words say,but what images do and what hands are doing while the words are still being found.















