A mouth click is over almost as soon as it begins. To most of us, it’s only a small, sharp sound. To an expert blind echolocator, the returning echoes can disclose the edge of a doorway, the path through a room, or even something about the material of a surface.

The precision of this ability is startling. What makes it scientifically interesting, though, isn’t just that humans can do it. It’s that evolution doesn’t appear to have built us specifically to do it.

Bats bear the marks of a long evolutionary history with echoes. Their ears and auditory systems are specialized for ultrasonic navigation, and young bats begin producing useful calls without having to be taught. Humans don’t. In fact, the number of human echolocators in history is so small that it can’t have had any evolutionary impact. Even so, blind adults often learn to echolocate with chiropteran skill and precision.

Given this, human echolocation can be seen as a warning about evolutionary inference. A capacity can be intricate, difficult, and astonishingly effective without having been directly designed by natural selection for its present use. Complexity is a fact about what an organism can do. It isn’t, by itself, a history of how the ability arose.

That warning matters for one of the oldest questions about language.

The wrong lesson from complexity

Human language is extraordinarily complex. Children learn to divide a continuous stream into words, connect those words to people and events, and discover patterns that let them understand sentences they have never heard before. No other species develops an open-ended linguistic system in the same way.

One tempting conclusion is that natural selection must have equipped us with specialized machinery for grammar. The reasoning resembles an argument sometimes made from bat echolocation: a complex, species-typical ability looks engineered, so it must be a direct adaptation for that function.

Human echolocation breaks the bridge in that argument. It shows that general neural systems can contain powerful capacities that remain hidden until experience gives them the right work to do. That doesn’t prove that language lacks specialized adaptations. Genetic, anatomical, and developmental evidence might still support them. But linguistic complexity alone can’t do so.

The more revealing puzzle isn’t how a general learner could ever acquire language. It’s why virtually every typically developing human child does, while echolocation and other remarkable human abilities remain rare.

The answer may lie less in what the brain can compute than in what it treats as worth learning from.

Not a spotlight, but a trust policy

A developing brain encounters far more information than it can learn from equally. Voices overlap. Objects move. Faces turn. Rhythms appear in speech, footsteps, music, and machinery. Some events are directed at the child; most aren’t.

Learning therefore depends on an admission policy. When the world violates an expectation, how much should the learner update? In predictive-processing terms, the answer depends on precision: an estimate of how reliable a particular error is. A high-precision mismatch deserves a substantial update. A low-precision mismatch can be dismissed as noise.

The attentional-bias hypothesis proposes that human infants inherit a structured policy for assigning that trust. The policy gives unusual weight to a particular kind of event: communication that is directed at the infant, organized rhythmically across several timescales, and bound to a shared scene.

The proposed high-precision event
Addressedgaze · turn-taking · prosody
Rhythmicpatterns nested across timescales
Situatedform · agent · action · context
Bound eventevidence worth updating on

No cue is sufficient by itself. The claim concerns the way the learner binds them into a single consequential episode.

The binding is essential. A caregiver doesn’t present a child with an isolated sound wave. They look toward the child, change their prosody, gesture, act on something in view, and wait for a response. In signed communication, visual form, bodily rhythm, gaze, and the scene can play the corresponding roles. The proposal is that the infant treats these features together as especially reliable evidence.

This isn’t a claim that infants possess grammar in advance. Nor does it claim that they’re simply more attentive than other animals. The inherited contribution is narrower: a bias concerning which prediction errors matter, at which levels, and across which relations. Culture supplies the words, constructions, and communicative practices. General learning mechanisms extract the regularities. The bias keeps steering those mechanisms toward the episodes in which linguistic structure is most densely connected to the world.

Over thousands of such encounters, a small difference in weighting could create a very different developmental path. The learner doesn’t receive a new processor. It receives a different diet of consequential evidence.

Attention is not a volume knob

It would be easy to weaken this idea into the slogan that more attention produces more learning. That’s not the proposal, and it’s probably false.

Turning up every signal can make irrelevant regularities look important. Clear sounds can still be overheard. Strong rhythms can come from events that communicate nothing. If the learner treats all of them as equally trustworthy, it may acquire false boundaries and misleading patterns.

The paper’s own stripped-down computer exercise demonstrates the danger. In a simple word-segmentation procedure, multiplying the weight on selected tokens didn’t improve performance. Segmentation became slightly worse, and false boundaries increased. The exercise is too crude to test the full hypothesis: it uses a single scalar boost where the theory requires a structured profile binding form, rhythm, addressee, and scene. But its failure is instructive. There’s no magic attention dial that can simply be turned upward.

The relevant evolutionary change would have been closer to a policy than an amplifier. It would specify when gain rises, which relationships receive it, and which learning systems are allowed to update.

Other animals notice different things

The distinction helps make sense of how other animals learn. Zebra finches learn much more effectively from live tutors than from recordings. Dogs pay attention to human gaze and pointing; wolves don’t.

What differs may be the target of the amplification.

A songbird can assign high value to matching a tutor’s song. A dog can treat a human gaze or pointing gesture as a privileged clue. Those gain profiles support impressive achievements, but they aren’t necessarily the human profile.

The proposed human bias targets a larger bound event: this form, from this agent, addressed to me, at this moment, bearing on this shared situation. It treats sound or sign, temporal structure, gaze, action, and context as mutually constraining evidence. Acoustic imitation can be part of that event without being equivalent to it.

This is also why caregiver behaviour matters. Human infants receive vastly more directed vocal communication than infant apes do. The hypothesis doesn’t replace that cultural environment; it makes the environment more consequential. Caregivers supply unusually rich addressed input, and infants are unusually disposed to learn from it. Each side may have helped the other evolve.

Four cradles

A useful hypothesis should say what would prove it wrong. The most direct test begins not with children who already understand words, but with newborns.

Imagine four cradles, each representing one of four carefully matched events. One is both rhythmically organized and ostensively directed at the child. A second is directed at the child but has its multi-timescale rhythm flattened. A third preserves the rhythm but removes the directed communicative framing. A fourth has neither feature.

The binding account predicts an interaction, not a general rise
low rhythm · low ostension Baseline
high rhythm · low ostension Rhythm alone
low rhythm · high ostension Ostension alone
high rhythm · high ostension Predicted premium

A rhythm-first account, a social-pragmatic account, general arousal, and the null each predict a different arrangement of the four cells.

The binding account predicts an interaction: the event with both ostension and rhythm should elicit a response premium beyond the separate contribution of either cue. A rhythm-first theory predicts that rhythmic events should dominate whether or not they are addressed to the infant. A social-pragmatic theory predicts the reverse. A general-arousal account predicts that the contrasts should largely disappear once intensity and excitement are matched. A reliable null result would weaken the proposed birth-onset bias.

Comparable studies in other species would test the claim of human specificity. If a suitable non-human comparison species showed the same interaction with species-appropriate signals, the proposed human difference would be undercut.

Two further tests approach the hypothesis from different directions. In a computational replay, two capacity-matched learners would receive exactly the same communicative stream. Only the precision-allocation policy would vary. The relevant outcomes would be speed of learning, sample efficiency, and transfer to new combinations—not mere imitation. A neonatal brain-imaging study could ask whether responses in temporal cortex follow the predicted ordering for infant-directed, adult-directed, and acoustically scrambled speech.

What such results could show

That newborns weight a bound communicative event differently, and that a structured precision policy can alter learning trajectories when capacity and input are held fixed.

What they could not show

That a newborn understands intentions, that preference already constitutes language, or that every account involving specialized circuitry has been eliminated.

None of these findings would, by itself, demonstrate that a newborn recognizes communicative intentions or guarantee later language. A preference is a measure of early weighting, not a miniature grammar. Even a successful set of results would support the attentional account without eliminating every theory involving specialized circuitry.

That caution is part of the proposal. The experiments haven’t yet established the mechanism. They’re what could support, constrain, or refute it.

A small evolutionary change with a long reach

The attention-bias hypothesis occupies a middle ground in the language-evolution debate. It doesn’t treat the infant mind as a blank slate. Something important is inherited. But what is inherited doesn’t need to be a store of grammatical knowledge or a processor built exclusively for syntax.

It may instead be a prejudice about evidence.

That prejudice would be modest in its genetic specification and enormous in its consequences. It would draw infants toward the structured episodes that caregivers and communities continually provide. It would also shape production: once a communicative target receives high precision, the mismatch between an infant’s own babble and the attended form becomes worth correcting. Across development, learning and action would keep converging on the same cultural system.

Human echolocation reveals how much general neural machinery can do when experience persistently directs it toward the right cues. Language may reveal the complementary evolutionary story. Rather than constructing a new machine for grammar, natural selection may have changed which moments the old machinery could not ignore.

Selection, on this account, didn’t write language into the brain. It made us notice where language was.

This is a plain-language companion to Brett Reynolds, “An attentional-bias account of language emergence,” Evolutionary Linguistic Theory (2026). Read the published article or the public manuscript.