Why Do Only Humans Speak? The Most Important Moment in Human History | Documentary For Sleep

Why Do Only Humans Speak? The Most Important Moment in Human History | Documentary For Sleep

Somewhere between 200,000 and 2 million years ago, science still cannot narrow that range, something happened in Africa. A creature that looked almost like us opened its mouth and produced not a cry, not a howl, not a growl, but something fundamentally different. A sound that referred to something [music] specific.

A sound that another creature understood not because it resembled the thing it described, [music] but because both creatures had agreed it meant [music] that. The first word. Nobody knows what it was. Nobody knows who said it. And nobody knows exactly when it happened. But that event changed [music] this planet more profoundly than any geological event in the last hundred million years.

Today, we go looking for it. Before we continue, please subscribe to this channel. Your support makes this work possible, and there is much more to explore. Thank you. Let us begin with the body. Because the story of how speech emerged is, before it is anything else, a story about anatomy. And the anatomy of human speech is genuinely strange.

Strange in a way that becomes obvious the moment you start comparing us to the animals most similar to us. A chimpanzee is, by [music] any reasonable measure, an intelligent animal. Chimpanzees use tools. They solve complex problems. They recognize themselves in mirrors. They form political alliances, deceive rivals, grieve their dead.

Their cognitive abilities in domain after domain are far beyond what most people expect when they encounter them for the first time. And their brains are, in terms of basic architecture, remarkably similar to ours. The same [music] regions, the same connectivity, scaled somewhat differently, but fundamentally recognizable.

If you were designing an experiment to test the hypothesis that intelligence alone is what enables language, the chimpanzee would seem like an ideal candidate. But, chimpanzees do not speak. They never have. Not for lack of trying. Researchers have spent decades [music] attempting to teach chimpanzees some form of human language, with results that are interesting, but ultimately limited.

Chimpanzees can learn to use sign language gestures or press symbols on keyboards to communicate with surprising complexity. Kanzi, a bonobo who has been studied extensively at the Great Ape Trust, understands hundreds of spoken English words and can respond to novel sentences he has never encountered before.

The cognitive capacity for linguistic-like communication is clearly present to some degree in our closest relatives. But, they cannot speak. They cannot produce the sounds of human speech. And the reason is not in their brains. It is in their throats. The human vocal tract is anatomically unique among primates in a specific and consequential way.

In most mammals, [music] including chimpanzees, gorillas, dogs, cats, horses, the larynx sits high in the throat, close to the back of the mouth. This high position means that the airway and the food passage are well separated. When a chimpanzee swallows, food goes one way and air goes [clears throat] another, and the risk of food entering the airway is minimal.

The high larynx also means that chimpanzees can breathe and drink simultaneously, a feat that human infants can perform for the first few months of life before the larynx begins [music] its descent. High larynx position is, in many ways, the safer and more mechanically sound design. It keeps food and air where they belong.

In adult humans, the larynx sits much lower in the throat, lower than in any other primate, lower than in almost any other mammal. This descended larynx creates the long pharyngeal space above the vocal cords that is essential for producing the full range of human speech [music] sounds. The pharynx acts as a resonating chamber, and its length and shape allow the subtle modulations that distinguish vowels from each other, and that give human speech its acoustic richness.

>> Without the descended larynx, the repertoire of sounds an organism can produce is dramatically limited. A chimpanzee’s high larynx simply cannot produce the acoustic variety that human language requires. But, the descended larynx comes with a cost, a serious one. Because when the larynx moves down, the separation between the food and air passages that evolution had carefully maintained for hundreds of millions of years is compromised.

In humans, food and air share the pharynx before diverging. Food going forward into the esophagus, air going backward into the trachea. This arrangement requires precise neuromuscular coordination every time we swallow. When that coordination fails, when a piece of food takes the wrong path, the result can be choking.

And choking can be fatal. Humans are, uniquely among mammals, at genuine risk of dying from eating. We have traded mechanical safety for acoustic flexibility. And the fact that evolution made this trade tells us something important about the selective advantage that speech provided. The benefit was so large that even the cost of potential death from swallowing was worthwhile.

The laryngeal descent in humans is not present at birth. Human infants [music] have a high larynx like other primates, and it descends during early childhood, completing its descent sometime around age three or four, which is roughly when children’s speech rapidly expands [music] in complexity. The timing is not coincidental.

The anatomy creates the possibility. The neurology then learns to exploit it. >> [music] >> Now here is where the fossil record enters the story, and where the [music] difficulty of studying the evolution of speech becomes apparent. Language leaves almost no direct fossil evidence. [music] Words don’t fossilize.

Sounds [music] don’t fossilize. The neural activity of communication doesn’t fossilize. What we’re left with is indirect [music] evidence. The bones and casts of the skull and vocal tract that can tell us something about what an ancient hominid was anatomically capable of producing. The most informative bone for this purpose is the hyoid.

The hyoid is a small, horseshoe-shaped bone that sits at the base of the tongue in the throat, attached to nothing. It floats, held in place entirely by muscles and ligaments. It’s the only bone in the human body with no direct articulation to any other bone. And it serves as an attachment point for the muscles of the tongue and the floor of the mouth.

Because the hyoid is so closely connected to the muscles involved in speech production, its shape can provide clues about the vocal capabilities of extinct hominids. And unlike soft tissue, bone fossilizes. In 1989, a remarkably complete Neanderthal skeleton was excavated at Kebara Cave in Israel. The skeleton, >> [music] >> dated to approximately 60,000 years ago, was exceptional in its preservation.

 And crucially, it included the hyoid bone. When researchers examined the Kebara hyoid, they found something that generated significant debate in the paleoanthropology community. It was essentially identical to a modern human hyoid. The shape, the proportions, the muscle attachment surfaces, all of them matched modern humans closely enough that if you showed the bone to someone without telling them where it came from, they might not immediately identify it as anything unusual.

This finding has important implications. If Neanderthal hyoid morphology matches modern human hyoid morphology, and if hyoid morphology is related to vocal tract anatomy, then Neanderthals may have had a descended larynx similar to ours, and therefore may have had the anatomical substrate for human-like speech production.

The Kebara hyoid alone doesn’t prove Neanderthals could speak. The relationship between hyoid morphology and vocal capabilities is complex. And having the right bone shape doesn’t guarantee >> [music] >> the right neural control. But it removes one of the anatomical objections to Neanderthal speech capability. There is another line of fossil evidence [music] worth examining.

The canal through which the hypoglossal nerve passes in the base of the skull. The hypoglossal nerve controls the tongue. Its size reflects the number of nerve fibers, and nerve fiber number reflects how much neural control the tongue required for its owner. Modern humans have a significantly larger hypoglossal canal than chimpanzees, reflecting the elaborate neural control needed for the precise [music] tongue movements of speech.

When researchers examined this canal in early Homo specimens and in Neanderthals, >> [music] >> they found it to be within the modern human range, suggesting that the neural machinery for fine tongue control was in place long before modern humans appeared. Then there is the brain itself. We cannot directly study the brains of extinct hominids, brain tissue doesn’t survive fossilization.

But, the inside of the skull preserves an impression of the brain’s surface, an endocast, that reveals something about the brain’s gross anatomy, including the development of particular regions. Broca’s area, in the left frontal lobe, is the region most directly associated with speech production in modern humans.

People with damage to Broca’s area lose the ability to speak fluently. They can understand language, but struggle to produce [music] it in a condition called Broca’s aphasia. The impression of Broca’s area can sometimes be detected in cranial endocasts, >> [snorts] >> and its presence in early Homo species has been used to argue for some form of language capability in our ancestors.

The evidence from endocasts is suggestive, but not conclusive. The gross anatomy of a brain region tells us less than we would like about its function. And the relationship between Broca’s area and speech is more complex than was once thought. But, when you add the hypoglossal canal data, the hyoid evidence, and the endocast data together, you get a picture of Neanderthals and late archaic Homo sapiens >> [music] >> as creatures with anatomy at least consistent with speech capability.

Whether they actually used that capability for something recognizable as language is a separate question. And it leads us to the most surprising evidence of all. In 2002, >> [music] >> a research team led by Wolfgang Enard at the Max Planck Institute for Evolutionary Anthropology published a finding that changed the field of language evolution overnight.

They had identified a gene F O X P 2 that appeared to be directly involved in the fine motor control required for speech. The finding had come from the study of a British family with severe speech and language difficulties. Members of this family across three generations had trouble with the precise oral motor movements required for speech.

Their lips, >> [music] >> tongue, and jaw movements were insufficiently coordinated for fluent speech production. When researchers analyzed the family’s genetics, they found a mutation in a gene called F [music] O X P 2 forkhead box P 2 present in all affected family members and [music] absent in unaffected ones.

F O X P 2 was quickly dubbed the language [music] gene. A label that turned out to be an oversimplification, but that captured something genuinely important. F O X P 2 is a transcription factor. A protein that controls the activity of other genes. And it plays a role in the development and function of circuits in the brain and elsewhere that are involved in learning and performing complex sequential movements.

 Speech is, among other things, a complex sequential movement. A precisely coordinated sequence of lip, tongue, jaw, larynx, [music] and breathing muscle actions performed at extraordinary speed. Fox 2 appears to be important for the neural machinery that learns and executes such sequences. But here is what made the discovery truly fascinating.

Fox 2 is not a uniquely human gene. It is found in birds, in crocodilians, in fish, in virtually all vertebrates. It is ancient. It has been doing something in vertebrate nervous systems for hundreds of millions of years, long before anything resembling speech existed. In songbirds, mutations in the equivalent gene disrupt song learning.

The birds can still vocalize, but cannot learn the complex songs characteristic of their species. In mice, Fox 2 mutations produce animals with abnormal vocalizations and reduced ability to learn motor sequences. The gene does something important across many different vertebrate lineages, in many different contexts involving complex learned movement.

What is unique about the human version of Fox 2 >> [music] >> is two specific amino acid changes. Two positions in the protein where the human sequence differs [music] from the sequence in all other known mammals except one. The human protein differs from the chimpanzee protein at exactly [music] two positions.

These two changes appear to have occurred somewhere in the hominid lineage after our split from chimpanzees approximately six [music] to seven million years ago. When researchers introduced the human version of Fox 2 into mice, genetically engineering mice to carry the human amino acid sequence rather than the mouse sequence, the humanized mice showed changes in the neural circuits of the basal ganglia, a brain region involved in learning motor sequences.

They didn’t suddenly start speaking, but their motor learning was altered in ways consistent with the human version conferring [music] some advantage for precisely coordinated sequential movement. When did the human FOXP2 mutations appear? This is a critical question for understanding the timeline of speech evolution.

The answer comes from population genetics. [music] By analyzing the variation in the FOXP2 gene sequence [music] across modern human populations, researchers can estimate when the current human variant [music] was subject to positive selection, when it spread rapidly through the human lineage because it provided an advantage.

The analysis suggests positive selection on FOXP2 occurred sometime in the last 300,000 [music] years, possibly more recently. This would place the human FOXP2 [music] variant as a relatively recent innovation in the hominid lineage, consistent with the emergence of fully modern human anatomy and behavior. But then came the twist.

Ancient DNA extracted from Neanderthal fossils revealed that Neanderthals carried the same two amino acid changes in their FOXP2 protein as modern humans. The human version of FOXP2 is not uniquely human. It is shared with Neanderthals, who diverged from our lineage somewhere between 500,000 and 700,000 years ago.

This means the FOXP2 [music] mutations either appeared in the common ancestor of modern humans and Neanderthals before the lineages split or they were transferred between the lineages through interbreeding, [music] which we now know from ancient DNA evidence did occur. Either way, the Neanderthal FOXP2 [music] data adds to the growing body of evidence that Neanderthals were not the speechless brutes of earlier popular imagination.

Let us step back [music] from the specific evidence and think about the larger question. Why did speech evolve at all? What was it for? The intuitive answer is that speech evolved for the communication of information, to share knowledge about where food is, where predators are, how to make tools, and this is certainly true as far as it goes.

But Robin Dunbar, a British anthropologist and evolutionary psychologist, has argued that this intuitive answer, while not wrong, misses the deeper social function that language may have originally evolved to serve. Dunbar’s starting point is a puzzle about group size. In primates, there is a well-documented relationship between the size of the neocortex, [music] the outer layer of the brain responsible for complex cognition, >> [snorts] >> and the size [music] of the social group the species typically lives in.

Species with larger neocortices tend to live in larger, more complex social groups. The relationship is robust across many primate species and makes intuitive sense. Managing complex social relationships requires cognitive sophistication. More relationships means more cognitive load. Dunbar’s measure of group size for humans, the number of relationships a person can meaningfully maintain, comes out to approximately 150, a number now known >> [music] >> as Dunbar’s number.

This figure appears repeatedly in human social organization, the size of traditional hunter-gatherer bands, the size of effective [music] military units, the size of villages in pre-industrial societies. It seems to reflect something real about the cognitive limits of human social relationship management. But how do primates maintain their social relationships? The answer for most primates is physical grooming.

Primates groom each other, picking through fur for parasites, cleaning wounds, touching. And this physical contact is not primarily about hygiene. It’s about social bonding. Grooming releases endorphins in both the groomer and the groomee, creating positive social associations, reinforcing alliances, and managing the complex social hierarchies of primate groups.

In many primate species, individuals spend up to 20% of their waking hours grooming others. Here is Dunbar’s key insight. >> As group size increases, the time required for physical grooming to maintain all necessary social relationships increases proportionally. For very large groups, groups of the size that Dunbar’s number suggests humans were [music] maintaining, physical grooming would require an impossibly large fraction of the waking day.

You simply cannot groom 150 individuals daily [music] or even weekly without sacrificing most of the time needed for everything else. Something needed to replace physical grooming as the social bonding mechanism. Something that could bond with multiple individuals simultaneously >> [music] >> at a distance without requiring physical contact.

Language, in Dunbar’s hypothesis, is vocal grooming. Speech evolved not primarily to convey factual information, but to maintain social bonds across larger groups than physical grooming could manage. And the specific content of early language, in this view, was social information. Who is allied with whom? Who did what to whom? Who can be trusted? Who is cheating? Information that is directly relevant to navigating complex [music] social hierarchies.

In modern terms, gossip. Dunbar estimates that approximately 60% to 65% of human conversation consists of social topics. The exchange of information about people and relationships. The proportion is remarkably consistent across different cultures and contexts. We are a deeply, constitutively social species, [music] and our language reflects that.

If Dunbar’s hypothesis is correct, then the selective pressure that drove the evolution of language was not [music] environmental. Not the need to communicate about predators or food sources or tool making, but social. Language was selected for because it enabled larger social groups, and larger social groups provided advantages.

More individuals to share the work of survival, more defenders against threats, more pairs of eyes watching for danger, more cognitive resources pulled for problem solving. The benefits of group living are substantial, and language extended the upper limit of viable group size. This reframing has a specific implication for the question of what the first words were.

If language evolved for social bonding and social information exchange, then the first meaningful vocalizations were probably not the names of objects or places or actions. They were social calls, expressions of alliance and affiliation, and eventually names. Names are among the most basic social linguistic tools.

The ability to refer to a specific individual by a unique identifier allows you to talk about someone who is not present, which is essential for the kind of social information exchange that Dunbar’s hypothesis predicts would have been the original function of language. You can’t gossip without names. The names hypothesis is speculative, deeply speculative, [music] because the earliest stages of language leave no fossil record.

But it is consistent with both the social origins hypothesis and with what we know about language acquisition in children. Children learn names for people and objects very early. The capacity to associate specific sounds with specific reference and to learn new such associations rapidly is fundamental to human language and appears to be present even in infants.

Now, we must address the hardest question in this field. When? When did language emerge? The answer depends critically on what you mean by language. If you mean the full modern capacity, the recursive, grammatically structured, unlimited combinatorial system that allows modern humans to produce and understand an infinite number of novel sentences, then the evidence points towards something in the range of 200 to 300,000 years ago, associated with the emergence of anatomically modern Homo sapiens in Africa.

This is the view sometimes called the late emergence hypothesis. But if you mean something more modest, meaningful vocalizations, >> [music] >> shared reference, some capacity for symbolic communication, then the evidence from the fossil record, from the FOXP2 data, and from the archaeology of Neanderthal behavior pushes the timeline back considerably.

And this is where the scientific debate becomes most interesting, most unresolved, and most revealing about the nature of what we’re trying to explain. The archaeological evidence for symbolic behavior is the most direct window we have into cognitive sophistication in ancient hominids. Language, as we experience it, is fundamentally symbolic.

Sounds that stand for things by convention, not by natural resemblance. If we can find evidence that ancient hominids were using symbols, creating things that referred to something beyond their physical selves, we have indirect evidence for at least some of the cognitive infrastructure that language requires.

The oldest evidence for unambiguous symbolic behavior in modern humans dates to roughly 75 to 100,000 years ago in South Africa. At sites like Blombos Cave, archaeologists have found ochre, a red iron oxide pigment that appears to have been deliberately processed and used for coloring, probably for body decoration or marking objects.

They have found pierced shells that may have been strung as [music] beads. They have found geometric engravings on pieces of ochre. These objects appear to be symbolic. They don’t serve any obvious practical function. >> [music] >> And the deliberate choice to create them suggests an organism capable of representing meaning through physical artifacts.

But here, the archaeologists face a challenge. The earliest evidence for symbolic behavior in modern humans predates the behavioral explosion often associated with the Upper Paleolithic Revolution in Europe, which begins around 40 to 50,000 years ago, >> [music] >> and which produces such a dramatic florescence of art, sophisticated tools, and personal ornaments that at once seemed to represent the sudden emergence of modern cognitive capability.

We now know that this picture was distorted by the geographic bias of early archaeology. Africa has yielded much earlier evidence for symbolic behavior. And the so-called Upper Paleolithic Revolution now looks less like a sudden leap and more like a gradual process whose early stages occurred in Africa long before humans reached Europe.

Meanwhile, Neanderthal behavior has also been revised. Neanderthals in Europe, dating to well before their contact with modern humans, produced pigments, possibly made jewelry, and may have created cave art. The evidence is disputed. Some of it has been questioned by researchers who argue that apparent Neanderthal symbolic objects are either misidentified or represent copying of modern human behavior rather than independent invention.

But the overall picture has shifted significantly from the earlier view of Neanderthals as cognitively limited compared to modern humans. If both modern humans and Neanderthals showed some capacity for symbolic behavior, and if both groups shared the same FOXP2 variant, then either symbolic behavior and FOXP2 associated language capability were both present in the common ancestor of the two lineages, which would push meaningful language back to at least 500,000 years ago, or both innovations arose independently

in the two lineages, which seems less parsimonious. The middle position, increasingly favored by researchers synthesizing the anatomical, genetic, and archaeological evidence, is that language did not emerge in a single moment. It evolved incrementally over hundreds of thousands, or perhaps >> [music] >> millions of years, with different components of the full modern language faculty appearing at different [music] times, providing selective advantages at each stage, and only achieving something like the modern full system

in the period associated with behaviorally modern humans. This incremental view avoids the implausibility of complex language appearing suddenly and fully formed. But, it raises its own challenge. If language was built incrementally, what were the stages? What did proto-language look like? How did you get from no language to something language-like? And from something language-like to full language? Michael Corballis, a New Zealand cognitive psychologist, has argued that language evolved from gesture.

That the communicative foundations of language were manual, not vocal. And that speech evolved later as a more efficient and less [music] physically demanding way to convey the same communicative content. The gesture-first hypothesis has the advantage of explaining why language is so closely associated with manual gesture even in fully spoken language.

Humans gesture constantly while speaking, and gesture in ways that are not merely redundant with speech, but that carry semantic content complementary to the spoken content. Deaf communities spontaneously develop full sign languages, complete with grammar and syntax, without any spoken component. The neural machinery for language seems to work equally well with hands as with voice.

The gesture first hypothesis also connects language evolution to the evolution of tool use, which requires precise manual control and complex sequential planning. The same neural systems, including possibly components of Broca’s area, that support language production, may have been co-opted from earlier systems supporting manual action and tool use.

 Language, in this view, is a form of action, sequential, hierarchically organized, learned through practice, and it may have exploited neural architecture that evolved originally for a different kind of action. But the honest answer to the question of what proto-language looked like is that we don’t know. The transition from primate vocalizations to human language is the most significant cognitive evolutionary transition in the history of our lineage, and it left almost no direct trace.

We can see the end points, the vocalizations of living primates on one side, modern human language on the other, but the intermediate stages are invisible to us. We have to reconstruct them from fragments, the bones of the throat and skull, the genetics of language-related genes, the archaeological traces of behavior that language makes possible.

What we can say with confidence is this. By the time anatomically modern humans appear in the fossil record, around 300,000 years ago in Africa, the anatomy for full modern speech was in place. The brain, the hyoid, the descended larynx, the FOXP2 protein, all of them present in their modern form. And the behavior that begins to appear in the archaeological record from that period onward, the pigments, the ornaments, the long-distance trade in materials, the increasing complexity and standardization of tool forms,

is consistent with the species using language to coordinate complex, social, and technological activities. Whether the vocal anatomy and the behavior appeared simultaneously, or whether the anatomy preceded the full behavioral expression of language by hundreds of thousands of years, is unknown. It is entirely possible that hominids possessed the anatomical machinery for complex speech long before they fully exploited it.

Just as modern humans possess brains capable of far more than most of us exercise in daily life, the capacity and the use of the capacity are not the same thing. Consider what language, once it was present in something like its modern form, would have made possible. Consider it as a technology, because that is, in important ways, what it is.

Before language, the knowledge possessed by any individual member of a group died with that individual, or could be passed on only by direct demonstration, watching and imitating. After language, knowledge became transmissible through description. You could tell someone how to do something you had done.

 You could warn them about something you had experienced. You could pass on information about events that had happened before they were born. Language is a technology for the accumulation and transmission of knowledge across time. And this accumulation is what we call culture. The extraordinary thing about cumulative culture is that it allows each generation to build on what the previous generation learned, rather than starting [music] from scratch.

A lone human being born into isolation would be cognitively formidable by animal standards, but would be unable to reinvent agriculture, metallurgy, writing, >> [music] >> or any of the other technological innovations that make civilization possible. These innovations emerged over thousands of generations of incremental improvement.

Each generation building on what the previous one had established. Language made this accumulation possible. Without language, cumulative culture doesn’t exist. Without cumulative culture, the modern world doesn’t exist. This is why the emergence of language is, in terms of its planetary consequences, the most significant event in the last several million years of Earth’s history.

It changed not just the trajectory of one species, but the trajectory of the entire biosphere. The moment the first hominid produced a sound that another hominid understood as referring to something specific, however rudimentary that exchange was, however far it was from what we do when we speak, was the moment that led, through an unbroken chain of incremental developments, to the world we inhabit.

Cities, writing, science, music, law, philosophy, all of it downstream from a sound made in Africa between 200,000 and 2 million years ago. We don’t know what that sound was. We don’t know who made it. We don’t know exactly when it [music] happened. The fossil record is silent on these specific questions and it may always be.

But the questions themselves, the attempt [music] to trace the single most consequential innovation in the history of animal cognition back to its origins, are among the most important questions science has ever asked. And the answers we have assembled so far, fragmentary and contested as they are, paint a picture that is both more complicated and more interesting than the simple story of a sudden moment of enlightenment.

Language did not arrive like lightning. It grew slowly over vast spans of time in the anatomy and the neural circuits and the social lives of creatures who were, generation by generation, becoming something new. By the time the full modern system was assembled, by the time recursive grammar and symbolic reference and the infinite combinatorial possibilities of human language were all present and operating, the creatures using it were already us.

Already building fires, making art, burying their dead, trading across distances that would have seemed inconceivable to their ancestors. Already, in the most essential sense, human. Language did not arrive like lightning. It grew. And whatever first sparked [music] that growth remains for now beyond our reach.

We are those creatures. And we are still listening. >> But let us slow down at a specific point in the story. The moment where anatomy meets behavior. Because this intersection is where the evidence is richest and the implications most dramatic. The descended larynx, as we discussed, is one of the clearest anatomical signatures of speech capability in the fossil record.

But there is another anatomical feature that reveals something equally important about the evolution of speech. The shape of the thoracic vertebral canal. The thoracic vertebral canal houses the spinal cord as it passes through the chest region. And it contains the nerves that control the muscles of breathing, the intercostals and the diaphragm.

For ordinary breathing, the demands on these muscles are modest. For speech, they are remarkably complex. When we speak, we must precisely modulate our exhalation, varying the pressure and flow of air through the vocal tract in ways that correspond to the sounds we’re producing. This requires fine neural control of breathing muscles that doesn’t exist in ordinary respiration.

And fine [music] neural control requires more nerve fibers, which means a larger spinal canal. When Ann MacLarnon at the University of London compared the thoracic vertebral canal size in humans versus non-human primates in the 1990s, she found what she expected. Modern humans have a significantly larger thoracic spinal canal than chimpanzees, consistent with the greater neural control of breathing that speech requires.

When she then applied the same analysis to Neanderthal and archaic Homo sapiens vertebrae, to the extent that these specimens were available, she found that some archaic specimens had thoracic canals within the modern human range. The neural infrastructure for the precise breathing control that speech requires appears to have been present in hominids considerably earlier than modern humans.

>> This is one piece of a puzzle that, assembled carefully, begins to suggest that the anatomical prerequisites for speech were in place remarkably early in the Homo lineage, perhaps as early as Homo heidelbergensis or even Homo erectus, well before the emergence of anatomically modern humans 300,000 years ago.

The implication is not necessarily that these earlier hominids had language in the modern sense. It is that the anatomical toolkit was assembled piece by piece over a very long period, with each piece providing some advantage and being [music] retained by selection, before the whole system was finally operating at modern efficiency.

The question of Homo Erectus and language is particularly fascinating and contentious. Homo Erectus, which first appears in the fossil record in Africa approximately 1.8 to 2 million years ago and spread across much of the Old World, is associated with the Acheulean tool tradition. A specific and highly standardized form of stone tool, the hand ax, characterized by a symmetrical teardrop-shaped form that required considerable skill to produce and that remained essentially unchanged across enormous geographic distances and spans

of time up to a million years or more. The standardization of Acheulean tools across such vast distances and time frames is difficult to explain without some form of learned cultural transmission. A way for individuals to acquire the specific manufacturing technique from others rather than reinventing it independently.

Direct demonstration and imitation alone could in principle account for this transmission without requiring language. But the precision and [music] consistency of the tools and the fact that their form was maintained across populations that had no contact with each other for geological spans of time have led some researchers to argue that some form of proto-language must have been involved.

That the information required to produce these tools reliably was too complex to transmit by imitation alone and required the kind of explicit, sequential instruction that only something like language could provide. Others disagree strongly. They argue that great apes can learn complex tool-making sequences by observation alone without language.

And that the consistency of Acheulean tools may reflect constrained design space. There are only so many ways to make an effective hand ax, and selection for effectiveness naturally produces convergent forms rather than linguistic transmission. The debate is genuinely open, and it illustrates the fundamental challenge of this field.

Because language leaves no direct traces, its presence in ancient hominids must always be inferred from indirect evidence. And intelligent researchers can look at the same indirect evidence and reach different conclusions. What changed with the emergence of anatomically modern Homo sapiens in the range of 200 to 300,000 years ago that might have given modern humans a cognitive advantage over earlier Homo species, including Homo erectus and eventually Neanderthals? The anatomical differences are relatively modest.

Modern human skulls are rounder with more globular brain cases. The face is flatter. The chin is more pronounced. These differences reflect changes in brain shape that, under analysis, turn out to involve expansion of specific regions. The parietal lobes, which are involved in spatial reasoning and numerical cognition, and the temporal lobes, which are heavily involved in language processing.

 The expansion of the temporal lobe in modern humans is particularly [music] significant for the language story. The left temporal lobe contains Wernicke’s area, the region most directly involved in language comprehension, the ability to understand and process the meaning of speech. People with damage to Wernicke’s area can speak fluently, but their speech is semantically incoherent.

Words are produced in grammatically plausible arrangements, but they don’t make sense. And the patients are often unaware that they are not communicating meaningfully. The expansion of the temporal lobe in modern humans, detectable even in endocasts, suggests an enhancement of language comprehension capability that may have given modern humans a linguistic advantage over archaic Homo species with somewhat different brain proportions.

The brain organization story is complex and still actively researched, but the broad outline is becoming clearer. Modern human cognition is not simply a scaled-up version of archaic hominin cognition. It is differently organized, with specific regions involved in language, social cognition, and symbolic thought more developed relative to other brain regions.

These differences in organization, rather than simple brain size increase, may be what gave modern humans the full modern language faculty. Now, let us consider the social context more carefully, because the Dunbar hypothesis of language as vocal grooming deserves fuller treatment than the summary it often receives.

The hypothesis makes a specific and testable prediction. If language evolved to maintain social bonds in large groups, then the content of human language should be predominantly social. Dunbar and his colleagues tested [music] this by systematically recording and analyzing conversations in cafeterias, restaurants, train stations, and other public spaces in Britain.

>> [music] >> They classified conversational topics and found that approximately 65% of conversations were social in content, talking about people, relationships, experiences. This proportion was consistent across different [music] settings and different participant groups, suggesting it reflects something fundamental about what language is for rather than what any particular culture happens to talk [music] about.

The social language hypothesis also makes predictions about the neural correlates of language processing. If language evolved from social grooming, we might expect the neural systems supporting language to overlap with the neural systems supporting social cognition. And there is growing evidence that this is the case.

The default mode network, a set of brain regions that become active during rest and that are associated with social cognition, including mentalizing, theory of mind, and thinking about other people, overlaps substantially with the neural networks supporting language processing. When you understand a narrative, you are using some of the same neural machinery you use when you think about what other people are thinking and feeling.

Theory of mind, the ability to attribute mental states to others, [music] to understand that other individuals have beliefs and intentions that may differ from your own is deeply intertwined with language. To understand language, you must understand that the speaker intends to communicate something specific to you.

To produce language effectively, you must model what your listener knows and doesn’t know, what they’ll understand and what requires explanation. Every act of communication involves modeling the mind of your interlocutor. And theory of mind, conversely, is greatly enhanced by language. The ability to describe mental states verbally, to reason about them explicitly, to communicate about what others might be thinking.

The relationship between theory of mind and language is so intimate that some researchers have argued they are fundamentally the same cognitive system, or at least that language and theory of mind coevolved, each enhancing the other in a positive feedback loop. As language allowed more precise communication about mental states, theory of mind became more sophisticated.

As theory of mind became more sophisticated, language could become more complex and nuanced, able to express subtler distinctions of meaning and intention. The two capabilities pulled each other upward, each advance in one creating [music] the conditions for an advance in the other. Children develop language and theory of mind on roughly the same developmental timeline.

Before about 18 months, children show limited theory of mind. They don’t consistently pass tests that require understanding that someone else has a false belief. Between 18 months and 4 years, both language and theory of mind develop rapidly, >> [music] >> with the two capabilities closely intertwined throughout.

Researchers who have studied children with language delays often find corresponding delays in theory of mind development. The developmental coupling of the two systems suggests deep [music] evolutionary coupling as well. There is another cognitive capacity that is uniquely or almost uniquely human, and that is deeply connected to language.

The ability to think about time. Specifically, to mentally travel to past and future events. Episodic memory. The ability to recall specific episodes from one’s past exists in some form in other animals. But the human capacity for mental time travel, for constructing detailed simulations of past and future scenarios, is far more developed than in other species, and is intimately linked to language.

We think about the past and future [music] largely in words. Language provides the scaffolding for temporal reasoning. >> Without language, it is difficult [music] to represent future events with the specificity required for [music] complex planning. You cannot communicate a plan to others without language. And it is arguably difficult even to hold a complex plan in mind.

To represent a multi-step future action sequence in a form precise enough to guide actual behavior. Without the kind of symbolic representation that language provides. The connection between language, planning, and temporal reasoning is another reason why language may have been so powerfully selected for in the hominin lineage.

In an environment where cooperation, planning, [music] and complex social organization were increasingly important for survival, the ability to represent future scenarios, communicate about them, and [music] coordinate collective action around them would have provided an enormous advantage. The first sentence ever spoken, whenever it was, by >> [music] >> whoever said it, was almost certainly about something that was not present at the moment of speaking.

That is, in a deep sense, what language is for. Language [snorts] allows us to communicate [music] about things that are absent, past, or future. Without language, communication is limited to the immediate present, [music] to what can be pointed at, demonstrated, or expressed through tone and gesture [music] in the moment of experience.

With language, the scope of communication expands to encompass [music] everything that has ever happened, everything that might happen, everything that exists anywhere, [music] and everything that doesn’t exist anywhere, but might be imagined. That is an extraordinary expansion of the communicative horizon, and it made possible everything that followed.

The children who grow up speaking are not aware of the extraordinary thing they are doing. [music] They acquire language with an ease that suggests it is natural, inevitable even, that human beings communicate this [music] way. And in a sense, it is natural. The human brain is prepared for language acquisition in ways that no other known brain is.

 The ease of language acquisition in children, across all cultures and all circumstances, is one of the strongest arguments that the human brain was shaped by selection specifically for language. That the neural infrastructure was built by millions of years of evolution to make [music] this particular cognitive feat as effortless as possible.

We are adapted for language the [music] way birds are adapted for flight. It looks easy because we were built for it. But it was not always easy. It was not always natural. There was a time when no creature on Earth did this. When every sound made by every living thing was either a direct expression of an emotional state or a learned signal with a fixed meaning.

An alarm call for predators. A contact call for conspecifics. A mating call for potential partners. There was no arbitrary reference, no symbolic meaning, no grammar, no recursion. The world that existed before language was, in its own way, enormously complex. Billions of organisms interacting in intricate ecosystems, sophisticated behaviors evolved over hundreds of millions of years.

But it was a world without stories, without plans, without arguments, >> [music] >> without questions asked simply out of curiosity about how things are. The hominid lineage spent several million years building toward language. The anatomy changed. The genetics [music] changed. The brain changed. And somewhere in that process, at a moment we cannot pinpoint and may never pinpoint, the threshold was crossed.

 [music] The first meaningful sound was made and understood. And the world was never the same. We have been talking ever since. Across every culture that has [music] ever existed on Earth, in every climate and every environment, humans talk. We talk constantly, almost compulsively. We talk to ourselves when no one else is present.

We talk in our dreams. We think in language. Or at least, we use language as a primary medium for thought, for planning, for self-reflection, for [music] understanding what we feel and who we are. Language is not just something humans do. It is something humans are. Strip it away and you have an animal, however intelligent, that is fundamentally different from the creature that built cities and wrote symphonies.

Let us take a different approach to the timeline question. Not through anatomy or genetics, but through archaeology. Because there is one class of archaeological evidence that bears directly on language capability in a way that bone and DNA [music] cannot. The evidence of coordinated, large-scale, collective [music] behavior.

Language is not just a tool for communication between two individuals. [music] It is the mechanism that makes large-scale social coordination possible. And large-scale social coordination leaves archaeological signatures that we can detect. Consider long-distance trade in raw materials. For most of human prehistory, the materials people used, the stone for their tools, the pigments for their art, the shells for their ornaments, came from the immediate local environment.

You used what was available nearby. The archaeological record of the Middle Stone Age in Africa, however, begins to show something different. Materials that had traveled significant distances from their geological sources to the sites where they were used. Ocher from sources tens or hundreds of miles distant.

Shells from the coast appearing at inland sites. >> [music] >> Marine shells found far from any shoreline. This long-distance movement of materials implies social networks. Relationships between groups in different locations that enabled the exchange of goods across distances [music] that no individual would travel regularly.

And social networks of this complexity, maintained over generations, are very difficult to sustain without [music] language. You need to be able to communicate about the terms of exchange, about who owes what to whom, about the reliability of trading partners, about the quality of the materials being exchanged.

 [music] You need to be able to talk about people you haven’t met, places you haven’t been, agreements reached in the past that have implications for the future. This is language doing exactly what Dunbar’s hypothesis predicts it should do. Maintaining social bonds and enabling cooperation across distances that physical proximity alone cannot sustain.

The earliest evidence for systematic long-distance material [music] exchange in Africa dates to approximately 100 to 150,000 [music] years ago. This is well within the period of anatomically modern humans, but the cognitive capacity it implies [music] and the social organization it requires is considerably more advanced than what most people imagine >> [music] >> when they think about ancient humans.

These were not isolated groups of primitive hunters. They were participants in exchange networks spanning hundreds of miles, maintaining relationships with groups they may have encountered only occasionally, operating [music] within social systems of considerable complexity. And complexity requires coordination.

Coordination requires communication. Communication of this level [music] of sophistication requires language. The long-distance trade networks are one of the most compelling indirect arguments for fully modern language capability in the populations that created them. There is another archaeological signature worth examining.

The evidence of collective coordinated hunting. Modern humans and their recent ancestors appear to have organized coordinated hunts at a level of complexity not seen in other primates. The hunting of large, dangerous prey, mammoths, buffalo, horses, required multiple individuals acting [music] in concert with specific roles, responding to changing circumstances in real time.

This kind of hunting is [music] not just difficult physically. It is cognitively demanding in ways that depend critically [music] on communication. You need to be able to plan before the hunt, to agree on roles and strategies. [music] You need to be able to coordinate in the moment.

 To signal what you’re seeing and what action you’re taking. You need to be able to debrief after, to learn from what worked and what didn’t. Non-human primates do cooperate in hunting. Chimpanzees hunt colobus monkeys in coordinated groups with individuals appearing to take different roles. But the chimpanzee hunts lack the precision and complexity of human hunts.

And they appear to be organized opportunistically, rather than planned in advance. The difference between chimpanzee group hunting and the organized drives of modern human [music] hunters, where animals are systematically moved toward waiting hunters through coordinated action across a landscape, is substantial.

And language appears to be a large part of what makes that difference. The evidence [music] for sophisticated organized hunting appears in the archaeological record at various points in the African Middle Stone Age, from around 200,000 years ago. Sites in South Africa show evidence of systematic [music] exploitation of specific species, suggesting selective hunting, rather than opportunistic scavenging.

 [music] Sites in East Africa show evidence of hunting large, dangerous prey in ways that imply coordinated group action. These behavioral signatures are consistent with populations using language to organize collective activity. Now let us examine one more dimension of the language evolution question that tends to receive less attention than the anatomical and genetic evidence.

The role of music. Music and language share a striking number of features. Both [music] are organized sequences of sounds. Both involve pitch, rhythm, and timing. Both engage overlapping neural systems. And both appear to be universal features [music] of all known human cultures. No human culture has ever been found without language.

And none has been found without some form of music. The two are as universal as any feature of human biology. The connection between music and language goes beyond superficial similarity. The neural systems that process the hierarchical structure of language, the way phrases are embedded within phrases, the way grammatical rules organize words into meaningful sequences, substantially overlap with the neural systems that process musical structure.

Brain regions involved in rhythm processing contribute to the processing of speech rhythm. Regions involved in melodic processing contribute to the processing of speech prosody, [music] the rises and falls of pitch that carry emotional content and [music] disambiguate meanings in speech. The auditory processing pathways that analyze the complex acoustic structure of music and the complex acoustic structure of speech share substantial neural infrastructure.

Some researchers, including Steven Mithen in his book The Singing Neanderthals, have proposed that music and language evolved together. That both emerged from a common ancestral system of emotional vocalization that was both musical and [music] protolinguistic. The calls of many primates are pitch-varied [music] and rhythmically organized.

 They have properties that we associate with music as well as properties we associate with language. Perhaps the earliest stages of language were more song-like than speech-like. Melodically organized vocalizations whose meaning was carried as much by acoustic properties as by arbitrary symbolic reference. This proposal connects interestingly to the evidence from Neanderthals.

 [music] If Neanderthals had the vocal anatomy, the FOXP2 genetics, and some of the neural organization associated with modern human speech, but nevertheless had a different kind of linguistic capability, perhaps less fully recursive, less symbolically complex, more emotionally organized, >> [music] >> then they might represent something close to the proto-language stage that Mithen imagines.

>> A communication system that was more musical, more emotionally immediate, less [music] abstractly referential than modern language, but that was more than simple primate vocalization. We will never know what Neanderthals sounded like when they communicated with each other. We will never know whether they told stories, asked questions, made promises, or argued about abstract ideas.

Their language, [music] if they had it, is gone as completely as they are. But the evidence suggests they were not silent. They were not the brutish, inarticulate creatures that early depictions showed. They buried their dead, cared for their injured, made pigments, and possibly made music. Whatever they communicated, they communicated something.

Something real, something social, something that required more than grunts and gestures to express. The question of what distinguished modern humans from Neanderthals cognitively, if modern humans were indeed cognitively superior in relevant ways, may never [music] be fully resolved. But the archaeological record does seem to show that after the behavioral revolution associated with modern humans around 40 to 50,000 years ago in Europe, the pace of cultural change accelerated dramatically.

New technologies appeared, spread, and were replaced [music] by better technologies on timescales much shorter than anything seen in the preceding hundreds of thousands of years. Art became more elaborate and widespread. Social networks became larger and more complex. In short, cumulative culture, the defining feature of modern human cognition, the capacity to build knowledge across generations, appears to have accelerated.

Cumulative culture requires language. It requires the ability to transmit precise, >> [music] >> complex information across generations in a form that can be built upon, and its acceleration in the archaeological record is the clearest evidence we have that the full modern language capacity, whatever precisely distinguishes it from what came before, was in place and operating effectively by the time of the Upper Paleolithic.

So, here is the timeline as best we can currently reconstruct [music] it. The anatomical prerequisites for speech, the descended larynx, the neural control of breathing and tongue movement, [music] the hyoid morphology, were assembled gradually over the past million or more years of hominin evolution with much of the machinery in place by the time of Homo heidelbergensis, perhaps 500,000 [music] years ago.

The genetic changes associated with language capability, particularly in FOXP2, were present in the common ancestor of modern humans and Neanderthals, suggesting significant language-related neural evolution before the lineages split more than 500,000 years ago. Some form of symbolic behavior, pigment use, ornament making, appears in the archaeological record by at least 100,000 years ago, and possibly as early as 300,000 years ago.

And fully modern, rapidly cumulative culture, with all that it implies for language sophistication, appears to be operating at full capacity by 40 to 50,000 years ago. The most probable conclusion, given this evidence, is that language evolved incrementally over hundreds of thousands of years with early stages of proto-language present in archaic Homo, including Neanderthals, and the full modern system established in Homo sapiens, perhaps 200 to 300,000 [music] years ago, with its effects on behavior becoming increasingly pronounced over the

subsequent tens of thousands of years. The first word, the first sound that meant something specific [music] by agreement rather than by natural association, was spoken by one of our ancestors. That ancestor was alive on this planet, walking across African landscapes we can still visit today, [music] looking at a sky that was already ancient and would persist long after they were gone.

They did not know what they were doing [music] when they produced that sound. They were not making history. They were not conscious of standing at a threshold in the story of life on Earth. They were doing what animals do, communicating, coordinating, surviving. But the sound they made, whatever [music] it was, was different from what had come before.

And because of that difference, [music] the story of life on Earth took a turn that no geological force, no physical process, >> [music] >> no previously existing biological capability could have produced. The turn toward language. The turn toward culture. The turn toward the world in which you are now receiving these words, understanding them, perhaps being moved by them.

Everything that has ever been said is the echo of that first sound. Every word spoken in every language that has [music] ever existed, every story told by every fire in every encampment across a hundred thousand years of human history, every poem, every argument, every declaration of love, every question asked in genuine curiosity about how things are, every word you will speak today, before this day is [music] done, all of it descends from a sound made in Africa in the deep past >> [snorts] >> by a creature who was almost us,

who was, in the most essential sense that exists, already us. The ape spoke. [music] The world changed. And we have been talking ever since.

 

Disclaimer: This story is fictional and created for entertainment purposes only. Any names, characters, places, or events are fictitious or used fictitiously. No real person or organization is intended to be portrayed.

Recommended for You

View Archive arrow_forward