top of page

Essentials: The Science of Learning & Speaking Languages | Dr. Eddie Chang

Aug 26
6 min read

Speaking feels effortless until illness or injury disrupts it. Yet every intelligible sentence depends on astonishing coordination: the brain must choose meaningful words, organize their sounds, direct dozens of muscles, monitor the result, and make corrections—often within fractions of a second.

In this conversation, neuroscientist and neurosurgeon Dr. Eddie Chang explains that hidden machinery and describes efforts to restore communication for people with severe paralysis. His discussion ranges from the distinction between speech and language to stuttering, neural decoding, expressive avatars, and the ethical questions surrounding human augmentation.

Speech Is a Signal; Language Carries Meaning

Chang begins by separating two ideas that are often treated as interchangeable. Speech is a physical communication signal: patterned sound produced by the body. Language is the larger system that gives those signals meaning.

Language includes semantics, or what words signify; syntax, or how words are arranged; and pragmatics, or how context affects interpretation. Speech is therefore one expression of language, not language itself. Reading and sign language also convey linguistic structure without requiring a spoken voice.

The distinction matters clinically. A person may know exactly what they want to say but be unable to produce intelligible speech. Conversely, someone may retain the physical capacity to make sounds while losing aspects of linguistic comprehension or organization. Identifying where the process has broken down is essential for understanding neurological disorders and designing assistive technology.

Chang also distinguishes speech from vocalization. Crying, laughing, groaning, and other spontaneous sounds depend on air passing through the larynx, but they do not necessarily engage the same neural systems used to construct spoken language. Some people with injuries to speech-related brain regions can still vocalize, which suggests that the brain maintains partially separate pathways for these behaviors.

How the Body Turns Breath Into Words

Speech begins with airflow. The larynx brings the vocal folds together, allowing passing air to generate vibration. This provides the raw acoustic energy that the vocal tract then transforms into recognizable consonants and vowels.

Chang notes that vocal-fold vibration commonly falls in the range of roughly 100 to 200 hertz. Anatomical differences in the size and shape of the larynx contribute to differences in voice quality and pitch; on average, male voices have a lower fundamental frequency and female voices a higher one. These are broad physiological tendencies rather than rigid categories.

Sound generated at the larynx is only the beginning. The lips, jaw, tongue, pharynx, and other structures must continuously reshape the vocal tract. Tiny changes in position can determine whether listeners hear one phoneme or another. The movements also overlap: while the mouth is finishing one sound, it is already preparing for the next.

For Chang, this makes speaking one of the most intricate motor behaviors humans perform. Fluent conversation conceals that complexity because the underlying movements become largely automatic. We usually focus on ideas and social context, not on consciously placing the tongue or timing the vocal folds.

When a Person Is Cognitively Present but Cannot Speak

The stakes become especially clear in locked-in syndrome. A brain-stem stroke, amyotrophic lateral sclerosis, or another severe neurological condition can leave cognition and awareness intact while removing nearly every reliable channel for voluntary expression.

Chang discusses a participant in the BRAVO clinical trial who had lived with profound paralysis for approximately 15 years following a car accident and brain-stem stroke. He could not move his arms or legs or produce intelligible speech. Before entering the study, he communicated slowly by making limited neck movements to select letters on a screen.

The trial’s goal was direct but technically formidable: detect activity in the cerebral cortex while the participant attempted to speak, then translate that activity into words on a computer. Surgeons placed an electrode array over cortical regions involved in controlling the lips, tongue, jaw, larynx, and broader vocal tract. A skull-mounted connection allowed the implant to transmit recorded signals to external equipment.

This approach does not read unrestricted thoughts. It focuses on neural activity associated with attempted speech—the motor patterns produced when a participant deliberately tries to say something. That distinction is central both to how the system works and to responsible discussion of its capabilities.

Decoding Attempted Speech With Machine Learning

Electrical recordings from the cortex contain complex patterns rather than neatly labeled words. Researchers must identify subtle relationships between those patterns and the speech movements a participant intends to make.

In the trial Chang describes, the recorded activity was converted into digital data and processed by machine-learning models. Training took weeks and initially centered on a vocabulary of 50 words. Although limited, that vocabulary could be combined into numerous useful sentences and provided a foundation for expansion.

Decoding is not perfect. Movements unrelated to attempted speech can alter neural activity or disturb the signal. Chang recalls that the participant reacted emotionally when the system began producing words: he shook, moved his head, and giggled. His delight was deeply human, but the movement complicated the system’s interpretation of what he tried to say next.

Language modeling can compensate for some errors. Just as phone keyboards use likely word sequences to correct mistyped letters, a speech neuroprosthesis can consider which decoded combinations make linguistic sense. This does not replace the brain signal; it helps resolve uncertainty when several possible outputs resemble the recorded pattern.

Chang places this work within a longer history of brain-machine interfaces. Earlier systems often concentrated on moving robotic arms or controlling computer cursors. Speech decoding extends the same broad principle toward a distinctly social human need: participating in conversation.

Restoring More Than Words

Communication is not merely a stream of text. Facial expression, timing, gaze, mouth movement, and vocal tone all influence what another person understands. Visible movement of the jaw and lips can improve intelligibility, while a listener’s reactions help a speaker adjust in real time.

That is why Chang’s vision extends beyond displaying decoded words. Researchers are investigating systems that connect attempted speech with animated faces capable of reproducing relevant mouth movements and expressions. A digital avatar could potentially speak on a user’s behalf while restoring some of the nonverbal information lost in text-only communication.

Embodiment may also make these systems easier to learn. When users see an avatar respond immediately to their intentions, the feedback can help them feel that they are directly controlling the result. The interface becomes less like operating a detached device and more like acquiring a new expressive channel.

This work could be particularly important in digital social spaces, where an expressive avatar may allow a person with paralysis to participate more naturally. The objective is not cosmetic realism for its own sake. It is to return agency, emotional nuance, and conversational presence.

What Stuttering Reveals About Fluency

Chang uses stuttering to illustrate how precisely the speech system must coordinate its parts. People who stutter generally possess language knowledge and know what they want to communicate. The difficulty lies in producing speech fluently, including initiating and coordinating vocal-tract movements.

Anxiety may intensify stuttering, but Chang cautions against treating anxiety as its fundamental cause. The condition appears to involve brain function and speech coordination, while its complete mechanism remains unresolved. It can also vary considerably: a person may stutter in one situation yet speak fluently in another.

That variability offers an important clue. If fluent speech remains possible at certain moments, researchers can ask what changes in the brain allow the system to coordinate successfully. Chang compares normal speech production to a symphony: numerous components must enter at the correct time, and a small disruption can affect the whole performance.

The auditory system is one crucial participant. Speaking is not a one-way process in which the brain simply issues motor commands. It also listens to the resulting voice and uses that feedback to regulate production. Changing what speakers hear can either improve or worsen stuttering, demonstrating how closely perception and articulation interact.

Early support may be especially valuable because the developing brain has strong plasticity. Speech therapy can address initiation, introduce strategies that create more favorable conditions for fluency, explore auditory feedback, and help with anxiety that has developed around speaking. Treatment is therefore less about forcing speech and more about helping the system find reliable coordination.

From Medical Restoration to Human Augmentation

Once an implanted interface can restore a lost function, a harder question follows: should similar technology enhance abilities beyond the ordinary range?

Chang observes that augmentation is not a new human impulse. People already use caffeine, nicotine, medication, education, and countless tools to change attention, endurance, memory, or performance. Neurotechnology nevertheless raises distinctive concerns because it may involve surgery, direct access to neural signals, and unequal availability.

Potential enhancements could include faster communication or improved cognition, but biological speech remains extraordinarily difficult to surpass. Human communication rests on neural systems containing millions of neurons and refined through evolution, development, and lifelong learning. Present technology cannot reproduce that full bandwidth.

If enhancement arrives, Chang suggests it may emerge through modest steps rather than a single dramatic leap. That gradual path makes questions of consent, safety, privacy, ownership, and access more urgent, not less. A technology can become socially consequential before it appears revolutionary.

The most immediate case for speech neuroprosthetics remains restorative. For someone who is fully aware but unable to express a sentence, even a small vocabulary can reopen a path to other people. The long-term ambition is richer: communication that is faster, more expressive, and sufficiently embodied to feel like one’s own.

Sources

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page