Oct 1, 2026

What a voice reveals about cognitive health: Two approaches to catching early signs of dementia

Image for What a voice reveals about cognitive health: Two approaches to catching early signs of dementia

What if a single phone call could flag early signs of cognitive decline?

“How have you been doing lately?”


Everyday conversations like this carry more than pleasantries. What we say and how we say it subtly reveal the state of our cognitive health.


According to the World Health Organization (WHO), 57 million people worldwide were living with dementia as of 2021. The earlier it’s detected, the better the chances that timely intervention can slow its progression. But the early signs of cognitive decline are subtle and easy to miss in daily life. When decline is suspected, the next step is usually a detailed examination such as brain imaging or, more recently, a blood test. Both require a visit to a medical facility and can be costly. What if, before it ever came to that, we could routinely check for signs of cognitive decline using nothing but a voice on the phone?


NAVER Cloud has been pursuing this question, and two papers from that work were accepted at Interspeech 2026, the leading international conference in speech technology. The two studies take different approaches, but when evaluated on the same public dataset, both surpassed the previous state of the art at the time of publication and landed on exactly the same score. Study 1 analyzes sound (acoustics) and language separately and then combines them, while Study 2 hands every clue to a single large language model (LLM). Here’s how each one reads the signs of cognitive health in a voice.


Two clues in the voice: What we say, and how we say it

When cognitive function declines, speech changes in two broad ways.


The first is linguistic change. Words don’t come easily, so people talk around them (“uh… you know, that thing”) or start a sentence and correct themselves midway (“no, that’s not it”). Vocabulary grows more repetitive and sentence structure simpler.


The second is acoustic change. Pauses between words grow longer and speech slows down. What changes isn’t the content of the speech but the way it’s delivered.


Both studies turned these two kinds of change into a form a model can read. To capture linguistic change, Study 1 had an LLM rate language patterns: lexical diversity, grammatical completeness, fluency, and how often the speaker self-corrects or hedges. Study 2, meanwhile, put numbers on the acoustic side, specifically speaking rate and pauses. Participants with cognitive decline spoke more slowly than those without (roughly 2.9 vs. 3.1 words per second). The gap between the two groups was clearest in pauses lasting 0.5 seconds or longer. Very short pauses are common for everyone, so they weren’t much help in telling the groups apart.


The challenge is that these two clues are different kinds of information. Look only at the text and the hesitations disappear. Listen only to the audio and you miss the changes in vocabulary and narrative structure. Each study proposes its own way of using both clues together.


A quick detour: What is the “picture description task”?

Both studies were evaluated on two public datasets, ADReSS and ADReSSo. These are English speech datasets released for dementia research, and they serve as benchmarks that let research teams around the world compare results. ADReSS, released in 2020, contains 156 participants, with equal numbers with and without cognitive decline and matched age and sex distributions. The design forces a model to tell the two groups apart by how they speak, not by age or sex. ADReSSo followed a year later with 237 participants and provides only the recordings, so the text has to be transcribed with speech recognition. Despite these differences, participants in both datasets look at the same picture and describe it freely: a kitchen scene where a child stands on a stool reaching for a cookie jar while the sink overflows beside them. It’s known as the “Cookie Theft” picture.



Because everyone describes the same picture, participants can be compared directly on what they mention, the language patterns they use, and how much of the picture they cover. One participant might give a well-organized account: “The boy is up on the stool getting cookies, and the water’s overflowing.” A participant with cognitive decline is more likely to hedge and self-correct (“um… this is… what was it… there’s a child…”) and to dwell on a few areas rather than surveying the whole scene. Beyond references to individual objects, the task also shows how well a speaker ties the different parts of the scene together.


Study 1: Listening between the lines—getting both an “ear” and a “feel for language” from one speech recognition model

The first paper, “Listening Between the Lines,” started from the idea of using Whisper, the automatic speech recognition (ASR) model released by OpenAI, as more than a transcription tool.


As Whisper converts speech to text, its encoder holds a wealth of information that never makes it into the final transcript: acoustic cues such as clarity of pronunciation, intonation, rhythm, and hesitation. Rather than stopping at the transcript, the study also puts these nonverbal elements to work. To do so, it designed a three-part architecture with separate acoustic and linguistic pathways whose signals are merged at the end.


1. Acoustic pathway—where to listen closely

The acoustic representations from Whisper’s encoder are processed by a sequential neural network (a BiLSTM) and then summarized with attention pooling. Attention pooling scores each small segment of the utterance on how much it matters for classification, then weights the segments by those scores and combines them into a single summary. No one tells the model which parts matter. It learns that on its own during training, which makes it a mechanism that automatically “leans in” to listen.


2. Linguistic pathway—the language signals an LLM reads

The transcript of the same utterance goes to an LLM for analysis. It extracts 46 interpretable linguistic features covering lexical diversity, grammatical completeness, semantic coherence, narrative structure, and more. The 29 most useful for classification were kept.


The study also used an LLM to automatically build an eight-topic taxonomy for analyzing picture descriptions. Earlier work relied on scoring checklists drawn up by hand decades ago. Here, the LLM looked at the picture and defined for itself the topic areas people naturally group together when they describe it. The taxonomy has seven scene-content topics, such as the child’s actions and the overflowing sink, plus one “meta-discourse” topic that captures hedging and self-correction (“um… what was it…”).


3. Gated fusion—weighing the two signals

Finally, a gated fusion network combines the acoustic and linguistic signals. For each sample, the gate decides on its own whether the acoustic or the linguistic clues are more trustworthy, and adjusts the weights accordingly.



How did it perform?

Measured by F1 score, a classification metric that balances precision and recall, the model reached 89.47% on ADReSS and 90.14% on ADReSSo. The ADReSSo result surpasses the previous best of 87.32%. On ADReSSo, acoustic information alone scored 83.08% and linguistic information alone 76.06%, so fusion lifted performance by 7.1 and 14.1 percentage points respectively. The numbers bear out the hypothesis that what we say and how we say it are complementary clues.


One finding stood out. A model that used only the statistically significant features reached just 78.82%, while one that selected features by how well they performed together reached 90.14%. A clue that’s weak on its own can become a strong signal when combined with others.


Study 2: Four views on a single card—letting an LLM make the overall call

The second paper, “LoRA-Tuned Large Language Models for Dementia Detection via Multi-View Speech-Derived Features,” started from a single question: “Could a single LLM take in all of these clues at once and make the call, without multiple models and a complex fusion architecture?”


Most approaches so far, Study 1 included, have trained an acoustic model and a text model separately and then combined their outputs. This study goes the other way. It packs four views of a single utterance into one structured prompt and hands the whole thing to an LLM.


Reading one utterance from four views

  1. Transcript with pause markers: Wherever the speaker paused for 0.5 seconds or longer, the transcript is marked with a <pause> token. In a line like “there’s a road <pause> and then…”, you can see where the speaker hesitated just by reading the text.
  2. Speech-flow statistics: Overall speech flow summarized in numbers, such as words per second, number of pauses, and average pause length.
  3. Narrative topic information: Using the eight-topic taxonomy from Study 1, each utterance is labeled by which part of the picture it describes, or whether it stays in meta-discourse (“what was it…”).
  4. Phoneme sequence: Pronunciation recorded at the level of phonemes, the smallest units of sound. A transcript tends to clean up wobbly pronunciation into the correct word. The study instead used HuPER, a phoneme recognizer inspired by the way people fill in unclear sounds from context, to recognize phonemes reliably even in disordered speech. That captures subtle changes in articulation (the movements of the mouth and tongue) that never show up at the word level.


LoRA: Turning a large model into a specialist with a light touch

These prompts are used to train open-source LLMs from the Qwen3 and Gemma3 families with an efficient technique called low-rank adaptation (LoRA). Instead of retraining the whole model, LoRA trains only a small set of add-on adapters. It’s like furnishing the rooms you need rather than remodeling the entire house. Since dementia speech datasets are small, this approach is especially effective: it preserves what the pretrained model already knows while tuning it to the task.



How did it perform?

On the ADReSSo benchmark, the model reached an F1 score of 90.14%, beating the best result reported up to publication (87.32%). That’s the same score as Study 1, and we’ll come back to what it means that two different approaches landed in the same place. What sets this study apart is that it got there in a single model pass, with no separate acoustic encoder and no complex fusion architecture.


Adding the views one at a time is just as telling. With the transcript alone, performance was 81.48%. Adding narrative topic information pushed it to 87.29%, a jump of 5.81 percentage points and the largest single contribution. Speech-flow statistics (+1.44 points) and phoneme information (+1.41 points) followed, bringing the total to 90.14%. The result shows how much a discourse-level clue (which parts of the picture an utterance covers) matters for dementia detection. Performance also improved consistently as model size grew from about 4 billion (4B) to 14 billion (14B) parameters, while even the smaller models stayed competitive with existing systems.


What the two studies show together

Two architecturally different approaches from the same research team landed on the same 90.14% on the same benchmark, ADReSSo. One is a two-branch neural network joined by gated fusion. The other packs four views into a prompt and hands it to a single LLM.


The convergence suggests that the key to performance lies less in any particular model architecture than in reading what we say and how we say it together. In both studies, performance rose consistently with each clue added, and models that used only one kind of clue fell clearly behind.


The other common thread is that LLMs can be a useful tool for speech-based cognitive health analysis. In both studies, an LLM played a central role: building a topic taxonomy to replace hand-made scoring criteria, extracting interpretable linguistic features, and synthesizing disparate clues into a single inference.


These studies do have a limitation: they were validated on public English datasets. That’s the usual order of research, though. Methods are first confirmed against a common benchmark, and validation in real-world settings follows as a separate step. Next, the team plans to extend the work to multilingual benchmarks and to the wide range of conversation topics that come up in a real service, not just picture description.


What comes next for cognitive health research

Both papers were presented at Interspeech 2026 in Sydney, Australia.


NAVER Cloud already offers CLOVA CareCall, a service in which an AI calls older adults to check in on them. This research lays the technical groundwork for those calls to grow beyond a wellness check into a safety net that can also pick up early signs of cognitive decline. The barrier to a hospital examination remains high. The goal is to spot early signs in everyday phone conversations before a hospital visit is ever needed, and to help those who need it reach specialist care sooner.


The first step is already underway. NAVER Cloud is running a clinical study with Emocog, a startup building digital solutions for dementia, at SMG-SNU Boramae Medical Center in Seoul, where cognitive assessments are conducted over the phone. Step by step, the day when a single CareCall can check on cognitive health is drawing closer.


A short check-in call becoming the first safety net for an older adult’s health: that’s the future NAVER Cloud wants to build with speech AI.


Note: The two studies described in this post are research-stage results and are not a substitute for diagnosis by a medical professional.


Papers

Listening Between the Lines: Joint Learning of ASR Embeddings and LLM-Augmented Linguistics for Dementia Detection
Olivier Jiyoun Jung, Jonghyeon Park, Myungwoo Oh
https://arxiv.org/abs/2606.30675


LoRA-Tuned Large Language Models for Dementia Detection via Multi-View Speech-Derived Features
Jonghyeon Park, Olivier Jiyoun Jung, Myungwoo Oh
https://arxiv.org/abs/2606.28445