The earlier cognitive decline is found, the more time there is to respond. But the diagnostic methods in use today are hard to receive often. Clinical assessments take a long time, and brain imaging is expensive and requires specialized equipment. Older patients are also often reluctant to undergo long examinations.
If screening were possible from a few minutes of producing sounds such as “ah,” “ee,” and “puh-tuh-kuh,” the burden of testing would fall considerably. The question this post addresses is as follows. Can cognitive status be distinguished using only the acoustic features of sounds that Korean speakers produce in short phonation tasks?
We read one paper that answers this question.
- Title: Exploring Voice Acoustic Features Associated with Cognitive Status in Korean Speakers: A Preliminary Machine Learning Study
- Journal: Diagnostics
- Year: 2024
- DOI: 10.3390/diagnostics14242837
The paper’s method can be summarized in three steps. First, 223 patients perform eight phonation tasks, and the researchers extract acoustic features from the recordings. Next, four types of machine learning models are trained to predict cognitive decline from these features. Finally, the researchers interpret which features the models relied on. The authors call this study a preliminary study.
The terms used throughout this post are as follows.
- Acoustic feature: A number computed from a recorded waveform. It quantifies properties of the voice such as tremor, pitch, and resonance position.
- Fundamental frequency (F0): The lowest frequency in a periodically repeating voice waveform. It corresponds to the pitch of the voice.
- Jitter: A value indicating how much the period length fluctuates between adjacent vibration cycles. It represents frequency instability.
- Shimmer: A value indicating how much the amplitude (loudness) fluctuates between adjacent vibration cycles.
- Korean Mini-Mental State Examination (K-MMSE): The Korean version of a test that assesses cognitive function. It is the criterion used to divide groups in this paper.
- Precision-Recall Area Under the Curve (PR-AUC): A model performance metric. It is mainly used when group sizes are imbalanced.
- Shapley Additive Explanations (SHAP): An interpretation method that computes how much each feature contributed to a model’s prediction.
Motivation: We Need a Cheap, Low-Burden Cognitive Screening Tool
Why Voice-Based Screening Matters
The paper summarizes the barriers of existing diagnostic methods as follows.
- They are costly.
- They are hard to access.
- Patients find them uncomfortable, and older patients in particular are reluctant to undergo multi-step examinations.
Studies that apply artificial intelligence (AI) to brain imaging data such as magnetic resonance imaging (MRI) or positron emission tomography (PET) report high diagnostic accuracy. However, they require specialized equipment and facilities, which makes wide deployment difficult. Speech involves attention, memory, language formulation, and motor control together. The authors therefore regard speech analysis as a complementary tool. The goal they envision is a screening tool for regularly checking cognitive function in daily life.
Prior Research and This Paper’s Contribution
There are several prior studies on the relationship between voice and cognitive function. Some report that voice perturbation features such as jitter and shimmer are useful for detecting early cognitive decline. Other studies find that shimmer is a sensitive indicator of advanced cognitive decline.
Other studies address additional features.
- A report that a lower harmonic-to-noise ratio (HNR) is associated with early cognitive decline
- A report that prosodic features such as pitch variation and speech rate are useful for distinguishing mild from severe impairment
- A report that changes in higher formants (formant: the resonance frequencies of the vocal tract) reflect reduced articulatory precision
The authors summarize their study’s contributions in three points.
- Whereas prior studies mainly examined English speakers, this study analyzes the speech of Korean speakers.
- It extracts 184 metrics from 8 tasks and analyzes them together.
- It uses SHAP analysis to examine the relative importance of acoustic features by level of cognitive status.
Method: Extract 184 Metrics from Eight Phonation Tasks and Classify with Four Models
The paper presents the research process in four steps: collecting speech and extracting features, dividing groups by K-MMSE score, building models, and interpreting with explainable AI. In this post, the first step is split into recording and feature extraction, giving five steps. This follows the usual order of an experimental or clinical study (participants, tasks, measurement, models) with an interpretation step added.
Data and Participants
The participants are 223 Korean patients who visited a neurosurgery outpatient clinic because cognitive decline was suspected. The inclusion criteria are as follows.
- Able to understand and follow the task instructions.
- No history of other neurological disease affecting speech.
- Consented to participate in the study.
A speech-language pathologist with more than 5 years of experience screened every participant in advance. The items checked were understanding of instructions, intelligible speech, and performance of all tasks. The study was approved by the Institutional Review Board of Kyungpook National University Chilgok Hospital.
Step 1: Divide Groups by K-MMSE Score
- Input: K-MMSE scores of 223 patients
- Processing: Divide patients into three groups by score range and define three binary classification problems
- Output: 72 severe, 54 mild, 97 normal
- All patients take the K-MMSE under the observation of medical staff.
- Following existing guidelines and studies, scores of 0–19 are classified as severe cognitive decline, 20–23 as mild cognitive decline, and 24–30 as no cognitive decline (normal).
- Three classification problems are built from the three groups: severe vs. normal, mild vs. normal, and severe+mild vs. normal.
The K-MMSE assesses orientation to time and place, memory registration, attention and calculation, recall, language, and visuospatial cognition. The authors explain that these score ranges have been validated in studies of Koreans.
There were differences in age and years of education between groups. The authors did not control for these differences. In the screening stage, their judgment is that cognitive decline caused by aging may also be a precursor of dementia, so it is better to detect all of it. The mean age and mean years of education by group are as follows.
- Severe group (72 patients): mean age 72.44 years, education 7.10 years
- Mild group (54 patients): mean age 71.11 years, education 8.24 years
- Normal group (97 patients): mean age 64.80 years, education 10.93 years
Step 2: Record the Eight Phonation Tasks
- Input: 223 participants
- Processing: Record 4 vowel tasks and 4 diadochokinetic (DDK) tasks under the same conditions
- Output: wav recordings of the 8 tasks for each participant
The tasks are as follows.
- Sustained vowel phonation /α/, /i/, /u/: Hold each vowel for 2–3 seconds. This assesses phonatory stability.
- Vowel prolongation /α-α-α/: Hold /α/ as long as possible. This assesses maximum phonation time and breath control.
- Alternate motion rate (AMR) /puh-puh-puh/, /tuh-tuh-tuh/, /kuh-kuh-kuh/: Repeat a single syllable rapidly. This assesses speech motor function and coordination.
- Sequential motion rate (SMR) /puh-tuh-kuh/: Produce different syllables in sequence. This assesses motor planning and sequencing ability.
There is a reason for using both types of tasks. Vowel tasks reveal fine motor control of the phonatory organs. DDK tasks require motor planning, sequencing, and timing control, so they are sensitive to cognitive-motor impairment.
The recording conditions were fixed as follows.
- Ambient noise below 50 dB
- Standing 55 cm from the screen showing task instructions
- Unidirectional microphone placed 30 cm from the mouth
- Sampling rate of 44,100 Hz, saved in wav format
- A speech-language pathologist with more than 5 years of experience conducted all recordings
Step 3: Extract Acoustic Features
- Input: Recording files of the 8 tasks
- Processing: Compute 23 acoustic features for each task
- Output: 184 metrics per participant (23 features × 8 tasks)
The features fall into two categories.
- Voice quality features: 5 jitter measures (local, absolute, RAP, PPQ5, DDP), 6 shimmer measures (local, localdB, APQ3, APQ5, APQ11, DDA), HNR
- Prosodic features: speech duration, mean F0, standard deviation of F0 (stdevF0), mean formant, and the mean and median of F1–F4
Jitter is computed by summing the absolute differences in length between all pairs of adjacent periods, dividing by (number of periods − 1), dividing by the mean period, and multiplying by 100. Shimmer applies the same computation to amplitude instead of period length. For both, larger values mean a less stable voice.
The variant measures differ in the range of computation. APQ3, APQ5, and APQ11 average amplitude variation over ranges of 3, 5, and 11 periods, respectively. DDA shimmer averages the differences between successive amplitude differences. As shown later, this DDA shimmer emerges as a key feature in the results.
HNR is the logarithm of the ratio of harmonic energy to noise energy, multiplied by 10. A higher value means a clearer voice with less noise. Harmonics are integer multiples of F0. As in the paper’s example, if F0 is 100 Hz, the second harmonic is 200 Hz and the third harmonic is 300 Hz.
Python 3.10.14 was used for feature computation. The authors did not examine vowel-task features and DDK-task features separately but fed them together into one model. The design is intended to let the model learn even interactions between features.
Step 4: Model Training and Evaluation
- Input: Three classification problems and 184 metrics
- Processing: Train four types of models and compare them by PR-AUC
- Output: The best-performing model for each classification problem
- Prepare four types of models: logistic linear classifier (LM), random forest (RF), gradient boosting (GBM), and deep neural network (DNN).
- Find hyperparameters (model settings) with automated machine learning (AutoML).
- Reduce overfitting with 5-fold cross-validation (a method that splits the data into five parts and validates on each in turn).
- Split the data 8:2 into training and test sets, and compute performance on the test set.
There is a reason for not using recent architectures such as transformers or convolutional neural networks: the data are small and the features are already in numeric form. The numbers of participants per classification problem are 72 and 97 for severe vs. normal, 54 and 97 for mild vs. normal, and 126 and 97 for severe+mild vs. normal.
PR-AUC was chosen as the main metric because group sizes are imbalanced and the test set is small. Under these conditions, looking only at precision or recall, which depend on a particular threshold, makes judgments unstable. PR-AUC is the area under the curve of precision and recall plotted while varying the threshold.
PR-AUC has a baseline. The baseline is the number of positive samples divided by the total number of samples, and a properly trained model should exceed it. For severe vs. normal it is 72÷(72+97)=0.426. For mild vs. normal it is 54÷(54+97)=0.358. These two values match the baselines reported in the paper.
For severe+mild vs. normal, the calculation does not add up. Using the numbers from the Methods section (126 and 97), 126÷223=0.565. However, the baseline reported in the paper is 0.677, and the Results section gives the numbers as 151 and 72. Since 151÷223=0.677, the reported baseline matches the numbers in the Results section. Which of the two figures in the paper is correct cannot be determined from the evidence alone.
Step 5: Interpret Important Features with SHAP
- Input: The model with the highest PR-AUC for each classification problem
- Processing: Compute SHAP values for each feature
- Output: Top 20 important features for each classification problem
SHAP transfers the Shapley value from cooperative game theory to model interpretation. It computes how much the prediction changes when a feature is included versus excluded. This computation is repeated over all possible feature combinations and averaged.
In the result figures, red indicates high feature values and blue indicates low ones. A positive SHAP value means the model uses that feature to predict the target group, and a negative value means it uses it to predict the non-target group. The authors analyzed the top 20 features from the model with the highest PR-AUC for each classification problem.
Results: Acoustic Features Alone Outperformed the Baseline, but Age and Education Differences Are Mixed In
Group Characteristics: The Normal Group Is Younger and Has More Years of Education
Before looking at model results, the researchers compared group characteristics. Because the data were not normally distributed and variances differed, they used the Kruskal–Wallis test. For age, Levene’s test gave p = 0.017.
- Age: Differences among the three groups were significant (H = 24.07, p < 0.001).
- Years of education: Differences were significant (H = 32.83, p < 0.001).
- Sex: Differences were not significant (χ2 = 2.79, p = 0.248).
In post hoc comparisons with Bonferroni correction, age differences appeared in severe vs. normal (p < 0.001) and mild vs. normal (p = 0.003). Severe vs. mild showed no difference, at p = 0.956. Years of education follow the same pattern. Severe vs. normal and mild vs. normal were p < 0.001, and severe vs. mild was p = 0.532. This means that only the normal group is younger and has more years of education.
Model Performance Comparison: DNN on Two Problems, RF on One
The PR-AUC of the three classification problems by model is summarized below.
| Model | Severe vs. Normal | Mild vs. Normal | Severe+Mild vs. Normal |
|---|---|---|---|
| DNN | 0.737 | 0.726 | 0.659 |
| RF | 0.516 | 0.630 | 0.715 |
| GBM | 0.716 | 0.583 | 0.682 |
| LM | 0.632 | 0.597 | 0.680 |
Source: Reconstructed by extracting only the PR-AUC values from Table 5 of Lee, J. et al., “Exploring Voice Acoustic Features Associated with Cognitive Status in Korean Speakers,” Diagnostics (2024). The original paper is published under a CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) license.
By PR-AUC, DNN was highest for severe vs. normal and mild vs. normal, and RF was highest for severe+mild vs. normal. The ranking changes with other metrics. For severe vs. normal, the area under the receiver operating characteristic curve (AUC) was 0.730 for GBM, higher than 0.716 for DNN. For mild vs. normal, accuracy was 0.800 for RF and GBM and 0.667 for DNN.
The precision and recall of the model with the highest PR-AUC are also worth examining. The DNN for mild vs. normal had recall of 1.000 and precision of 0.524. The RF for severe+mild vs. normal had precision of 1.000 and recall of 0.521. The paper does not discuss this difference separately. In this writer’s view, recall, which avoids missing patients, matters in a screening tool, so from this perspective the two models serve different uses.
Improvement over Baseline: Largest for Mild vs. Normal
| Classification problem | Top PR-AUC model | PR-AUC | Baseline |
|---|---|---|---|
| Severe vs. Normal | DNN | 0.737 | 0.426 |
| Mild vs. Normal | DNN | 0.726 | 0.358 |
| Severe+Mild vs. Normal | RF | 0.715 | 0.677 |
Source: Reconstructed from Table 5 and the main text figures of Lee, J. et al., Diagnostics (2024). The original paper is published under a CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) license.
The authors reported the improvement over baseline as 73% for severe vs. normal, 103% for mild vs. normal, and 6% for severe+mild vs. normal. By calculation, 0.737÷0.426=1.73, 0.726÷0.358=2.03, and 0.715÷0.677=1.06. For the last problem, the baseline is already high, so the improvement is small.
The authors themselves note that the performance differences among the three problems are not large. Comparison with prior studies also requires caution. A study using Framingham Heart Study data reported an AUC of 0.942 with acoustic, linguistic, and demographic variables together. Another study reported that an AUC of 0.773 using only demographic variables rose to 0.812 when acoustic features were added. This paper used only acoustic features and chose PR-AUC as the main metric, so it is hard to compare the numbers directly.
Important Acoustic Features: DDA Shimmer of /i/ and stdevF0 of /puh-tuh-kuh/
The top SHAP features for each classification problem are as follows.
Severe vs. normal (DNN)
- The top feature is DDA shimmer in the /i/ task. Higher values led to a prediction of normal.
- A high mean F3 in the /tuh-tuh-tuh/ task led to a prediction of severe.
- A high DDA shimmer in the /α-α-α/ task led to a prediction of normal.
Mild vs. normal (DNN)
- The top feature is again DDA shimmer in the /i/ task. Higher values led to a prediction of normal.
- Shimmer-family features in the /kuh-kuh-kuh/, /α/, and /u/ tasks also had a large influence.
- The higher the APQ5 shimmer in the /u/ task, the more the model predicted mild.
Severe+mild vs. normal (RF)
- The top feature is stdevF0 in the /puh-tuh-kuh/ task. Low values led to a prediction of cognitive decline, and high values to normal.
- For DDA shimmer in the /i/ task, higher values led to a prediction of cognitive decline.
- For DDP jitter in the /puh-tuh-kuh/ task, higher values led to a prediction of cognitive decline.
- The median F2 in the /i/ and /u/ tasks and the median F4 in the /kuh-kuh-kuh/ task generally led to a prediction of cognitive decline when higher.
- For APQ3 shimmer in the /tuh-tuh-tuh/ and /kuh-kuh-kuh/ tasks, higher values led to a prediction of normal.
The authors note that DDA shimmer in the /i/ task was important across several classification problems. The vowel /i/ requires precise control of tongue position and keeping the vocal tract narrow. /puh-tuh-kuh/ requires rapid alternation among lip, alveolar, and velar consonants, so the cognitive-motor load is high. The authors interpret the importance of stdevF0 as possibly reflecting this load.
In this writer’s view, one point needs to be flagged. DDA shimmer in the /i/ task was important in all three problems, but the direction of its influence differs by problem. The two DNNs predicted normal when the value was higher, and the RF predicted cognitive decline when the value was higher. A high shimmer value means an unstable voice, and prior studies also report that shimmer increases with age. To use this feature as an indicator of cognitive status, the direction must therefore be rechecked first.
Significance and Limitations
What Changes
What this paper newly did is as follows.
- It attempted to classify cognitive status from the speech of 223 Korean speakers.
- It extracted 184 metrics from 8 vowel and DDK tasks and fed them together into one model.
- Using only acoustic features, without demographic variables or linguistic content, it obtained a PR-AUC above the baseline in all three problems. However, the size of the improvement differed by problem, and it was only 6% for severe+mild vs. normal. The authors also noted that age and education differences between groups may be mixed into this result.
- It separated severe from mild and compared important features by classification problem.
- It reported importance at the level of task and feature combinations. The representative combinations are DDA shimmer of /i/ and stdevF0 of /puh-tuh-kuh/.
Usefulness for Practitioners
For practitioners planning a screening service, this paper provides a concrete recording procedure. The content of the 8 tasks, the 30 cm microphone distance, ambient noise below 50 dB, and the 44,100 Hz sampling rate are all documented. If you follow these conditions as they are when collecting your own data, you can compare results with this paper.
The evaluation approach is also worth drawing on. When group sizes are imbalanced, accuracy alone makes judgment difficult. Reporting PR-AUC together with the baseline can be applied directly when evaluating other screening models as well.
Usefulness for Researchers
For researchers, there is now a list of hypotheses to test. The candidates for which features in which tasks relate to cognitive status have been narrowed. The follow-up research directions proposed by the authors are as follows.
- Age-matched groups and analyses stratified by education level
- Longitudinal studies following the same people over several years
- Analyses that statistically adjust for demographic variables
- Studies comparing multiple languages to check whether voice metrics differ by language
Limitations Stated by the Paper
The paper states six limitations itself.
- The normal group is younger and has more years of education. The important features may reflect aging or education effects rather than cognitive decline.
- Group differences in each feature were not statistically tested.
- Performance was not compared with existing screening methods.
- The time to perform all 8 tasks may be burdensome for patients.
- The technical requirements for recording and analysis need to be standardized, and the reliability and reproducibility of measurement need to be confirmed in multiple clinical settings.
- With a sample of 223, detailed analyses such as sex differences require larger studies.
For these reasons, the authors wrote that the results should be read as a proof-of-concept rather than an established biomarker. The data are part of ongoing research and are not publicly available.
Conditions for Application, in This Writer’s View
The following are not content written in the paper but conditions this writer adds with practical application in mind.
- The age and years of education of the comparison groups must be matched. Without matching, it cannot be confirmed whether the model is distinguishing younger voices from older voices.
- The size of the test set must be checked. The paper does not separately state the number of test participants for each problem. Because training and test data were split 8:2, the test set is only a fraction of each problem’s participants. At this scale, performance figures are prone to fluctuate.
- Whether the direction of important features is the same across classification problems must be checked first. Features whose direction flips, like DDA shimmer of /i/, are hard to use as screening rules.
- This paper contains no comparison with other data such as sleep or gaze. The paper only addresses the contrast with brain imaging. Whether voice is cheaper and more accurate than other data can be judged only by collecting multiple data types from the same subjects and comparing cost and performance.
- Moving to a home screening tool changes the recording environment. This study recorded in a controlled room with noise below 50 dB, so whether the same performance holds in everyday environments must be checked separately.
What to Do Next
The first steps for following this paper’s procedure with your own data are the following three.
- Compute the PR-AUC baseline for each classification problem first. Divide the number of positives by the total number. For this paper’s severe vs. normal, it is 72÷(72+97)=0.426. Always report model performance alongside this baseline.
- Right after dividing the groups, compare age, years of education, and sex. Using the Kruskal–Wallis test and Bonferroni-corrected post hoc comparisons as in the paper lets you immediately check whether only the normal group is younger.
- If it is hard to do all 8 tasks at once, start by recording just two tasks: sustained /i/ and repeated /puh-tuh-kuh/. Keep the conditions of a 30 cm microphone distance, noise below 50 dB, and 44,100 Hz, and check in which direction DDA shimmer and stdevF0 differ by group.