As people age, memory declines a little for everyone. That makes it hard to tell normal aging apart from early cognitive impairment that may lead to dementia. Screening tests diagnose only the state on the day of the test. Approaches based on patient data may end up relying on incomplete data if patients refuse to participate.
The voice analysis covered earlier required a separate recording task. This time, we look at the sleep records that a ring-type device worn on the finger leaves every night. The question is this: can sleep records that a wearable accumulates every day identify people with cognitive impairment on their own?
We read one paper that answers this question.
- Title: Prediction of Cognitive Impairment Using Sleep Lifelog Data and LSTM Model
- Journal: Mathematics
- Year: 2024
- DOI: 10.3390/math12203208
The method the paper proposes has two parts. The first part converts daily sleep records into bundles of 3, 4, and 5 consecutive days and feeds them into a Long Short-Term Memory (LSTM) model. This model is a neural network that learns by remembering values as they arrive in date order. The second part uses SHapley Additive exPlanations (SHAP) to calculate which sleep indicators moved the predictions.
The terms used throughout this post are as follows.
- Lifelog: Data about daily life that sensors record automatically. In this post, it means sleep records.
- Mild Cognitive Impairment (MCI): An intermediate stage between normal cognition and dementia that is not severe enough to interfere with daily life but may lead to dementia.
- Normal Cognition (NC) and Dementia (DE): The diagnostic labels in the data. The paper combines MCI and DE and treats them as a “cognitive impairment group.”
- Sequence: A learning unit that bundles several consecutive days of one person’s records into one.
Motivation: A Record from a Single Point in Time Misses Gradual Cognitive Decline
Why Early Detection of Cognitive Impairment Matters
The reasons the paper gives are as follows.
- As the elderly population grows rapidly, the number of patients with cognitive impairment has also grown.
- Each year, 5–10% of MCI patients progress to dementia. This rate is much higher than in the general population.
- Not all MCI leads to dementia. Starting medication or rehabilitation early can slow the decline or return a person to normal.
- The paper holds that early detection can prevent dementia and contribute to lower social and economic costs. However, there is no clear standard for comparing normal aging with MCI.
The most widely used screening tool is the Mini-Mental State Examination (MMSE), and in Korea the Korean version, K-MMSE, is used. The paper points out that this tool has low sensitivity and specificity for mild symptoms. It also notes that the tool diagnoses only the state on the day of the test and cannot capture usual behavior.
What Prior Studies Did Not Cover
Studies that look for cognitive impairment using data fall into two directions. The first uses data obtained directly from patients. This includes studies that applied pattern recognition to electroencephalography (EEG) and studies that distinguished patients from healthy people using acoustic features of voice signals. The paper holds that such data burden older adults and are hard to keep collecting.
The second direction uses lifelogs. One study trained an Artificial Neural Network (ANN) on activity and sleep data from a wristband to classify MCI. Another fed nursing facility sensor data into a Multilayer Perceptron (MLP) to detect abnormal behavior in dementia patients. The gaps the paper identifies are as follows.
- Lifelog studies also used only values from a single point in time as input. They did not exploit the temporal changes in data that accumulate daily.
- They focused on improving classification performance. Interpretation of which factors are related to cognitive impairment was lacking.
- A model that cannot explain its basis is hard to use directly in medical diagnosis.
Method: Bundle Daily Sleep Records into Multi-Day Sequences, Classify with an LSTM, and Interpret with SHAP
The paper carried out the study in three stages. It collects sleep lifelogs and reconstructs them as time series, classifies cognitive impairment with an LSTM, and uses explainable AI to find the sleep factors that had an influence. This post splits the first stage in two, giving four steps. Selecting the sleep indicators and labels, and building sequences and splitting training and test data, are treated separately.
Data and Participants
The data are the “Wearable Lifelog of Dementia High Risk Group,” provided by AI-Hub of the National Information Society Agency of Korea. The data can be downloaded from AI-Hub. They are sleep records left by 300 people aged 55 or older wearing a ring-type wearable. Diagnostic labels were assigned from a specialist physician’s comprehensive diagnosis.
Each person’s recording period ranges from 35 to 122 days. The data include information such as sleep duration, blood pressure, heart rate, and respiration. Missing values and outliers had already been handled by the providing organization. So the researchers did not clean the data separately.
The analysis used 174 people and 12,183 sleep records. The number of people by diagnosis is as follows.
- Normal Cognition (NC): 111 people
- Mild Cognitive Impairment (MCI): 51 people
- Dementia (DE): 12 people
The cognitive impairment group is 51+12=63 people. That is 63÷174=36% of the total, and the normal cognition group is 111÷174=64%. The criterion for using 174 of the 300 people in the analysis is not stated in the source.
Step 1: Selecting Sleep Indicators and Diagnostic Labels
- Input: Sleep records with one row per day and the diagnosis (NC, MCI, DE)
- Processing: Select 32 sleep indicators, redefine the labels into two groups
- Output: Daily data with 32 variables and a 0/1 label
- The researchers selected 32 sleep indicators. They drew on prior studies showing that cognitive impairment appears together with day-night reversal, difficulty falling asleep, and frequent nighttime awakenings.
- The indicators were divided into two groups: 26 sleep quality indicators and 6 statistical indicators.
- The labels were changed to normal (0) and cognitive impairment (1). MCI and DE were combined into one group.
There are two reasons for combining the labels. With only 12 dementia patients, splitting into three groups carried a high risk of lower performance because of imbalance. The study’s goal was also to determine whether cognitive impairment is present, not to distinguish stages.
The sleep quality indicators include the following.
- Awake, deep sleep, light sleep, and Rapid Eye Movement (REM) sleep durations (seconds)
- Sleep efficiency: total sleep time ÷ sleep time × 100
- Time taken to fall asleep, tossing and turning ratio (%), skin temperature deviation
- start1–6, end1–6: which of six segments, dividing the day into 4-hour blocks, the time of falling asleep and the time of waking fall into (0 or 1)
The statistical indicators are the average respiratory rate per minute, the mean, minimum, maximum, and median of heart rate per minute, and the mean heart rate variability (rmssd_average, milliseconds).
A row of the raw data can be seen in Table 2 of the paper. The first row is one night of an MCI patient. Bedtime start is 18:38:28 and end is 05:10:28, awake time is 8,700 seconds, and total sleep is 29,220 seconds. Daily rows like this, containing the 32 indicators and the diagnostic label, become the training data.
Step 2: Reconstruct into Multi-Day Sequences and Set Aside Test Data
- Input: The 12,183 daily records from Step 1
- Processing: Bundle consecutive 3, 4, and 5 days per person, separate the last week as test data, undersample the training data
- Output: Training sequences and test sequences for each length
- The researchers bundled each person’s records from consecutive dates into 3-day, 4-day, and 5-day units.
- The last week of each participant’s records was set aside as test data. The rest was used for training.
- The normal group is larger in the training data. So the researchers balanced the sizes of the two groups with simple undersampling, keeping only part of the normal group.
Sequences were used because cognitive impairment appears gradually. The aim is to find, in multi-day bundles, the flow of living patterns that single-day values do not reveal.
Step 3: LSTM Training and Comparison Models
- Input: The training sequences by length from Step 2
- Processing: Determine hyperparameters with grid search, train the LSTM, train 4 comparison models that do not use time series
- Output: 3 LSTMs (3-day, 4-day, 5-day) and 4 comparison models, 7 models in total
Each time the LSTM receives a day’s input, it updates the cell state, a “memory box.” Broken into four steps, the process is as follows.
- Forget gate: Decides, as a value between 0 and 1, how much to erase from the memory carried up to the previous day.
- Input gate: Creates candidates to add from today’s input and decides how much to add.
- Cell state update: Adds the retained memory and the new information to form today’s memory.
- Output gate: Decides, from today’s memory, the value to pass to the next step.
Thanks to this structure, the LSTM reduces the problem of ordinary recurrent neural networks forgetting the distant past. That is why it suits sleep records that change over time.
- The researchers ran a grid search combining 64, 128, and 256 LSTM units, 32, 64, and 128 dense layer units, and learning rates from 0.001 to 0.01.
- The final model has an LSTM layer with 128 units, a dense layer with 64 units, and an output layer for binary classification.
- The optimizer is Adam and the learning rate is 0.001. This setting gave the best balance between accuracy and validation loss.
- For comparison, they built a Support Vector Machine (SVM), Logistic Regression (LR), Random Forest (RF), and XGBoost. These models do not reflect time-series characteristics. Hyperparameters were set with the H2O AutoML library (H2O 3.46.0.1).
With the data small at 174 people, there was a risk of overfitting (a phenomenon in which a model fits only the training data and performs worse on new data). So early stopping was applied during training.
Step 4: Performance Evaluation and SHAP Interpretation
- Input: The 7 trained models and the test sequences
- Processing: Calculate classification metrics, calculate top-K precision, apply Deep SHAP to the best-performing model
- Output: Performance table by model, precision@100, importance and direction of influence by indicator
The classification metrics are defined as follows.
- Sensitivity: The proportion of actual cognitive impairment cases that the model correctly identified as cognitive impairment. TP ÷ (TP + FN)
- Specificity: The proportion of actual normal cases that the model correctly identified as normal. TN ÷ (TN + FP)
- Precision: The proportion of cases the model called cognitive impairment that are actually cognitive impairment
- F1 score: A value reflecting both precision and recall (sensitivity)
- AUC (Area Under the Curve): The area under the Receiver Operating Characteristic (ROC) curve drawn by varying the decision threshold. The closer to 1, the better
Here, TP is a case where cognitive impairment is correctly identified as cognitive impairment, and FN is a case where cognitive impairment is missed as normal. TN is a case where normal is correctly identified as normal, and FP is a case where normal is wrongly seen as cognitive impairment.
The model outputs a prediction score between 0 and 1. A score above 0.5 is judged as cognitive impairment, and below that as normal. Top-K precision (precision@K) is the proportion of actual cognitive impairment patients among the top K people after sorting by prediction score from highest to lowest.
SHAP divides each indicator’s contribution using the Shapley value from cooperative game theory. It starts from a baseline value in which all indicators are absent. It computes how much the prediction changes when indicators are added one at a time, and averages over all possible orders of addition. Adding the SHAP values of all indicators to the baseline gives the original prediction. The researchers used Deep SHAP, which is adapted for neural networks.
Variations and Scenarios
The condition the paper varied is the sequence length. The same LSTM architecture was trained on 3-day, 4-day, and 5-day bundles. Hyperparameters were also found separately for each length through grid search. How performance differed by length is covered in the results.
Results: The 5-Day Sequence LSTM Had a Sensitivity of 0.89 and an AUC of 0.92, Higher Than the Comparison Models
Comparing the LSTM with Models That Do Not Use Time Series
The researchers first compared the 7 models on sensitivity, specificity, and F1 score. When values were similar, they compared again by AUC. The two tables below reconstruct Table 5 of Hong et al. (2024), split in two to fit mobile screens. The original paper is published under a CC BY 4.0 license.
| Model | Sensitivity | Specificity | AUC |
|---|---|---|---|
| LSTM (3 days) | 0.87 | 0.74 | 0.88 |
| LSTM (4 days) | 0.89 | 0.77 | 0.91 |
| LSTM (5 days) | 0.89 | 0.80 | 0.92 |
| XGBoost | 0.68 | 0.76 | 0.81 |
| Random Forest | 0.67 | 0.77 | 0.81 |
| Logistic Regression | 0.59 | 0.60 | 0.63 |
| SVM | 0.62 | 0.59 | 0.64 |
Source: Reconstructed from Table 5 of Junhee Hong et al., “Prediction of Cognitive Impairment Using Sleep Lifelog Data and LSTM Model”, Mathematics (2024). CC BY 4.0
| Model | Accuracy | Precision | F1 |
|---|---|---|---|
| LSTM (3 days) | 0.81 | 0.77 | 0.82 |
| LSTM (4 days) | 0.83 | 0.80 | 0.84 |
| LSTM (5 days) | 0.85 | 0.82 | 0.85 |
| XGBoost | 0.72 | 0.74 | 0.71 |
| Random Forest | 0.72 | 0.75 | 0.71 |
| Logistic Regression | 0.60 | 0.60 | 0.60 |
| SVM | 0.61 | 0.60 | 0.61 |
Source: Reconstructed from Table 5 of the same paper. CC BY 4.0
The largest difference appeared in sensitivity. The 5-day LSTM is 0.89 and XGBoost is 0.68. On this test data, the LSTM’s sensitivity was higher. However, this is a single measurement on the same test data, and the design has the same people in both training and testing.
For specificity, the 5-day LSTM is 0.80, Random Forest is 0.77, and XGBoost is 0.76. The 3-day LSTM’s specificity of 0.74 is lower than both tree-based models. In the author’s view, the difference in specificity was not as large as in sensitivity.
Differences by Sequence Length
Comparing the LSTMs with each other, the metrics improved as the bundled days grew longer. AUC is 0.88 for 3 days, 0.91 for 4 days, and 0.92 for 5 days. Specificity rose to 0.74, 0.77, and 0.80. Sensitivity is the same 0.89 for 4 and 5 days.
In this test, longer lengths had higher specificity and AUC. However, each length was tested once, so it is hard to conclude that this is an effect of length itself. The paper chose the 5-day LSTM as the final model. The paper constructed only 3-day, 4-day, and 5-day sequences.
Precision for the Top 100
The researchers checked how many actual patients there were when they picked 100 people in order of highest prediction score. The LSTM’s precision@100 is 96%. Of the top 100 people, 96 actually had cognitive impairment (96÷100=96%).
This metric suits situations where people to receive a detailed examination must be prioritized. Among the comparison models, Random Forest and XGBoost found them relatively well even without time series. The exact value for each model is not given as a number in the source.
There are also examples of prediction scores. The scores of two MCI patients were 0.999999. The paper’s example table includes a case in which a normal person with a score of 0.756962 was misclassified as cognitively impaired.
Sleep Indicators That Became Signals
The researchers applied SHAP to the 5-day LSTM. It produced two outputs.
- Summary plot: Plots the distribution of SHAP values for each indicator, marking high indicator values in red and low values in blue. The further a point spreads to the right, the greater its influence in pushing toward cognitive impairment.
- Bar plot: Draws the mean of the absolute SHAP values for each indicator as a bar. This is the importance for the overall prediction.
In order of importance, the top five indicators are as follows.
- Average respiratory rate per minute
- Heart Rate Variability (HRV)
- REM sleep duration
- Deep sleep duration
- Tossing and turning ratio
The direction of influence was also confirmed. The higher the respiratory rate and the more frequent the tossing and turning, the higher the model judged the risk of cognitive impairment. The shorter the REM sleep and deep sleep and the longer the light sleep, the higher the risk. The lower the HRV, the higher the predicted risk.
The medical evidence for the top two indicators is weak. For respiratory rate, ranked first, the paper gave no separate medical evidence in its discussion of results. It only cited prior studies in the indicator selection stage as a physiological indicator of sleep quality. For HRV, ranked second, it went no further than mentioning literature that it is closely tied to sleep and overall health.
The clinical study the paper cited is indirect evidence about REM sleep. There is a study showing that red blood cells entering cerebral cortical capillaries increase greatly during REM sleep. Based on this study, the paper interpreted it as supporting the view that REM sleep deprivation accelerates cognitive decline. There is also a study linking sleep deprivation to increases in amyloid and insoluble tau protein.
One caution is needed. The paper’s summary paragraph lists respiratory rate, REM sleep, tossing and turning, and light sleep as the indicators with large influence. Its composition differs slightly from the top five in the bar plot. When reading, it is safer to use the bar plot order as the reference.
Significance and Limitations
What Changes
- The model was trained by converting daily sleep records into bundles of 3, 4, and 5 consecutive days. This differs from lifelog studies that used single-time-point values.
- It was compared with SVM, logistic regression, random forest, and XGBoost, which do not use time series, on the same data. As a result, the LSTM’s sensitivity and AUC were higher.
- Deep SHAP presented the importance and direction of influence of sleep indicators together.
- No separate test task was set. Only data that the device collected automatically (passively), without the patient’s active participation, were used.
Usefulness for Practitioners
Voice analysis had to go through the task of recording. Sleep lifelogs accumulate just by wearing the device and sleeping. The paper holds that because wearables collect data without the patient’s active participation, the burden is low and sustainability is high.
The author believes that the precision@100 result of 96% suggests the possibility of use for screening, to first narrow down candidates for detailed diagnosis. However, the same people were included in both training and testing, so generalization to new people has not been confirmed. The proportion of cognitive impairment in the test data is also not given in the source. The paper also states that it did not validate in actual medical settings.
SHAP results can serve as reference material that explains which indicators the model responded to. However, clinical validation has not been done, and what is revealed is association, not causation. What the paper presented is also a summary of importance over the whole data, not an individual-level explanation.
Usefulness for Researchers
When the sequence length and the hyperparameters matched to it were changed, AUC varied from 0.88 to 0.92. This suggests that the bundling period can be an experimental variable. However, it is hard to conclude without repeated measurements.
The data used in the analysis can be downloaded from AI-Hub. The paper states that it accessed the AI-Hub data as of May 1, 2023. The follow-up task the paper proposes is testing other models such as the Gated Recurrent Unit (GRU) or attention architectures.
Limitations the Paper States
The paper states three limitations itself.
- The sample is 174 people. The risk of overfitting is high, so early stopping was used. More data are needed to reduce imbalance and to distinguish stages of cognitive impairment.
- Only one to two months of records were used. This is short considering the period of progression from mild to severe. Using multi-year data together with brain imaging, voice, and genetic information could improve accuracy.
- It was not tested in actual medical settings. It must be validated in a clinical environment.
Conditions for Application the Author Sees
The following is not what the paper states but conditions the author adds with practical application in mind.
- The test data are the last week of each participant. The same people are in both training and testing. Whether the same performance holds for people never seen before must be rechecked by splitting the data by person.
- The data used had missing values and outliers already handled. In real monitoring, there will be nights when the device is taken off. A rule for handling broken 5-day consecutive bundles must be set first.
- The model was trained on indicators from a ring-type wearable. Other devices may calculate REM sleep or the tossing and turning ratio differently. If the device changes, it is safer to retrain.
- SHAP results are associations the model saw. They must not be read as evidence of a causal claim that respiratory rate causes cognitive impairment.
- The decision threshold of 0.5 is fixed. For screening use, the threshold should be lowered to reduce missed patients, and how much specificity drops then should be checked as well.
What to Try Right Away
- Organize the sleep records into a table with one row per day. From the wearable data you have, start by making columns for awake time, deep sleep, light sleep, REM sleep, tossing and turning ratio, average respiratory rate, and average HRV. Check how many of the paper’s 32 indicators you can obtain.
- Compare a single-time-point model and a 5-day sequence model side by side. Set aside each person’s last week as test data. Train a random forest on single-day values and an LSTM on consecutive 5-day bundles, then put sensitivity, specificity, and AUC in the same table.
- Check the top five indicators. Apply SHAP to the best-performing model and draw the bar plot. Compare how much the order overlaps with the paper’s respiratory rate, HRV, REM sleep, deep sleep, and tossing and turning.
The Business Intelligence Lab blog unpacks one paper at a time that reads technology and people through data. In the next post as well, we will follow the steps of the method and the numbers of the results together.