Normal Pressure Hydrocephalus: Reading the Gaze Behind Wrong Answers

What can gaze reveal at the moment a patient with normal pressure hydrocephalus gets a picture naming item wrong? After reading a paper on webcam eye tracking, this post summarizes the group differences that grew larger on incorrect responses and the limits of interpretation.

Translated from the Korean original. Korean original

Language assessment usually ends with “how many did the person get right?” Yet two people with the same score can reach their answers by different routes. Incorrect items are usually dropped from the score, and what happened at that moment is not recorded.

Reports suggest that patients with normal pressure hydrocephalus have difficulty with naming. If we analyze eye movements at the moment of an incorrect response, can we capture language processing characteristics that accuracy alone does not reveal?

Let’s read one paper that addresses this question.

  • Title: Language processing characteristics in normal pressure hydrocephalus: insights from eye-tracking analysis of incorrect responses
  • Journal: Frontiers in Aging Neuroscience
  • Year: 2025
  • DOI: 10.3389/fnagi.2025.1527962

The paper is open access under a CC BY license. The tables in this post are reconstructed from selected values in Tables 2, 3, and 4 of Kim et al. (2025).

The researchers asked patients with normal pressure hydrocephalus (NPH) and healthy elderly adults (Healthy Elderly, HE) to look at pictures and name them. Eye movements were recorded with a webcam during the task. The task is called the Lexical Retrieval Task (LRT). Unlike a typical naming test, the LRT scores only the first response. Even if a person first answers wrongly and then corrects themselves, the item is scored as incorrect.

The key terms used in this post are as follows.

  • Eye-tracking: a method that records where on the screen a person looks and for how long.
  • Fixation: a state in which the gaze stays at one point for 70 ms or more.
  • Saccade: a rapid movement of the gaze between fixations. In this paper, a movement faster than 30°/s was classified as a saccade.
  • Correct and incorrect responses: if the first response is right, it is a correct response. No response, a wrong answer, and a self-corrected answer are all incorrect responses.

Motivation: Accuracy Alone Cannot Show How NPH Patients Name Objects

Why Language Processing in NPH Matters

The paper gives the following reasons for examining language problems in NPH.

  • NPH is one of the reversible dementias. If it is found early, symptoms can improve.
  • In NPH, cerebrospinal fluid circulation worsens and the ventricles enlarge. The enlarged ventricles press on the occipital, temporal, and frontal lobes and the hippocampus. As a result, visual perception, lexical access, phonological encoding, and semantic memory processing may weaken.
  • Chung (2018) reviewed the records of 529 NPH patients who underwent shunt surgery. Of them, 23.1% had a language disorder. About 75% of those with a language disorder had deficits in expressive language, especially in naming objects.
  • NPH and Parkinson’s disease (PD) share major symptoms. Differential diagnosis is therefore needed to tell them apart.

Object naming proceeds through several stages. Visual recognition, semantic processing, lexical access, and phonological processing occur in sequence. People with neurocognitive disorders often have weakened vocabulary, semantics, and visual perception, and they frequently struggle with naming.

What Earlier Studies Did Not Address

Naming results in NPH patients have differed across studies. In Kim et al. (2024), the NPH and PD groups performed significantly worse on naming than healthy elderly adults. The NPH group also responded more slowly and made more semantic errors and nonword errors. By contrast, Saito et al. (2011) found no significant differences among idiopathic NPH, Alzheimer’s disease (AD), and healthy elderly adults.

Eye-tracking studies also exist. Lehtola et al. (2022) compared idiopathic NPH, AD, and healthy elderly adults on the King-Devick test, in which numbers are read quickly. Both the NPH and AD groups had smaller saccade amplitudes than healthy elderly adults. Ungrady et al. (2019) had patients with primary progressive aphasia (PPA) and AD name 200 pictures and analyzed correct and incorrect responses together.

The gaps the paper identifies are as follows.

  • Distinguishing outcome from process: One view holds that language deficits in NPH come from broader cognitive impairment rather than from the language system itself. If so, difficulty may show up more in the processing than in the final score. The differing results of earlier studies may also be due to differences in tasks.
  • Few eye-tracking studies in NPH: Eye-tracking studies of NPH are rare.
  • Analysis focused on correct answers: Earlier studies mainly looked at quantitative differences in correct responses. Few compared the gaze patterns of correct and incorrect responses qualitatively.

Method: Only the First Response Was Scored, and Gaze Was Compared Separately for Correct and Incorrect Responses

The paper’s research aims are a group comparison, the relationship between scores and metrics, a comparison by correct and incorrect responses, and gaze visualization. This post introduces the participants first and then organizes the method into five steps.

Figure 1. Five-step procedure from task design to grid analysis, comparing gaze for correct and incorrect responses.

Data and Participants

There were 26 participants in total. The NPH group consisted of 14 patients aged 65 or older, and the HE group had 12 people.

The selection criteria for the NPH group are as follows.

  • Symptoms began gradually.
  • Symptoms lasted 3–6 months or longer.
  • No other condition explained the symptoms or imaging findings.
  • The ventricles were enlarged.
  • Gait disturbance, cognitive impairment, and urinary disturbance were present.
  • Cerebrospinal fluid pressure did not exceed the diagnostic range (70–245 mmH2O).
  • The patient had not yet had shunt surgery.

A neurosurgeon and a neurologist selected the NPH participants by consensus. The HE group was defined as people with no sensory, neurological, or physical impairment and a Korean Mini-Mental State Examination-2 (K-MMSE2) score at or above the 16th percentile.

ItemNPH groupHE groupGroup difference
Mean age (years)79.1476.33p = 0.078
Mean years of education (years)7.0010.08p = 0.126
K-MMSE2 score19.9326.58p = 0.001

Source: Section 2.1 of the paper

The two groups did not differ significantly in age, sex, or years of education. Only the K-MMSE2 score differed significantly.

Step 1: Designing and Scoring the Lexical Retrieval Task

  • Input: the word list created by Kim et al. (2024)
  • Processing: select 30 words by matching frequency, syllable count, and semantic category, and score only the first response
  • Output: each participant’s LRT score, correct/incorrect marks for each item, and error types
  1. Divide the words by frequency. Following the criterion of Jo (2002), words with a frequency below 15 were classified as low frequency and those of 15 or above as high frequency. The result was 13 high-frequency and 17 low-frequency words.
  2. Balance the syllable count. There are 5 two-syllable words, 13 three-syllable words, 11 four-syllable words, and 1 five-syllable word. The share of three- and four-syllable words was raised to make the task harder.
  3. Divide the words into 15 semantic categories and distribute them in a balanced way. The categories include land animals, insects, fruits, vehicles, musical instruments, flowers, tools, and birds.
  4. Score only the first response. No cue is given even when there is no response.
  5. Classify incorrect responses by type. Following Dell et al. (1997), errors were classified as semantic, formal, mixed, unrelated, and nonword errors. These data were used in a supplementary analysis to aid interpretation.

The researchers chose the words this way to elicit errors. When many similar words exist in the same semantic category, it is easy to select the wrong word at the semantic level. When phonology is complex, errors are likely to arise at the retrieval stage.

Step 2: Gaze Recording and Preprocessing

  • Input: gaze coordinates and visual information recorded during the task
  • Processing: calibration, classification of fixations and saccades, exclusion of data outside the area of interest
  • Output: a list of fixations and a list of saccades within the area of interest
  1. Set up the equipment. Participants sat 55 cm from the monitor and 30 cm from the microphone. The monitor was a Dell E2422HS (1,920 × 1,080 pixels), and the camera was a Pengca Webcam 1080p at 30 frames per second. VisualCamp’s SeeSo was used as the eye-tracking software. The actual sampling rate was about 28.43 Hz.
  2. Perform a 5-point calibration. The participant looks in turn at points at the center and four corners of the screen. The gaze must stay within 1.7° for 5 seconds at all five points before moving on to the task. If even one point fails, calibration starts over from the beginning.
  3. Classify fixations and saccades. Gaze that stayed 70 ms or longer is a fixation, and movement faster than 30°/s is a saccade.
  4. Clean up the coordinates in fixation intervals. All coordinates within one fixation are replaced with the first coordinate of that fixation. Intermediate coordinates classified as saccades are removed from the analysis.
  5. Define the Area of Interest (AOI). Only the 1,483 × 707 pixel area at the center of the screen, where the picture appears, was analyzed. All gaze outside this area was removed to keep only the gaze that processed the picture.

Step 3: Computing Gaze Metrics

  • Input: the fixation list and saccade list from Step 2
  • Processing: compute six metrics for each item
  • Output: metric values by participant and response type

The researchers used the following six metrics.

MetricMeaningUnit
Fixation Duration (FD)Sum of fixation durationsms
Fixation Count (FC)Number of fixationscount
Saccade Count (SC)Number of saccadescount
Saccade Duration (SD)Sum of saccade durationsms
Saccade Scanpath Total (SST)Sum of saccade travel distancespixels
Saccade Amplitude Total (SAT)Sum of saccade angle changesdegrees

Source: Reconstructed from Tables 1 and 2 of Kim et al. (2025) (CC BY)

Each metric is calculated as follows.

  • FD (Fixation Duration): for each fixation, subtract the start time from the end time, and add up all the values. Only fixations of 70 ms or more are included.
  • SD (Saccade Duration): for each saccade, subtract the start time from the end time, and add them up.
  • SST (Saccade Scanpath Total): add up the straight-line distances between the starting point and the ending point of each saccade.
  • SAT (Saccade Amplitude Total): obtain each angle by applying the inverse tangent to the vertical change of a saccade divided by its horizontal change, and add up the angles.

The paper links time and frequency metrics to cognitive processing, and spatial metrics to visual perceptual ability. Accordingly, if time and frequency differ in the NPH group, this is interpreted as abnormal cognitive processing. If spatial metrics differ, it is interpreted as a visual perceptual deficit.

Step 4: Statistical Analysis

  • Input: each participant’s LRT score and gaze metrics
  • Processing: group comparison, correlation analysis, group × response type analysis
  • Output: test statistics and p values
  1. Group comparison: because the data did not satisfy the normality assumption, the Mann–Whitney U test was used to compare the two groups’ LRT scores and gaze metrics.
  2. Correlation analysis: Pearson correlations were computed between the LRT score and each gaze metric.
  3. Response type analysis: a two-way mixed ANOVA was run with group (NPH, HE) and response type (correct, incorrect) as the two factors. Bonferroni correction was applied to post hoc comparisons.
  4. Power check: post hoc power was calculated with G*Power 3.1.

Group is a between-subjects factor, with each person assigned to one group. Response type is a within-subjects factor, since each person produces both correct and incorrect responses. A mixed ANOVA tests the two main effects and the interaction at the same time.

Step 5: Visualization and Grid Analysis

  • Input: fixation data for two items with high accuracy and two items with high error rates
  • Processing: comparison of heatmaps and scanpaths, computation of 10 × 10 grid metrics
  • Output: grid metrics by item and group

The researchers selected “오토바이” (motorcycle) and “가지” (eggplant), which had accuracy above 90%, as correct items. “실로폰” (xylophone) and “에스컬레이터” (escalator), which had error rates above 60% in both groups, were selected as incorrect items. They compared the heatmaps (a figure that uses color to show where gaze lingered most) and scanpaths (the path the gaze traveled) for these four items.

To check the qualitative comparison numerically, the picture was divided into a 10 × 10 grid, that is, 100 cells. Five metrics were then computed.

  • Grid entropy: indicates how evenly fixations were spread across cells. The larger the value, the more widely the gaze was dispersed.
  • Spatial dispersion: the average straight-line distance between two consecutive fixations.
  • Grid occupancy: the proportion of cells that had at least one fixation (0–1).
  • Transition density: the proportion of distinct movements that actually occurred among the possible movements between cells.
  • Nearest Neighbor Index (NNI): below 1, fixations are clustered; above 1, they are dispersed.

Here is an example of the grid occupancy calculation. For the xylophone item, the NPH group’s occupancy was 0.72. Since there are 100 cells, this means fixations fell in 0.72×100=72 cells. The HE group’s was 0.59, which is 0.59×100=59 cells.

Results: The Gaze Differences in the NPH Group Were Larger at the Moment of Incorrect Responses

Lexical Retrieval Score: The NPH Group Was Much Lower

The mean LRT score was 52.14 (SD = 16.88) for the NPH group and 84.17 (SD = 7.26) for the HE group. The difference was significant (U = 2.500, p = 0.000). Post hoc power exceeded 90%. The authors believe this score difference should be interpreted together with the gaze results.

Overall Gaze Metrics: The NPH Group Was Larger on All Six

When all responses were pooled and the two groups compared, significant differences appeared on all six metrics.

MetricNPH meanHE meanp
Fixation Count (FC)22.6312.310.001
Fixation Duration (FD, ms)6139.763263.670.001
Saccade Count (SC)21.6511.310.001
Saccade Duration (SD, ms)4143.482018.580.000
Saccade Scanpath Total (SST, pixels)6064.022617.300.000
Saccade Amplitude Total (SAT, degrees)666.74311.280.000

Source: Reconstructed from Table 2 of Kim et al. (2025), extracting only the means and p values (CC BY)

The authors interpret the results as follows. Long fixation durations may indicate a weaker ability to shift attention elsewhere. A high fixation count means the picture was scanned less efficiently. Long saccade durations and high saccade counts suggest that visual information may have been processed piece by piece.

The authors point to compression of brain structures as the cause. The frontal lobe is involved in attention shifting, and the basal ganglia in controlling eye movements. If enlarged ventricles press on these structures, visual information processing weakens and lexical retrieval may also be affected.

When reading these results, the difference in cognitive level between the two groups must be considered as well. The authors judged that the NPH group’s cognitive function was significantly lower than that of the HE group but not severe (the discussion gives K-MMSE = 20.33). The mean K-MMSE2 in Section 2.1 was 19.93 for NPH and 26.58 for HE (p = 0.001). I therefore judge that this design cannot separate whether the gaze metric differences are specific to NPH or arise from differences in overall cognitive level. Years of education also differed, at means of 7.00 and 10.08 years (not significant).

Scores and Gaze Metrics: Negative Correlations

The LRT score was negatively correlated with all six metrics.

  • FC: r = −0.569 (p = 0.002)
  • FD: r = −0.566 (p = 0.003)
  • SC: r = −0.569 (p = 0.002)
  • SD: r = −0.554 (p = 0.003)
  • SST: r = −0.488 (p = 0.011)
  • SAT: r = −0.462 (p = 0.017)

Participants with larger gaze metric values tended to have lower naming scores. These results are correlational and do not show that eye movements cause lower scores. In the discussion, the authors interpreted the coefficients, mostly in the 0.40–0.60 range, as moderate correlations. The abstract says the coefficients are generally 0.50 or above, so the wording differs within the original paper. The actual r values were −0.462 to −0.569. Post hoc power exceeded 80% for most, but it was lower for SAT (74% or above) and SST (68% or above). The correlations for the two spatial metrics call for somewhat more caution in interpretation.

Correct and Incorrect Responses: Metrics Rose on Incorrect Responses in Both Groups

The means by response type are as follows.

Figure 2. Saccade duration rose on incorrect responses in both groups. The NPH group’s nearly tripled, a larger increase than the HE group’s.
MetricGroupCorrectIncorrect
FCNPH14.0534.23
FCHE9.9422.26
FD (ms)NPH3925.589216.46
FD (ms)HE2634.165663.05
SD (ms)NPH2249.656425.71
SD (ms)HE1658.814152.57
SST (pixels)NPH3599.139484.76
SST (pixels)HE2076.575236.22

Source: Reconstructed from Table 3 of Kim et al. (2025), extracting only the means for four metrics (CC BY)

In the two-way mixed ANOVA, the main effect of group was significant for all six metrics. For example, SD was F[1,24] = 13.600 (p = 0.001) and FD was F[1,24] = 8.281 (p = 0.008). The main effect of response type was also significant for all. SD was F[1,24] = 94.014 and FC was F[1,24] = 60.922 (both p = 0.000). This means that in both groups, gaze metrics rose on incorrect responses.

The group × response type interaction was significant only for SD (F[1,24] = 5.981, p = 0.022). This means that the increase in saccade duration from correct to incorrect responses was larger in the NPH group. Calculating the increase, it is 6425.71−2249.65=4176.06 ms for the NPH group and 4152.57−1658.81=2493.76 ms for the HE group. As a ratio, it is 6425.71÷2249.65≈2.86 times for the NPH group and 4152.57÷1658.81≈2.50 times for the HE group.

In a supplementary analysis of incorrect response types, semantic errors (e.g., tiger → lion) were the most frequent in the NPH group. Figures by error type were not presented in the main text. Based on this result, the authors infer that a deficit in visual processing, beginning with object recognition, may block semantic-level word retrieval. This inference is the authors’ hypothesis, resting on a supplementary analysis for which no figures were presented.

Visualization and Grid Analysis: Values Were Close for the Two Groups on Correct Items

In the heatmaps, the HE group looked at the center of the picture for both correct and incorrect items. The NPH group’s gaze scattered toward the edges on incorrect items. This difference was largest for the escalator picture.

Item (response)GroupSpatial dispersionGrid occupancy
Xylophone (incorrect)NPH5.200.72
Xylophone (incorrect)HE2.200.59
Escalator (incorrect)NPH4.300.60
Escalator (incorrect)HE1.350.38
Motorcycle (correct)NPH1.880.40
Motorcycle (correct)HE1.770.39
Eggplant (correct)NPH1.660.33
Eggplant (correct)HE1.510.29

Source: Reconstructed from Table 4 of Kim et al. (2025), extracting only two metrics (CC BY)

On incorrect items, the NPH group’s spatial dispersion was larger than the HE group’s, and on correct items the two groups’ values were close. However, Table 4 in the main text contains only group-level values for the four items, and no test statistics were presented (details are in Appendix 5). Grid entropy showed a similar trend. It was 5.58 and 5.45 for the incorrect item xylophone, and 4.95 and 4.86 for the correct item motorcycle.

Differences also appeared in the scanpaths. The HE group’s paths became more complex on incorrect items. The NPH group’s paths were already varied on correct items and became more complex on incorrect items. There were exceptions. Patient P04 showed a very simple path even while getting the xylophone wrong. Participants like this had low naming performance even within the NPH group. The authors interpret the simple paths seen on incorrect responses as possibly reflecting “inactive” processing rather than efficiency.

Significance and Limitations

What Changes

  • Analysis of incorrect responses: Incorrect responses were analyzed with gaze data rather than discarded. The group × response type interaction was significant for SD, and in the descriptive comparison of four items as well, group differences were larger on incorrect items.
  • Strict scoring: Only the first response was scored, and self-corrections were treated as incorrect. This targets deficits in the retrieval process rather than final success.
  • Quantification: The qualitative impression from heatmaps was quantified with five grid metrics.
  • Contrast with other conditions: In the discussion, the authors note that the PD and PPA patients in Ungrady et al. (2019) had fewer fixations on incorrect responses. The NPH group in this paper had more fixations on incorrect responses. The authors distinguish this difference as difficulty with “fixation itself” in PD versus “inefficiency of fixation” in NPH, and propose it as a candidate metric for differential diagnosis. However, the comparison groups from the same study are given as PPA and AD in the introduction and as PD and PPA in the discussion, so the wording differs within the original paper.

Practical Value for Practitioners

The authors’ proposals in the conclusion are as follows.

  • Assessment items: To assess lexical retrieval deficits in NPH patients, use multisyllabic low-frequency words.
  • Supplementary metric: Because scores and gaze metrics were correlated, the authors proposed that analysis of gaze efficiency could become a more sensitive metric. However, no comparison of sensitivity was made in this study. It should be taken as worth exploring as a supplementary metric.
  • Stepwise intervention: For patients who struggle with naming and show atypical gaze patterns, include an intermediate step that simplifies the gaze pattern before aiming for accurate naming.
  • Promoting exploration: The authors considered that patients who answer wrongly while showing an overly simple gaze path may be the most severely impaired. They proposed that visual search strategies may help such patients before retrieval training. This interpretation is a hypothesis drawn from case observations like P04, and the effect of intervention was not tested in this study.

Practical Value for Researchers

This paper presents a design that separates “outcome” from “process.” If correct and incorrect responses are set as a within-subjects factor, the interaction can test under which condition group differences grow.

The researchers recorded gaze with a webcam, a laptop, and commercial software rather than dedicated equipment. This study interpreted time and frequency metrics as cognitive processing and spatial metrics as visual perception. The paper emphasizes that gaze patterns should be interpreted in light of the task and context. Whether this interpretive framework can be applied to other conditions is worth examining with the task and context in mind (my opinion).

Limitations Stated by the Paper

The paper states three limitations of its own.

  • Sample size: The sample is small. Results need to be made more reliable with larger groups.
  • Comparison groups: Studies are needed that compare NPH with PD and AD groups, for which differential diagnosis is difficult.
  • Postoperative change: Because NPH can improve after surgery, it is necessary to track how word retrieval and visual processing change after surgery.

Conditions for Application, as I See Them

The following are conditions that I add, not content from the paper. The effect of the difference in cognitive level was addressed in the results section above.

  • Precision of gaze recording: The accuracy of webcam-based recording is about 1.7°, and the sampling rate is about 28.43 Hz. These conditions may not suit uses that require finely distinguishing short saccades.
  • Scope of visualization: The visualization and grid analysis are group-level values for four items, and no test statistics were presented. It is hard to use them directly as a criterion for individual diagnosis.
  • Basis of interpretation: The authors interpreted the SD interaction as meaning that the two groups’ SD did not differ on correct responses. However, in the descriptive statistics, the SD for correct responses was 2249.65 ms and 1658.81 ms, and the post hoc comparison figures are not in the main text.
  • Checking figures: Some values differ between the main text and the tables. The mean K-MMSE2 is 19.93 in Section 2.1 and 20.33 in the discussion. The U value for SD is 20.000 in the main text and 10.000 in Table 2. When citing, check the original directly.

What You Can Do Right Away

  • Rescore using only the first response. In naming records you already have, reclassify no responses and self-corrections as incorrect. Place the old and new scores side by side and check whose rank changes.
  • Compute means separately for correct and incorrect responses. If you have gaze records, compute the means of FC, FD, SC, and SD by group for each response type separately. Calculating the increase between the two response types for each group lets you check the trend before running an interaction analysis.
  • Compute grid occupancy. Select one item with high accuracy and one with low accuracy. Divide the picture into a 10 × 10 grid and compute, by group, the proportion of cells that received at least one fixation.

Keywords

Related posts