How Should We Assess Korean and English Speech Processing?

Can administering the same tasks in Korean and English help determine at which stage a child's speech sound processing weakness lies? We summarize the development and validation of a five-task system and the results from 30 typically developing children.

Translated from the Korean original. Korean original

When a child learning English in an after-school class keeps mispronouncing English words, parents and speech-language pathologists ask the same question. Is the child making errors because they are still learning, or is there a difficulty somewhere in how they process speech sounds? Conventional articulation tests mainly record which sounds were produced incorrectly and how. Even among children who show the same errors, the strengths and weaknesses in the underlying processing abilities can differ greatly.

Can we administer tasks with the same structure in Korean and English and identify at which stage a child has difficulty: hearing, storing, or producing speech sounds?

This post reads one paper that addresses this question.

  • Title: Development and Validation of a Korean (L1)-English (L2) Speech Processing Task System for Children with and without Speech Sound Disorders
  • Journal: Folia Phoniatrica et Logopaedica
  • Year: 2026

The authors developed a computer-administered Speech Processing Task (SPT) system. The SPT assesses, stage by stage, a child’s ability to discriminate sounds, to be aware of and manipulate sound units, to store the sound forms of words in the mind, and to plan and produce sound sequences. Five types of tasks are administered in each of Korean and English, so there are ten subtasks. The authors computerized the system so that scoring and result analysis are handled automatically.

The key terms used in this post are as follows.

  • Speech Sound Disorder (SSD): persistent articulation difficulty that interferes with effective spoken communication.
  • First language (L1) and second language (L2): in this paper, L1 is Korean and L2 is English.
  • English as a Foreign Language (EFL): an environment such as Korea’s, where English is learned as a school subject rather than used as an everyday language.
  • Typically Developing (TD): in this post, the term for the peers with whom the paper compared children with speech sound disorders.

Motivation: Articulation Accuracy Alone Makes It Hard to Find Each Child’s Processing Weaknesses

Why Stage-by-Stage Assessment Matters

The paper gave the following reasons.

  • Children with SSD appear to struggle with speech in similar ways on the surface, but severity, cause, and response to treatment differ greatly from child to child.
  • In clinical practice, the cause of the speech disorder is unknown for most children. A comprehensive assessment is therefore needed to design an effective intervention plan.
  • Earlier studies of speech perception produced mixed results. One study found lower perceptual accuracy in children with SSD, while another found no group difference.
  • In contrast, studies of phonological awareness, phonological representation, and motor programming have relatively consistently found processing deficits in children with SSD.
  • In an environment where English is learned as a foreign language, the structure and pronunciation habits of the native language strongly affect English acquisition. Children with SSD are therefore likely to face greater difficulty in learning English.

Earlier research (Kim and Ha) supports this last reason. Children with SSD had lower phonological short-term memory and lower English vocabulary acquisition scores than TD children. Nonword repetition scores significantly predicted receptive vocabulary learning ability. In a follow-up study, the SSD group’s English consonant accuracy was lower than that of their peers, and Korean error patterns were observed carrying over into English pronunciation.

The theoretical framework for the assessment is the speech processing model of Stackhouse and Wells. This model divides the processing pathway into the following stages.

  • Peripheral auditory processing
  • Discrimination of speech from nonspeech sounds
  • Phonetic discrimination
  • Phonological recognition
  • Phonological representation
  • Motor programming
  • Motor planning
  • Motor execution

The model holds that administering tasks matched to each stage makes it possible to identify a child’s strengths and weaknesses and reflect them in the intervention plan.

What Previous Research Has Not Addressed

In English-speaking countries, the CTOPP (Comprehensive Test of Phonological Processing) and the PAT-2 (Phonological Awareness Test 2) are widely used. Both instruments reliably assess individual components such as phonological awareness, phonological memory, and rapid naming. In Korea, the recently developed Korean articulation and phonology profile assesses motor programming and phonological memory with a Nonword Repetition Task (NRT). The Netherlands’ CAI (Computer Articulation Instrument) is an automated assessment system for Dutch children aged 2 to 7, and includes picture naming, word and nonword repetition, and a maximum repetition rate task.

The gaps the paper identified are as follows.

  • Instruments that cover only individual subskills make it hard to examine multiple stages of speech processing at once.
  • The NRT alone involves several stages, from auditory discrimination through phonological storage to motor planning. Pinpointing a motor programming deficit therefore also requires tasks that assess the other stages.
  • Psycholinguistically based assessments have been criticized as too procedurally complex for clinical use.
  • The CAI automated administration, but score calculation is manual. It is limited to Dutch-speaking children and lacks assessments of discrimination and phonological awareness.

In light of the English learning environment as well, the authors developed an assessment system for children with SSD in two languages, Korean and English.

Methods: Five Tasks Built in Two Languages and Validated with Expert and Child Data

The paper describes its methods as three sub-studies. Study 1 developed and computerized the tasks, Study 2 verified validity and reliability, and Study 3 applied the tasks to TD children. This post separates the item development and the scoring and computerization of Study 1 and describes them in four steps. Steps 1 and 2 correspond to Study 1, Step 3 to Study 2, and Step 4 to Study 3.

Figure 1. The SPT was built in this order: task development and computerization, validity and reliability verification, and application to typically developing children.

Data and Participants

Three kinds of data were used in this paper. A preliminary study came before them.

  • Preliminary study: Discrimination, phonological representation, and nonword repetition tasks were administered in Korean and English to 10 children with SSD and 20 TD children.
  • Expert panel: Ten experts were selected to evaluate content validity. The authors followed Anderson’s recommendation that a minimum of 10 to 15 experts can reliably assess content validity.
  • Reliability data: Item-level responses were collected from 70 children aged 5 to 7.
  • Application study participants: 30 TD children aged 5 to 7, of whom 23 were boys and 7 were girls.

The authors gave two grounds for targeting ages 5 to 7. Children reach about 100% speech intelligibility around age 4, and syllable deletion tasks are generally mastered in the second half of age 5.

The selection criteria for application study participants were as follows.

  • Raw score on the Receptive and Expressive Vocabulary Test (REVT) at or above 1 SD (original notation: ≥ 1SD)
  • Score of 85 or higher on the Korean version of the Coloured Progressive Matrices (K-CPM)
  • Percentage of Correct Consonants (PCC) on the Urimal Test of Articulation and Phonology 2 (UTAP2) at or above 1 SD (original notation: ≥ 1SD)
  • No history of neurological or hearing disorders according to parent report
  • At least one year of English classes, twice a week for at least 30 minutes per session

The participants’ mean chronological age was 77.87 months (±10.63). The mean nonverbal intelligence score was 113.17 (±16.85). English articulation was additionally assessed with the DEAP (Diagnostic Evaluation of Articulation and Phonology).

Step 1: Developing the Five Tasks and Items

  • Input: preliminary study results, prior literature, the speech processing model
  • Process: set a task for each processing stage and create Korean and English items separately
  • Output: five tasks in each of the two languages, ten subtasks in total
  1. The authors created preliminary items from the preliminary study results and a literature review.
  2. They organized five tasks to match the input, storage, and output stages of the speech processing model.
  3. The English tasks were not translated from the Korean tasks. They were newly created with vocabulary and stimuli suited to 5- to 7-year-old children learning English as a foreign language.

Table 2 of the paper classifies the five tasks by processing stage and assessment domain as follows.

  • Input stage: DT (perception), IPAT (phonological recognition and representation)
  • Storage stage: PRT (phonological representation)
  • Output stage: OPAT (phonological recognition and representation), NRT (motor programming)

Response modes differ by task: judgment, picture selection, and speaking. The task composition can be found in Table 2 of the paper.

The design intent of each task is as follows.

  • Discrimination Task (DT): The child hears two identical or similar sounds and answers whether they are the same or different. Because the preliminary study found neither group differences nor language differences, the authors raised the difficulty to increase sensitivity. Korean used pairs of nonsense three-syllable items, while English changed CV monosyllabic pairs (/sa/-/si/) to CVC monosyllabic pairs (/fæs/-/fæʃ/) and multisyllabic pairs. The final 12 pairs are 6 pairs of identical sounds and 6 pairs differing in phonetic features such as fricative versus stop.
  • Input Phonological Awareness Task (IPAT): A four-choice deletion task that requires no spoken response. For example, if “apple” is removed from “pineapple,” the child selects the picture of “pine.” The choices include the correct answer (fish), a sound-alike distractor (dish), a meaning-related distractor (turtle), and a second pictured word (star).
  • Phonological Representation Task (PRT): A picture and a sound are presented together, and the child judges whether they match. Of the 12 words, 6 are the original sound and 6 are a sound with a changed vowel (/εlifɔnt/-/εlifint/).
  • Output Phonological Awareness Task (OPAT): The structure is the same as the IPAT, but the child must say the remaining syllable aloud. For example, the child sees a picture of a school bus, hears /skuːl/, and answers /bʌs/.
  • Nonword Repetition Task (NRT): Assesses motor programming. Korean has 10 items of 2 to 4 syllables, and English has 10 items of 1 to 3 syllables.

There is a reason the phonological awareness task was split into input and output versions. Children with SSD may be assessed as having lower phonological awareness than they actually do on tasks requiring a spoken response.

Step 2: Scoring and Computerization

  • Input: the child’s selected responses and recorded speech
  • Process: apply task-specific scoring rules and process automatically in a client-server system
  • Output: raw scores, scores converted to a 100-point scale, detailed NRT indicators
  1. For the four tasks other than the NRT, a correct answer earns 1 point and an incorrect one 0 points to give the raw score, which is converted to a 100-point scale.
  2. The NRT uses an arrangement score that checks whether each syllable is in its proper place. Scores are assigned by checking syllable positions in two directions: from the front (F) and from the back (B).
  3. The authors, together with a planning, design, and development team and speech-language pathology professors and clinicians, determined the program content, presentation format, input method, and scoring rules.

The body of the paper gives the example of the target nonword /haɾʌʣi/ produced as [hamɾʌʣi]. Even with the added [m], the forward direction correctly begins with [ha], so it is counted as correct. The backward direction begins with [m], so 1 point is deducted. Other response examples and scores can be found in Table 1 of the paper.

This kind of scoring reveals how accurately syllable order was planned better than simple right/wrong scoring does.

The computerized system is divided into a client environment, which administers and evaluates the tasks, and a server environment, which manages data and task resources. The assessment proceeds in the following order.

  • Entering basic participant information
  • Selecting a task
  • Practice items
  • Main task
  • End screen

The IPAT, PRT, and OPAT have a preview function. With it, the examiner first checks whether the child knows the target vocabulary. On the results screen, correct answers are shown in blue and incorrect answers in red, and each task shows a raw score and a 100-point converted score. NRT results include indicators such as the recording file, target consonant, the child’s response, syllable score, order score, number of inserted consonants, and addition error rate.

Step 3: Verifying Content Validity and Internal Consistency

  • Input: survey responses from 10 experts, item-level responses from 70 children
  • Process: calculate the content validity ratio and index, and Cronbach’s alpha
  • Output: validity values by task and item, reliability coefficients by language
  1. The authors created a three-part survey using Google Forms.
    • Part 1: gender, age, education, clinical and research experience
    • Part 2: questions asking whether each task is appropriate for assessing the corresponding processing stage
    • Part 3: rating of 48 items on a 5-point scale
  2. From the responses, they calculated the Content Validity Ratio (CVR) and the Item-level Content Validity Index (I-CVI).
  3. From the responses of the 70 children, they calculated Cronbach’s α. It was computed for three sets: Korean items, English items, and the two languages combined.

The meaning of each indicator and the criteria for judging it are as follows.

  • CVR: indicates whether the experts reached sufficient agreement. Following Ayre and Scally’s panel-size criteria, a value of 0.80 or higher was taken as sufficient agreement.
  • I-CVI: the degree of expert agreement on whether an item fits the assessment purpose. For a panel of 10, the minimum acceptable value is 0.78.
  • Cronbach’s alpha: a value indicating how consistently the items measure the same construct. It ranges from 0 to 1, and a value above 0.70 is generally regarded as acceptable reliability.

Step 4: Application to Typically Developing Children

  • Input: performance data on the ten subtasks from 30 TD children
  • Process: administer in a fixed order and apply a 2×5 repeated-measures ANOVA
  • Output: main effects of language and task type, the interaction, and post hoc comparison results
  1. Participants first underwent screening tests, and then performed the tasks in the order DT, IPAT, PRT, OPAT, NRT.
  2. To aid comprehension, for each task the Korean version was administered first and the English version afterward.
  3. Before the IPAT, PRT, and OPAT, a vocabulary check was done to confirm that the child knew all the target words.
  4. All auditory stimuli were played at most twice.
  5. For the NRT, audio and video were recorded together. Two trained transcribers each transcribed independently, and disagreements were resolved by consensus.
  6. A repeated-measures Analysis of Variance (ANOVA) with two levels of language (Korean, English) and five levels of task was conducted.

When the same children perform all ten tasks, repeated-measures ANOVA separately tests whether score differences are due to language, to task, or to the combination of the two. Because the sphericity assumption was not met, the authors applied the Greenhouse-Geisser correction.

Results: Native-Language Advantage Was Clear, with DT the Only Exception

Content Validity: All Five Tasks Exceeded the Criteria

In the evaluation by the 10 experts, all five tasks had a CVR of 0.80 or higher and a CVI of 0.90 or higher. At the item level as well, every item had a CVR of 0.80 or higher and a CVI of 0.78 or higher, so no items needed to be removed. Values by task can be found in Table 2 of the paper.

Only the NRT, with a CVR of 0.80 and a CVI of 0.90, was lower than the other tasks. Even so, the NRT values did not fall below the CVR criterion of 0.80 or the I-CVI minimum of 0.78. The expert ratings were highest for the OPAT and IPAT at 4.70.

Internal Consistency: All Three Sets Exceeded 0.70

Cronbach’s alpha values calculated from the responses of the 70 children are as follows.

  • Korean items: 0.779
  • English items: 0.886
  • Items from both languages combined: 0.912

All three values exceeded 0.70, confirming the internal consistency of the items. It is worth noting that the coefficient for the Korean items is lower than that for the English items. The paper this post draws on offers no interpretation of this difference.

Baseline Assessment: English Consonant Accuracy Was Lower Than Korean

In the articulation test given before the tasks, a difference between the two languages was already evident. Mean consonant accuracy was 99.72% (±1.12) in Korean and 97.45% (±3.90) in English. A paired-samples t test showed that Korean consonant accuracy was significantly higher than English (t(29) = 3.734, p < 0.01). The mean receptive vocabulary score was 75.23 (±16.52), and the expressive vocabulary score was 79.13 (±14.48).

Performance by Language and Task: Korean Was Higher on Four of the Five Tasks

On the four tasks other than the DT, Korean scores were higher than English. Among the Korean tasks, the highest score was the PRT (93.60), and among the English tasks, the DT (93.40) was highest. In both languages the lowest score was the NRT.

Figure 2. Typically developing children scored higher in Korean than English on four tasks, all except the discrimination task.

The size of the language difference, calculated for the NRT and DT, is as follows.

  • NRT: Korean 86.18 points, English 78.55 points, a difference of 86.18−78.55=7.63 points.
  • DT: In the opposite direction, English was higher by 93.40−89.33=4.07 points.

The means and standard deviations by language for the remaining tasks are in Table 4 of the paper. The IPAT and OPAT had larger standard deviations, 17.88 to 25.73, than the other tasks. The author reads these values as meaning that individual differences were large on the phonological awareness tasks even among TD children. This interpretation is the author’s own and is not in the paper.

Repeated-Measures ANOVA: Language, Task, and the Interaction Were All Significant

The results of the analysis are as follows.

  • Main effect of language: F(1, 29) = 10.601, p < 0.01
  • Main effect of task type: F(1.999, 57.985) = 4.420, p < 0.05
  • Language × task interaction: F(2.628, 76.199) = 4.674, p < 0.01

In post hoc comparisons between tasks, scores on the DT, PRT, and OPAT were significantly higher than on the IPAT and NRT. The five tasks thus split into three easier tasks and two harder ones. The interaction was analyzed further with the COMPARE command. The results showed that on the DT, English was significantly higher than Korean (p < 0.05), and on the other four tasks Korean was higher than English.

The Paper’s Interpretation: Native-Language Advantage, Native-Language Interference, and Task Complexity

The authors interpreted the results under three themes.

  • Native-language advantage: They held that well-established Korean phonological representations make processing more efficient at the perception, storage, and output stages. They gave two reasons for the DT being the exception. The DT involves early perceptual processing rather than lexical access or motor programming, so the influence of native-language representations is small. In addition, the Korean stimuli were all three syllables, while the English stimuli were one to three syllables, so some English items may have been easier.
  • Native-language interference: On the English NRT, some children produced /ʧoʊvæg/ in a form closer to a Korean word, such as [ʧoʊbæg]. The interpretation is that they processed unfamiliar English sound sequences by relying on Korean phonological patterns.
  • Task complexity: The NRT involves auditory perception, phonological encoding, phonological memory, motor programming, motor planning, and articulation. The IPAT goes through multiple stages: accessing phonological representations through visual cues, retrieving the target sound, and manipulating it silently in the mind.

Distinguishing Second-Language Differences from Disorder: The Author’s Interpretation

The main study in this paper targeted only TD children. Direct comparisons with a disorder group therefore exist only in the preliminary study. In the preliminary study, the SSD group’s speech production ability was significantly lower than the TD group’s in both languages. By task, the results were as follows.

  • DT: There was no group difference.
  • PRT: The language difference was significant.
  • NRT: Both the language difference and the group difference were significant.

What follows is the author’s interpretation, not a conclusion of the paper, and it was not tested in the paper. In the main study, even TD children who had studied English for at least a year scored lower in English than in Korean on the NRT (86.18−78.55=7.63 points). The author thinks a language difference of this size may be a difference commonly seen in the course of learning English. A case like that of the preliminary study, where a group difference appears on the same task, might serve as a clue for distinguishing language learning differences from a disorder.

However, whether the SPT can be used for such a distinction requires follow-up research. The paper states that it did not compare a disorder group in the main study and could not sufficiently evaluate the effect of the SPT in children with SSD. The preliminary study also has a small sample, with 10 children with SSD and 20 TD children.

The authors also presented examples of interpreting score combinations across tasks. If IPAT and PRT scores are high but only the OPAT is low, the difficulty may lie in the speech output stage. If the IPAT is lower than the OPAT, the PRT score can be used to judge whether the problem lies in phonological awareness or phonological representation.

Significance and Limitations

What Changes

  • It created a system that assesses the input, storage, and output stages of the speech processing model together with five tasks.
  • It administers tasks with the same structure side by side in Korean and in English, a foreign language. The English tasks are not translations but were newly created for 5- to 7-year-old children learning English.
  • By splitting phonological awareness into a no-spoken-response input version and a spoken-response output version, it sought to reduce the problem of underestimating phonological awareness in children with SSD.
  • By scoring the NRT with a syllable arrangement score, it assesses syllable order planning ability more precisely.
  • Administration, scoring, and result display are automated. This function may reduce the scoring burden, but the actual effect has not been confirmed. The paper presents no measurement of how much administration time or effort was actually reduced.

Usefulness for Practitioners

Speech-language pathologists and teachers may be able to use the SPT to organize a child’s strengths and weaknesses into a stage-by-stage profile. However, the paper applied the tasks only to TD children and could not confirm criterion-related validity, so validity in clinical groups has not been confirmed. When working with a child who makes many English pronunciation errors, one can make an exploratory comparison of the gap between their Korean and English scores with the gaps of the 30 TD children. This comparison is only a rough reference and is not a criterion for judging whether a disorder is present. The results screen automatically organizes correctness and detailed indicators. This automatic scoring may reduce the burden of manual scoring, but the paper did not measure actual time savings.

Usefulness for Researchers

The three-stage procedure of task development, expert validity verification, and application to children can serve as a reference when building assessment instruments for other language pairs. The process of adjusting the difficulty of the DT, which showed no differences in the preliminary study, is an example of revising items to increase sensitivity. The author thinks the values in Table 4 of the paper can be used as a reference in later studies comparing with children with SSD. The paper did not propose this use.

Limitations Stated by the Paper

The paper itself states three limitations.

  • Because the main study sample consisted only of TD children, the effect of the SPT in children with SSD could not be sufficiently evaluated. The authors plan a follow-up clinical study including children currently diagnosed with SSD and children with a past history.
  • Because there is no comparable standardized Korean instrument for assessing English speech processing, criterion-related validity could not be confirmed.
  • Information on participants’ socioeconomic, cultural, and linguistic backgrounds and bilingual exposure was not collected.

Conditions for Application, According to the Author

What follows is not stated in the paper. It is the conditions the author adds with practical application in mind.

  • The target age and English experience must match. The tasks and reference scores were built on 5- to 7-year-old children who had had at least one year of English classes.
  • It is too early to use the means in Table 4 as norms. It is safer to use the means and standard deviations from 30 children for rough comparison and not as a criterion for determining whether an individual child has a disorder.
  • The administration order must be kept for comparisons to be meaningful. The main study always administered the Korean version first, so administering the English version first may change the scores.
  • Cross-language comparisons on the DT should be read with caution. The syllable structures of the two languages’ stimuli differ, so a higher English score is hard to interpret as better English perception.

What to Try Now

  • If you already have assessment data on children, place the same child’s Korean and English scores side by side by task, calculate the differences, and compare them exploratorily with the differences in Table 4 of the paper. Separate tasks like the NRT, where English is lower even in TD children, from tasks like the DT, where English is not lower.
  • If you have recorded nonword repetition responses, rescore them with the forward and backward arrangement scoring rules. As in the example in Table 1 of the paper, the first step is to check how omissions and additions are reflected in the score.
  • If you have administered phonological awareness tasks only in a spoken-response format, also administer a few picture-selection items. Then use the combination of the IPAT, PRT, and OPAT scores to judge whether the difficulty lies at the output stage or the representation stage.

Keywords

Related posts