Reading Multiple Emotions Mixed in a Single Utterance from Speech

Can multiple emotions mixed within a single Korean speech clip be detected at the same time? We compare seven models built on three types of speech features, and outline the best combination and the limitations of using it in practice.

Translated from the Korean original. Korean original

A single customer-service recording can contain both an angry voice and a tired one. If an emotion recognition model assigns only “anger” to this recording, the sadness it also contains never appears in the record. Until now, most speech emotion recognition research has chosen only one emotion among many. In real situations, however, emotions often appear mixed together.

This post addresses one question: can we detect multiple emotions contained together in a single clip of Korean speech at the same time?

We read one paper that answers this question.

  • Title: Multi-Label Emotion Recognition of Korean Speech Data Using Deep Fusion Models
  • Journal: Applied Sciences
  • Year: 2024
  • DOI: 10.3390/app14177604

The paper extracts three speech features from Korean speech and builds one deep learning model for each feature. It then builds deep fusion models that combine two or three of these models, and compares the performance of all of them. Speech Emotion Recognition (SER) is the technology that automatically detects emotion from recorded speech. Multi-class classification chooses only one emotion. Multi-label classification handles both the case of a single emotion and the case of mixed emotions.

The key terms used in this post are as follows.

  • Log-mel spectrogram: a two-dimensional image that represents how frequency and amplitude change over time on a scale matched to human hearing
  • Mel-Frequency Cepstral Coefficients (MFCCs): a set of coefficients obtained by applying the discrete cosine transform to a mel spectrogram, summarizing the distribution of energy across frequency bands
  • Voice Quality Features (VQFs): 22 acoustic measures describing properties of the voice itself, such as pitch, loudness, and the fluctuation of frequency and amplitude

Motivation: Single-Emotion Classification Misses Mixed Emotions

Why Recognizing Mixed Emotions Matters

The reasons the paper gives are as follows.

  • Emotion is a key element that reveals a person’s psychological state.
  • Automatic emotion recognition systems can be used to detect mental health problems such as depression and anxiety. Both the social system and the medical system benefit from such systems.
  • Detecting emotion from brain waves, heart rate, or pulse requires dedicated equipment. Speech is subject to few constraints of time and place, so data are easy to collect.
  • In Human–Computer Interaction (HCI), automatically detecting emotion makes it possible to provide services tailored to the user’s mood.
  • In real situations, emotions often appear mixed. Multi-label classification, which detects mixed emotions, is therefore needed.

What Previous Research Did Not Cover

Most previous SER research used multi-class classification and only one speech feature. One example is a study that fed spectrograms from the German EmoDB dataset into a deep Convolutional Neural Network (CNN). Another study fed MFCCs from the English RAVDESS dataset into a Long Short-Term Memory (LSTM) network. Some studies used two or more features but mainly focused on multi-class classification. Examples are studies that used MFCCs and mel spectrograms from the IEMOCAP dataset, and MFCCs and ERB spectrograms from a Korean speech database.

There were also multi-label classification studies. One study recognized mixed emotions from utterances in the English IEMOCAP dataset. Another extracted log-mel spectrograms and MFCCs from the multilingual RML dataset and classified mixed emotions with a model combining a CNN and an LSTM. The gaps the paper points out are as follows.

  • Few studies have performed multi-label classification on Korean speech. Studies using Korean speech remained at multi-class classification.
  • Speech signals used when expressing emotion, such as speaking rate, intonation, and morphemes, differ by national culture and language. Models therefore need to be built from the speech database of each language.
  • Most studies used only one of the log-mel spectrogram or MFCCs. A single feature cannot fully reflect the complexity of speech data, which makes it hard to detect mixed emotions.

Method: Training Three Speech Features Separately and Together for Comparison

The paper describes the method in four steps. This post divides the second step into data augmentation and feature extraction and describes five steps.

Figure 1. Procedure that processes hard-labeled Korean speech in five steps and compares the performance of seven models

Case and Data

The subject of analysis is the Korean speech emotion database provided by AI-Hub of the National Information Society Agency (NIA). This database contains recordings of speakers in their 20s through 50s with no restriction on age group or gender, and includes utterances from various situations. Five experts listened to each audio file and judged its emotion, and each file carries one to five emotions. The paper uses five emotions: Happiness, Sadness, Anger, Neutral, and Disgust.

The researchers determined labels by hard labeling. In this method, an emotion is kept as a label only when at least two of the five raters chose it. This removes ambiguous emotions and leaves one or two emotions per file. Checking with the paper’s example files gives the following.

  • Audio#1: Sadness by 2 raters, Neutral by 3 raters. Both emotions have at least 2 raters, so the label is “Sadness and Neutral.”
  • Audio#2: Neutral by 1 rater, Anger by 2, Sadness by 1, Disgust by 1. The label is “Anger.”
  • Audio#3: Happiness by 3 raters, Sadness by 1, Neutral by 1. The label is “Happiness.”

The combinations “Happiness and Disgust” and “Happiness and Anger” were excluded because they have little data and are hard to express at the same time. The final data comprise 40,645 samples, with labels of 5 single-emotion types and 8 two-emotion combinations. The number of samples per label differs greatly. The largest, “Sadness,” has 15,184 samples (37.4%), and the smallest, “Neutral and Disgust,” has 266 samples (0.7%).

Step 1: Preprocessing

  • Input: audio files after hard labeling
  • Output: speech data at least 3 seconds long with silence removed

The processing order is as follows.

  1. Convert the audio files to audio data with a sampling rate (how many times per second the sound is recorded) of 16,000 Hz.
  2. Use power_to_dB from the Python librosa library to convert the power of the frequency bands to dB, and remove data whose mean is lower than −70 dB. This is because it is hard to detect mixed emotions when the sound is quiet.
  3. Trim leading and trailing silence and keep only data at least 3 seconds long. The researchers judged that speech must be at least 3 seconds long to detect mixed emotions.

About 20% of the data was removed at this step.

Step 2: Data Augmentation

  • Input: preprocessed speech data
  • Output: 3-second data with a reduced difference in sample counts across labels

Imbalanced training data degrade the performance of deep learning models. Increasing usable data through data augmentation improves performance and also avoids underfitting and overfitting. The processing order is as follows.

  1. Randomly split the preprocessed data in advance into training, validation, and test datasets at a ratio of about 6:2:2.
  2. Apply shift augmentation to all three datasets. This technique creates multiple segments by shifting the original gradually to the right along the time axis. If a 3.5-second recording is given a window of 3.0 seconds and a shift time of 0.5 seconds, two segments result.
  3. No augmentation is applied to “Sadness,” the largest label. For “Happiness and Sadness” and “Neutral and Disgust,” which have little data, a short shift time is used to create many segments.
  4. To match the model input size, the length of every sample is unified to 3 seconds.

After augmentation, the data comprise 71,340 training, 11,968 validation, and 12,849 test samples. In the training data, the label with the highest share is “Anger” with 6,354 samples (8.9%), and the lowest is “Neutral and Disgust” with 4,835 samples (6.8%). The gap, which was 37.4% versus 0.7% before augmentation, shrank considerably.

Step 3: Feature Extraction

  • Input: 3-second speech data
  • Output: input vectors of sizes (100, 300, 1), (300, 40), and (22)

The three features carry different information. The log-mel spectrogram and MFCCs capture properties in the frequency and time domains, while VQFs capture the original signal and the quality of the voice itself. The processing order is as follows.

  1. Log-mel spectrogram: Obtain a spectrogram with the Short-Time Fourier Transform (STFT), which computes frequency components for each short segment. Convert it to a mel spectrogram with a mel filterbank and take the logarithm of each value. The number of mel filters is 100, the STFT hop length is 25 ms, and the overlap length is 15 ms.
  2. MFCCs: Apply the discrete cosine transform to the mel spectrogram and extract 40 coefficients with the same STFT settings. The vector is transposed (rows and columns swapped) for input to the 1D CNN-LSTM.
  3. VQFs: Extract 22 acoustic measures from the periodic waveform of the speech.

The mel scale reflects how humans hear frequency. People are sensitive to differences in low sounds and become less sensitive as frequency rises. The formula for converting hertz (f) to mel (M) is M = 2595 log(1 + f/700).

The 22 VQFs are divided into six groups.

  • Fundamental frequency (F0): F0 mean, F0 stdev, 2 features
  • Harmonics-to-Noise Ratio (HNR): 1 feature
  • Jitter (fluctuation of frequency): Local, Local absolute, Rap, PPQ5, DDP, 5 features
  • Shimmer (fluctuation of amplitude): Local, Local dB, APQ3, APQ5, APQ11, DDA, 6 features
  • Intensity: max, mean, min, dynamic range, Intonation variation, 5 features
  • Pitch: max, mean, range, 3 features

Adding the counts by group gives 2+1+5+6+5+3=22.

Step 4: Model Construction

  • Input: three feature vectors
  • Output: probabilities of five emotions produced by seven models

Each of the three single models pairs one feature with one architecture.

  1. Vision Transformer (ViT) and log-mel spectrogram: ViT divides an image into patches (small pieces) and learns relationships among patches through self-attention. The researchers divided the spectrogram into rectangular patches of size 100×10. Since 300÷10=30, there are 30 patches, and the embedding dimension of each patch is 100×10×1=1000. The patches pass through position embedding and 6 layers of a transformer encoder, and are then classified by a Multi-Layer Perceptron (MLP) with dropout applied.
  2. 1D CNN-LSTM and MFCCs: The numbers and lengths of filters in the three 1D CNN layers are (32, 3), (64, 3), and (128, 3). After convolution and max pooling, the (300, 40) input becomes (35, 128), and the LSTM layer converts it to a (1, 15) vector. The CNN captures short-range patterns, and the LSTM learns long-range sequential information.
  3. Deep Neural Network (DNN) and VQFs: The hidden layers have 32 and 16 neurons, and the output layer has 5 neurons. GELU activation functions, batch normalization, and dropout were placed between layers.

There are two reasons for choosing rectangular patches. First, making square patches requires changing the size of the original image, whereas rectangular patches can use the original as it is. Second, rectangular patches divided along the time axis capture short-term changes in frequency and amplitude, making it easier to detect emotional change. The researchers did not use a ViT pretrained on ImageNet; they trained a model tailored to Korean SER from scratch.

The four fusion models combine the features of the single models.

  • ViT + 1D CNN-LSTM
  • ViT + DNN
  • 1D CNN-LSTM + DNN
  • ViT + 1D CNN-LSTM + DNN

The outputs of the fully connected layers of the individual single models are concatenated and finally classified by an MLP.

The output layer is configured with 5 nodes for the five emotions, and a sigmoid function is applied to each node to produce a probability for each emotion. If the probability exceeds the threshold, the emotion is considered present; if not, it is considered absent. The output can therefore be any combination from zero to five emotions. The researchers set the threshold to 0.45 rather than the commonly used 0.5, because combinations of multiple emotions came out better at 0.45 and performance was higher.

The loss function is binary focal loss. The formula is BFL(y, p̂) = −αy(1−p̂)^γ log(p̂) − (1−y)p̂^γ log(1−p̂). Here y is the ground truth (0 or 1), p̂ is the predicted probability, α is the class balance weight, and γ determines how much weight to give hard examples. For examples that are already classified well, (1−p̂)^γ becomes small and the loss decreases, and a larger loss is assigned to emotions that are often wrong. This loss function was used because augmentation alone could not fully resolve the severe imbalance of the original data.

The Adam optimizer was used for training. The search ranges were learning rate [0.0005, 0.001], epochs [50, 150], batch size [32, 256], early-stopping patience [5, 8], and γ [0.5, 2.5], and the values with the best performance within these ranges were chosen.

Step 5: Performance Validation

  • Input: per-emotion predictions on the test data (present 1, absent 0)
  • Output: per-emotion and average binary accuracy, precision, recall, and F1 score

A confusion matrix is a table that compares predictions with actual values and divides them into true positives (TP), true negatives (TN), false positives (FP), and false negatives (FN). The researchers built a confusion matrix for each emotion and calculated four metrics.

  • Binary accuracy: (TP + TN) ÷ (TP + FN + FP + TN). The proportion of cases in which both the presence and absence of an emotion were predicted correctly.
  • Precision: TP ÷ (TP + FP). Among those predicted as “present,” the proportion that are actually present.
  • Recall: TP ÷ (TP + FN). Among those actually present, the proportion correctly identified as “present.”
  • F1 score: 2 × precision × recall ÷ (precision + recall). The harmonic mean of precision and recall.

Here is a worked example. For Happiness in the ViT single model, recall is 0.423 and precision is 0.582. The F1 score is 2 × 0.582 × 0.423 ÷ (0.582 + 0.423) = 0.492 ÷ 1.005 ≈ 0.490, which matches the value in the paper’s table. Average binary accuracy is the mean of the binary accuracies of the five emotions. For the ViT and 1D CNN-LSTM fusion model, it is (0.776 + 0.684 + 0.654 + 0.686 + 0.761) ÷ 5 = 3.561 ÷ 5 ≈ 0.712.

Results: The Model Combining Log-Mel Spectrogram and MFCCs Had the Highest Average Binary Accuracy, at 71.2%

Rectangular Patches and Square Patches

Before building the single models, the researchers first compared the patch shape of ViT. The rectangular patches divided a (100, 300, 1) image into 30 patches of (100, 10, 1). The square patches divided a (128, 128, 1) image into 16 patches of (32, 32, 1).

Average binary accuracy was almost the same: 0.678 for rectangular patches and 0.677 for square patches. Average F1 score was slightly higher for rectangular patches, at 0.512, than for square patches, at 0.499. The researchers interpreted this as rectangular patches better capturing the time, frequency, and amplitude characteristics of the spectrogram, which helped recognize mixed emotions. Rectangular patches were therefore used in subsequent models.

Comparison of Single Models and Fusion Models

Seven models were compared by average binary accuracy and average F1 score.

Figure 2. The fusion model combining log-mel spectrogram and MFCCs was highest on both average metrics
ModelAverage binary accuracyAverage F1
ViT (log-mel)0.6780.512
1D CNN-LSTM (MFCCs)0.7070.499
DNN (VQFs)0.5720.389
ViT + 1D CNN-LSTM0.7120.522
ViT + DNN0.6900.465
1D CNN-LSTM + DNN0.7010.478
All three models0.7060.509

Source: Park, S.; Jeon, B.; Lee, S.; Yoon, J. (2024), Applied Sciences, Tables 8 and 9. CC BY 4.0 (license). Reconstructed by the author, extracting only the average values.

Among the single models, the 1D CNN-LSTM had the highest average binary accuracy at 70.7%, and ViT had the highest average F1 score at 51.2%. The DNN was far lower than the other two. In particular, its recall for anger was only 0.093 and its F1 score only 0.149, so VQFs did not help much in detecting anger.

The fusion model combining ViT and 1D CNN-LSTM was the highest of the seven models on both metrics. Among the four fusion models, the model combining all three came next (average binary accuracy 0.706, average F1 0.509). Across all seven models, however, this model is slightly lower than the 1D CNN-LSTM single model (average binary accuracy 0.707) and the ViT single model (average F1 0.512). Fusion models that included the DNN did better than the DNN single model but did not surpass the ViT and 1D CNN-LSTM fusion model. The researchers interpreted this as the poorly performing DNN dragging down the performance of the whole fusion model.

Per-Emotion Performance of the Best Model

The per-emotion values for the ViT and 1D CNN-LSTM fusion model are as follows.

EmotionBinary accuracyRecallF1
Happiness0.7760.4710.531
Sadness0.6840.4810.546
Anger0.6540.6210.526
Neutral0.6860.5280.554
Disgust0.7610.3680.451

Source: Park, S.; Jeon, B.; Lee, S.; Yoon, J. (2024), Applied Sciences, Table 9. CC BY 4.0 (license). Reconstructed by the author, omitting the precision column.

Precision is 0.610 for happiness, 0.631 for sadness, 0.456 for anger, 0.582 for neutral, and 0.584 for disgust. For disgust, recall is 0.368 against a precision of 0.584, the lowest recall among the five emotions.

The Difference Between Binary Accuracy and F1 Score

Even the best model had lower recall and F1 scores than binary accuracy. The researchers judged that this model has a high True Negative Rate (TNR) for all emotions and is therefore strong at reducing false positives. On the other hand, in recordings where emotion is not distinct or several emotions are mixed, it was hard to find the actual emotions without omission. The mismatch between precision and recall is a difficulty commonly seen in multi-label emotion recognition.

This post takes the paper’s table values as the basis for per-emotion judgments. Recall for anger varied greatly by model. Following the column order of Table 8 (binary accuracy, recall, precision, F1), the recall for anger in the single models is 0.448 for ViT, 0.304 for 1D CNN-LSTM, and 0.093 for DNN. In the fusion model it rose to 0.621, but precision is 0.456.

For reference, the paper’s text states that recall for anger and disgust is low in this fusion model. By the Table 9 values, this statement holds only for disgust (0.368). This discrepancy does not affect the model ranking, which is based on average values.

Returning to the initial question, detecting multiple emotions at the same time from a single clip of Korean speech is feasible. The best model correctly predicted the presence or absence of each emotion 71.2% of the time on average. However, because the average F1 score is 52.2%, the author considers this level suitable only as an aid to human judgment.

Significance and Limitations

What Changes

  • Multi-label emotion classification was performed with a Korean speech database.
  • Models using the log-mel spectrogram, MFCCs, and VQFs one at a time were compared with models combining two or three of them, on the same data.
  • The combination of log-mel spectrogram and MFCCs performed best. Combinations including the DNN did not surpass the ViT and 1D CNN-LSTM combination. ViT + DNN had an average binary accuracy of 0.690, which is even lower than the 0.707 of the 1D CNN-LSTM single model.

Usefulness for Practitioners

The researchers expected this model to be applied to mental health care, customer service, entertainment, and smart service systems. However, performance on customer-service recordings was not validated in this paper. By emotion, recall for disgust is only 0.368 and precision for anger only 0.456.

This model should therefore be regarded as something that can serve as a reference in attempts to record emotions contained together, such as “Sadness and Anger.” This dataset has 1,625 samples labeled “Sadness and Anger.”

The researchers noted that the training data were recorded from speakers in their 20s through 50s with no restriction on age or gender, and judged that the model can be applied to various users in Korean-speaking settings. However, this result comes from a single database, and performance by age and gender was not reported. Applying it to other speaker groups or other settings requires separate validation. The setting of lowering the threshold from 0.5 to 0.45 so that combinations of multiple emotions come out better is worth referring to in practice.

Usefulness for Researchers

The researchers hoped this study would serve as a comparison baseline for later research. In the author’s view, the procedure, from hard labeling, shift augmentation, and binary focal loss to threshold setting, is written with numbers, so it is worth rerunning the experiment on the same data. However, for hyperparameters such as learning rate, epochs, batch size, and γ, only ranges are given and the final selected values are not stated. The data are available from AI-Hub.

The difference by patch shape was almost nonexistent in average binary accuracy and small in average F1 score. The researchers interpreted rectangular patches as able to compensate for a lack of inductive bias on small speech datasets. However, this experiment alone does not confirm that effect. Research dealing with patch design would do well to verify this point separately.

Limitations Stated by the Paper

The paper states four limitations itself.

  • MFCCs are sensitive to background noise, so performance may drop in noisy environments. A noise reduction technique needs to be added to the preprocessing step.
  • ViT has low inductive bias and may perform worse on small datasets. It is necessary to collect more data and try rectangular patches together with a large-scale database.
  • No technique for interpreting the model’s decision process was used. Future work plans to examine the influence of features on predictions with Grad-CAM or SHAP.
  • Research on multi-label emotion recognition is scarce, and many recent studies use multimodal approaches that combine facial expressions or text with speech. The number and types of target emotions also differ across studies, making it hard to compare performance directly with previous research.

In addition, the researchers noted that only the LSTM was used, adding that testing other recurrent neural networks such as CARU or GRU could handle complex temporal patterns better.

Conditions for Application, in the Author’s View

The following is not content from the paper but conditions the author adds with practical application in mind.

  • Labels must be judged by multiple people. Hard labeling is based on agreement by at least two of five raters, so it cannot be applied as is to data with only one rater.
  • Utterance length must be checked. For counseling data with many responses shorter than 3 seconds, the proportion removed in preprocessing grows.
  • For recordings mixed with noise, such as phone counseling, noise removal must be done first.
  • Metrics should be chosen according to the purpose of use. If reducing false alarms matters, use binary accuracy and precision as the standard; if finding emotions without omission matters, use recall.
  • These results concern the five emotions in the AI-Hub data. If the emotion distribution and speaker characteristics of your own recordings differ, retraining and revalidation are needed.

What to Try Right Away

As a first step toward applying the paper’s procedure as is to your own data, you can do the following three things.

  • If you have recordings to which multiple raters assigned emotions, apply hard labeling, keeping an emotion as a label only when at least two raters chose it. Tabulate the number of samples per label to check how large the imbalance is.
  • Sample the recordings at 16,000 Hz, remove data with a mean below −70 dB and leading and trailing silence, and keep only those at least 3 seconds long. From the remaining speech, extract a log-mel spectrogram with 100 mel filters and 40 MFCCs together.
  • If you have built a model that outputs a sigmoid probability for each emotion, compute per-emotion binary accuracy, recall, and F1 score at thresholds of 0.45 and 0.5 and compare them.

Business Intelligence Lab continues research that connects unstructured data such as speech and text to practical decision-making. Interested readers are encouraged to also read the blog’s other paper reviews.

Keywords

Related posts