The hardest moment for an environmental manager at a ferroalloy electric arc furnace plant is when the nitrogen oxides (NOx) concentration in the stack suddenly surges. In Korea, a tele-monitoring device installed on the stack collects NOx as 5-minute averages. A violation is recorded when the 30-minute average exceeds the permissible limit (e.g., 60 ppm) three times in a row, or eight times in one week. A violation brings administrative sanctions such as heavy fines or suspension of operations.
If a plant acts only after the limit has been exceeded, lowering NOx is harder and requires stronger, more expensive measures. So the question is this: can we spot the signal just before a NOx spike using process data alone?
Let us read a paper that answers this question.
- Title: Early detection of NOx spikes in ferroalloy electric arc furnace plants
- Journal: Process Safety and Environmental Protection
- Year: 2026
- DOI: 10.1016/j.psep.2025.108253
The paper proposes a method for identifying the state just before a NOx spike from the operating and measurement data of an electric arc furnace (EAF). The model is a Long Short-Term Memory Autoencoder (LSTM-AE). It learns only from normal operating data. When it fails to reconstruct a new input well, that point in time is judged to be an alarm.
The terms used in this post are as follows.
- Spike: A point at which the NOx concentration exceeds an upper limit set from its usual distribution.
- Early warning state: The point in time immediately before a spike occurs.
- Reconstruction error: The difference between the original input and the value the model produces by compressing the input and expanding it again. A large difference means the model is seeing a pattern it has not seen before.
Motivation: The Moments Above the Limit Are Short and Rare, So Existing Prediction Models Miss Them
Why We Need to Know About NOx Spikes in Advance
The reasons given in the paper are as follows.
- The temperature inside an EAF can exceed 1600 ◦C. At this temperature, nitrogen and oxygen in the air react to form thermal NOx.
- NOx released into the atmosphere becomes secondary pollutants such as fine particulate matter and ozone under sunlight. These substances contribute to respiratory and cardiovascular disease and to ecosystem damage.
- The European Industrial Emissions Directive, Korea’s Clean Air Conservation Act, and the Total Air Pollution Load Management System strictly regulate NOx.
- EAF flue gas passes through a Semi-Dry Reactor (SDR), which reduces sulfur oxides, and a bag filter, which removes dust. Most NOx, however, leaves through the stack untreated.
- When raw material charging is irregular and fuel composition and combustion conditions change, NOx spikes briefly. Such spikes cause environmental damage and economic losses.
Selective Catalytic Reduction (SCR), a dedicated NOx abatement system, incurs continuing costs for catalyst replacement and ammonia injection. The paper treats early detection as a complementary measure to be used together with SCR and Selective Non-Catalytic Reduction (SNCR). The judgment is that it helps maintain regulatory compliance at low cost.
What Previous Studies Did Not Address
Most NOx research has built regression models that predict the emission concentration itself. For EAFs, one study reduced noise with a Kalman filter, predicted NOx with an LSTM, and identified key variables with SHAP (Shapley Additive Explanations). In thermal power generation and incineration as well, several prediction models have appeared that use variable selection, outlier removal, and smoothing.
These models are useful for revealing the relationship between NOx and operating variables. But the paper points out the following gaps.
- Spikes make up a very small share of the data, so they contribute almost nothing to model training.
- Preprocessing such as outlier removal and smoothing erases the spikes themselves.
- Regression models are evaluated with metrics for continuous-value prediction, such as root mean square error (RMSE) and the coefficient of determination (R2). These metrics express overall predictive ability but do not separately evaluate the ability to catch rare events in advance.
- Early warning studies exist for paper machine failures, steam turbines, sand dust, and city gas pipeline leaks. However, few studies deal directly with NOx spikes in EAFs.
- The EAF environment is harsh, which makes it hard to install and maintain sensors. It is also hard to equip high-performance sensors that could capture the subtle changes just before a spike.
Method: A Model Trained Only on the Normal State Flags Unfamiliar Patterns as Alarms
The paper describes the method in four steps. In step 1, variables highly correlated with NOx are selected. In step 2, spikes are defined using the interquartile range (IQR), early warning labels are attached to the points just before spikes, and the data are split. In step 3, an LSTM-AE is trained only on normal data, and the alarm threshold is set with a precision-recall curve. In step 4, test data whose reconstruction error exceeds the threshold are classified as early warnings, and performance is evaluated with several metrics.
Case and Data
The subject of the analysis is a ferroalloy production site in Korea. The period is 297 days, from March 10 to December 31, 2023. Data were collected as 5-minute averages from three sources.
- Electric arc furnace: 8 variables, including electrode depth, power consumption, cooling water flow rate, and dust collection duct temperature.
- Flue gas treatment equipment: 6 variables, including SDR inlet and outlet temperatures, bag filter pressure, and induced draft fan power.
- Tele-Monitoring System (TMS): 6 flue gas variables: NOx, SOx, oxygen, dust, temperature, and flow rate.
In total there are 8+6+6=20 variables. The raw data contain 58,965 records.
The researchers excluded data from July 15 to October 1 because of maintenance and temporary shutdowns. They also treated stretches in which NOx was missing or 0 for more than 6 consecutive readings as abnormal and excluded them. The remaining missing values were filled by linear interpolation between neighboring values. As a result, 53,250 records remained, and about 10 % of the raw data had been removed.
Step 1: Selecting Variables That Move with NOx
- Input: 5-minute average time series of 20 operating and measurement variables
- Processing: Pearson correlation calculation, expert review
- Output: 7 variables to feed into the model
- Calculate the Pearson correlation coefficient between each variable and NOx.
- Select variables with an absolute value greater than 0.30 as candidates.
- Field engineers, process managers, and researchers review the physical plausibility of the candidates.
The correlation coefficient takes a value from −1 to +1. The closer it is to ±1, the more the two variables move together in a straight-line relationship, and the closer to 0, the weaker the relationship. The value 0.30 follows the convention of treating 0.30 to 0.70 as a moderate correlation. The researchers chose to intuitively narrow down candidates for experts to review quickly rather than to capture every nonlinear relationship.
Six variables passed the threshold.
- Flue gas oxygen: |r| ≈ 0.86
- Flue gas temperature: |r| ≈ 0.57
- Flue gas SOx: |r| ≈ 0.49
- Power consumption: |r| ≈ 0.37
- Flue gas dust: |r| ≈ 0.34
- Electrode depth: |r| ≈ 0.32
The expert review produced the following opinions. A higher flue gas temperature speeds up combustion reactions and increases NOx. Fluctuating power input makes arc intensity and temperature fluctuate, which affects NOx. Electrode depth reflects the operating conditions inside the furnace. The final input is 7 variables, including NOx itself.
Step 2: Defining Spikes and Labeling the Points Just Before Them
- Input: NOx time series of 53,250 records
- Processing: IQR upper bound calculation, time-shift labeling, segment-based splitting
- Output: Normal and early warning labels, and training, validation, and test data
- Obtain the first quartile (Q1) and third quartile (Q3) from the entire NOx distribution.
- Calculate the upper bound (UB) as Q3 + 1.5 × (Q3 − Q1), and mark points above the upper bound as spikes.
- Attach the early warning label to the point immediately before a spike, and remove the spike point itself from the data.
- Divide the data into 20 segments and randomly assign them to training, validation, and testing.
The IQR method does not assume that the data follow a normal distribution. It is therefore suitable for skewed emission data. NOx in this dataset had a mean of 28.32 ppm and a standard deviation of 17.60 ppm, with a range of 0 to 458.95 ppm. Q1 is 16.36 ppm and Q3 is 38.66 ppm. The IQR is 38.66 − 16.36 = 22.30, and the upper bound is calculated as 38.66 + 1.5 × 22.30; the paper reports 72.10 ppm.
Data above this upper bound made up 0.25 % of the total. The ratio to normal data is about 400:1. There were 105 spikes. The average duration was about 6 minutes, the median was 5 minutes, and the longest was 25 minutes.
The researchers also compared this with the 3-sigma method, which adds three standard deviations to the mean. The threshold from this method is about 81.12 ppm, far higher than Korea’s regulatory limit of 60 ppm. The judgment is that the IQR upper bound is closer to the regulatory limit and therefore better suited for use as a field standard.
Labels were attached by time-shifting. For example, if a spike occurs at time t and the shift length is 2 steps, then t−2 and t−1 become early warnings. One step is 5 minutes. The researchers tested all shifts of 1 to 4 steps (5 to 20 minutes).
The amount of early warning data depends on the shift length. With 1 step (5 minutes), there are 53,116 normal records and 106 early warning records (0.20 %). Extending to 4 steps (20 minutes) raises early warnings to 378 records (0.71 %). Whatever the length, early warnings are under 1 %.
There is also a reason for the splitting approach. In ferroalloy processes, output and raw materials differ from day to day and month to month, so early warnings cluster in particular periods. Splitting in time order could leave almost no warnings in the test segment. Therefore, of the 20 segments, 12 were randomly assigned to training, 4 to validation, and 4 to testing. In the representative result (random seed 50), there were 31,866 training, 10,620 validation, and 10,622 test records. After labeling, min-max normalization was applied to bring the variable ranges into line.
Step 3: Training the LSTM-AE on Normal Data and Setting the Threshold
- Input: Normal data from the training and validation segments (7 variables)
- Processing: LSTM-AE training, grid search, precision-recall curve construction
- Output: Trained model and alarm threshold
- The encoder compresses a short time series into a small vector.
- The decoder reconstructs the original time series from that vector.
- The weights are adjusted so that the mean squared error becomes smaller.
- A precision-recall curve is drawn on the validation data to set the threshold.
The mean squared error is the average of the squared differences between the original and reconstructed values. A model that has learned only normal patterns reconstructs normal inputs well, so the error is small. When an unfamiliar pattern comes in, the error grows.
The model architecture is as follows.
- Encoder: An LSTM layer with 32 units, followed by a layer with 16 units, produces a fixed-length vector.
- Repeat layer: This vector is replicated to the input length.
- Decoder: An LSTM layer with 16 units, followed by a layer with 32 units.
- Output layer: Returns the 7 variables at each time step so that the shape matches the input.
The grid search ranges are as follows: input length 1 to 8, learning rate 0.001 to 0.01, number of epochs 30, 60, 100, and 200, batch size 32, 64, 128, and 256, and early stopping patience 5, 6, and 8. The optimization algorithm was set to Adam.
The reason for not using a classification model that learns the ground-truth labels directly is the imbalance. Such models tend to lean toward the majority class and easily miss rare events. Transformers and Generative Adversarial Networks (GANs) were also considered. However, 53,250 records are too few to train such models, and short inputs are more effective for spike detection, so they were excluded.
The threshold was set at the point where the precision curve and the recall curve of the validation data meet. For the model with input length 2 and shift length 1 step, this value was 0.00595. If the reconstruction error exceeds 0.00595, it is an early warning; otherwise it is normal.
Step 4: Classifying Test Data and Evaluating Performance
- Input: Test data and the threshold from step 3
- Processing: Reconstruction error calculation, confusion matrix construction, metric calculation
- Output: Precision, recall, F1 score, AUC, MCC
- Feed the test data into the model to obtain the reconstruction error.
- Compare it with the threshold to classify each point as normal or early warning.
- Match against the actual labels and tally true positives, false positives, true negatives, and false negatives.
- Calculate five metrics.
The metrics are defined as follows.
- Precision: The proportion of alarms raised that are actual early warnings.
- Recall: The proportion of actual early warnings that the model found.
- F1 score: The harmonic mean of precision and recall. It is calculated as 2 × precision × recall ÷ (precision + recall).
- AUC (Area Under the ROC Curve): The area under the receiver operating characteristic curve drawn by varying the threshold. The closer to 1, the better it separates the two classes.
- MCC (Matthews Correlation Coefficient): A metric that uses all four values of the confusion matrix, ranging from −1 to +1. It is often used for imbalanced data.
Here is a worked calculation. The final model has a precision of 0.350 and a recall of 0.212. The F1 score is 2 × 0.350 × 0.212 ÷ (0.350 + 0.212) = 0.1484 ÷ 0.562 = 0.264. The researchers made the F1 score the primary criterion because looking at precision or recall alone can distort judgment with imbalanced data.
Variations and Scenarios
Three settings change the results.
- Shift length: The longer it is, the more early warning data there are (106 records at 1 step, 378 at 4 steps). Four of the top five models in the grid search used 1 step (5 minutes).
- Threshold selection: Sites that prioritize risk prevention can weight recall, and sites that prioritize process stability can weight precision. A fixed value matched to the regulatory limit can also be used.
- Period for the spike criterion: The researchers set a single criterion from the entire distribution. Setting separate criteria by day or month was left as future work.
Results: The Combination of 7 Variables and an LSTM-AE Achieved the Highest F1 of 0.264
Comparing Shift Length and Input Length Combinations
The researchers compared the models produced by the grid search in order of F1 score. First place was the combination of shift length 1 step and input length 2 steps. It has an F1 score of 0.264, AUC of 0.768, and MCC of 0.271. Second place (input 3 steps) had an F1 of 0.259, and third place (input 1 step) had 0.250. The combination with shift length 2 steps had an F1 of 0.217 and ranked fifth.
In this dataset, the combination of shift 1 step and input 2 steps had the highest F1. It may be because the signal 5 minutes before a spike was clearer than that 10 minutes before. However, this is the author’s conjecture and not something the paper demonstrated.
Comparing How to Choose the Input Variables
The following is a comparison with different input variable configurations. The model using only NOx had an F1 of 0.217 and an AUC of 0.622. The model using all variables had an F1 of 0.235 and an AUC of 0.770. The 7-variable model had an F1 of 0.264 and an MCC of 0.271, the highest in both F1 and MCC.
By F1 and MCC, the 7-variable model came out ahead. In AUC, the all-variable model (0.770) and the 7-variable model (0.768) were similar. The all-variable model had a higher F1 than the NOx-only model but a lower one than the 7-variable model. This comparison is a result from a single split of this dataset.
Comparing with Other Models
The LSTM-AE was compared with five baseline models. The baselines are LSTM, AE (Autoencoder), VAE (Variational Autoencoder), DNN (Deep Neural Network), and XGBoost. The LSTM-AE and LSTM take multiple time steps as input, while the others take only the 7 variables at a single time step.
The key figures are as follows.
- LSTM-AE: precision 0.350, recall 0.212, F1 0.264.
- LSTM: precision 0.161, recall 0.476, F1 0.241.
- DNN: precision 0.144, recall 0.485.
- VAE: AUC 0.782, F1 0.208.
The LSTM-AE was highest in F1 score and MCC (0.271). The LSTM and DNN have high recall but low precision, so they produce many false alarms. The VAE had the highest AUC but the lowest F1 at the fixed threshold. By F1, the LSTM-AE was highest, but the differences between models were not large. The F1 of the AE (0.235), which takes only one time step as input, was also almost the same as that of the LSTM (0.241).
Comparing with Using the Regulatory Limit Directly
The easiest alternative to come to mind in the field is to raise an alarm when NOx exceeds a certain value. The researchers evaluated this approach while varying the threshold from 40 to 70 ppm.
Lowering the threshold raises recall but lowers precision. With shift 1 step and a threshold of 40 ppm, recall is 0.788 but precision is 0.011, and F1 is only 0.021. At 60 ppm, precision is 0.057, recall is 0.364, and F1 is 0.098. The combination with the highest F1 in the table was 4 steps and 60 ppm, at 0.206.
Even at its best, the F1 score of this approach is 0.21. Exceeding the regulatory limit does not mean a spike is coming, so many false alarms occur.
The paper states that the F1 score of the LSTM-AE is 6~9 % higher than this approach. The paper does not specify whether this is a percentage-point difference or a relative ratio. Numerically, it matches a percentage-point difference in F1 values. On the test data, 0.264 − 0.206 = 0.058, a difference of about 6 %p (percentage points). Comparing the F1 of 0.30 on the additional data with the highest F1 of 0.21 from the simple threshold approach gives 0.30 − 0.21 = 0.09, a difference of 9 %p (percentage points).
Stability, Additional Data, and Computational Load
When the data were re-split with random seeds of 10, 42, 44, and 50, the variation in F1 and MCC was within ±1.5 %. On data from January 1 to 24, 2024, which were not used for training, precision, recall, and F1 score were all 0.30. Based on this result, the researchers judged that the model was not overfitted. However, because these are 24 days of data from the same site, more validation is needed before generalizing.
The model has about 16,871 parameters. On a general workstation, inference took an average of 20.6 ms per record, processing about 48 records per second. Since the TMS stores values every 5 minutes, this is fast enough for real-time operation.
An F1 of 0.26 to 0.30 looks low. The researchers explain that this is because the points just before a spike are very similar to normal ones and the early warning share is only 0.20 to 0.71 %. A prior study that detected failures in advance with time-shift labels also reported an F1 of 0.10.
Significance and Limitations
What Changes
- It changed the question in NOx research from “what will the concentration be” to “is a spike coming soon.”
- It made the spikes that used to be removed as outliers the object of analysis, and labeled the points just before them.
- It set the spike criterion with the IQR and confirmed its validity by comparing with the 3-sigma method and Korea’s regulatory limit.
- It handled the imbalance of about 400:1 with a reconstruction approach that learns only from normal data.
- It showed performance comparisons against five baseline models and a simple threshold approach.
Usefulness for Practitioners
This model runs on existing operating and measurement data collected as 5-minute averages. The paper suggests that alarms can be used to adjust power input, electrode depth, and cooling water flow rate to stabilize the furnace state. The effect of these adjustments was not validated. At sites with SCR or SNCR, ammonia injection can be adjusted in real time in response to alarms. Because the model is light, it is also easy to attach to an existing control system.
Usefulness for Researchers
The design combining time-shift labels and reconstruction error can be transferred to other high-temperature processes where emissions occasionally spike, such as cement kilns, incinerators, and thermal power plants. The procedure for setting the threshold from a precision-recall curve, and the procedure for comparing a distribution-based criterion with a regulatory criterion, can be used as they are in studies of other pollutants. The paper also provides comparison material for judging what level an F1 of 0.26 to 0.30 represents in an extremely imbalanced problem.
Limitations Stated in the Paper
The paper states six limitations itself.
- Data period and scope: Only ferroalloy EAF data were used. Data over a longer period are needed to reflect seasonal and process changes, and validation is needed in processes with different alloy types, furnace operation, and raw material conditions.
- Delay effects of variables: Process variables may not affect NOx immediately, but the delay times could not be tuned precisely.
- Model architecture: It relied on a single LSTM-AE. Transformers and GANs should be tested once more data accumulate.
- Measurement interval: A 5-minute average cannot capture all the intermediate changes in fast combustion reactions. Intervals of 1 second, 30 seconds, and 1 minute should also be tested.
- Variable selection: Pearson correlation looks only at linear relationships. Techniques that capture nonlinear relationships are needed.
- Fixed threshold: The spike criterion and the alarm threshold were each fixed to a single value. Period-specific criteria and dynamic thresholds are tasks for the future.
Conditions for Application, in the Author’s View
The following are not stated in the paper but are conditions the author adds with practical application in mind.
- There must be an operating procedure that can absorb false alarms. A precision of 0.350 means that only about 35 of every 100 alarms are actual early warnings (100 × 0.350 = 35). It must be decided in advance how operators will respond to the remaining 65 (100 − 35).
- The validation segments must contain enough spike cases. Because the threshold is set from the precision-recall curve, the threshold becomes unstable if the validation data have few early warnings.
- Real operation uses past data to judge the future. Rather than expecting the performance obtained from random segment splitting as it is, it is safer to run one more test split in time order.
- When raw materials or equipment change, the IQR upper bound and the threshold must be recalculated.
What You Can Do Right Away
You can follow the first steps of the paper’s procedure with your own plant’s TMS data.
- From the 5-minute average NOx data, obtain Q1 and Q3 and calculate the upper bound as Q3 + 1.5 × (Q3 − Q1). Compare this upper bound side by side with your site’s regulatory limit.
- Tally the number of spikes above the upper bound, their average duration, and their longest duration. Check what percentage of the total the early warnings make up when the 1 step (5 minutes) before each spike is labeled as an early warning.
- For each operating variable, calculate the Pearson correlation coefficient with NOx and select those with an absolute value above 0.30. Review this list with field engineers and keep only the variables that can be explained physically.