Policy departments and corporate strategy teams must read the articles and posts that pour in every day and judge what society is paying attention to right now. If they read topics in descending order of article count, the top of the list is usually occupied by old problems that everyone already knows. What the person in charge really wants to know is which topic suddenly drew attention this week, and when.
So can we analyze the connections among words in web-based text as a network and systematically find the social issues that people suddenly noticed, along with when they emerged?
This post reads a paper that answers this question.
- Title: A Network Analysis Approach to Detecting Social Issues with Web-Based Data
- Journal: Applied Sciences
- Year: 2023
- DOI: 10.3390/app13148516
The paper proposes a method that builds keyword networks from web data to detect social issues. The method does three things.
- It finds the period when an issue emerged from changes in network structure.
- It extracts issue candidates from the network of that period.
- It ranks the candidates by urgency.
The key terms used in this post are as follows.
- Keyword co-occurrence network: Keywords are represented as points (nodes), and two keywords that appear together are joined by a line (link). The weight of a link is the number of times the two keywords appeared together.
- Degree centrality: The number of connections attached to a node.
- Betweenness centrality: The number of times a node or link lies on the shortest paths between all pairs of nodes.
- Network structure entropy: A value showing whether a network is structurally ordered or disordered.
- Community: A group of nodes that are densely connected to one another and loosely connected to the outside.
- Impact score: A score showing how much an issue candidate contributed to the change in structure entropy.
Motivation: It Is Hard to Say in Numbers When an Issue Emerged and What Is Urgent
Why Social Issue Detection Matters
The paper defines an issue as a specific topic or problem that forms from the narratives of many individuals and groups and gradually gathers attention. The paper gives three reasons why detecting social issues is important.
- Numerous issues arise in many fields such as the economy, politics, and industry. Analyzing these issues has long been regarded as a core task for society.
- If issues are not handled properly, they can grow into national challenges and cause large economic losses. The paper gives climate change, food crises, and cybersecurity as examples.
- Issues change: they may appear briefly and disappear, or keep getting worse over time. Society as a whole therefore needs to be monitored continuously over time.
What Existing Research Has Not Addressed
Existing studies that detect events and issues from web data fall into three types. Feature-based approaches use features extracted from documents, such as keywords, hashtags, time, and location. Topic-model-based approaches use probabilistic models such as Latent Dirichlet Allocation (LDA) and treat each topic distribution as one event. Both approaches require parameters such as the number of issues and the monitoring period to be set in advance, and because issues keep changing, it is hard to set these values beforehand.
Incremental approaches addressed this problem by detecting existing issues and newly emerging issues together. However, most studies using this approach defined a cluster of similar documents as an issue or event. The paper points out three gaps.
- It is hard to state the time of issue emergence quantitatively. If a cluster of similar documents is defined as an issue, enough documents must accumulate before it is recognized as an important issue.
- When a few documents are grouped into a new issue, it is unclear whether the group came from a concentration of people’s attention or from a trivial event.
- It is hard to rank issues by urgency or importance. Existing studies stopped at identifying issues or presented frequently occurring events as important issues. With this approach, many issues must be reviewed continuously, so urgent issues are hard to find in time.
The paper views an issue as “a topic on which people’s attention suddenly converged,” and seeks to obtain the timing and the ranking together in line with that definition.
Method: Finding the Issue Period and Ranking from Structural Changes in the Keyword Network
The paper describes the method in four steps.
- Collect and preprocess web data to determine valid keywords.
- Build a keyword co-occurrence network for each period.
- Monitor structural changes with network structure entropy to find the period when an issue emerged.
- Define issue candidates in that period and prioritize them to detect the most urgent issue.
This post also explains the method in the same four steps.
Case and Data
The paper used the diesel exhaust fluid (DEF) problem in Korea as its case of a social issue. Korea experienced a DEF shortage in late 2021. DEF is a liquid essential for reducing harmful nitrogen oxide emissions from modern diesel vehicles. According to Ministry of Environment data, the vehicles that need DEF included passenger cars, freight trucks, and public transit buses. Without DEF these vehicles could not operate, which disrupted freight transport and commuting.
The data were collected from BIGKinds, a news big data analysis service run by the Korea Press Foundation.
- Collection scope: Articles from 11 national daily newspapers were collected, excluding regional dailies and specialized papers.
- Number collected: The publication date, outlet, and title of 129,747 articles were collected. The paper states the collection period as “September 1, 2021 to November 31, 2021.” Because November has only 30 days, this appears to be a date error.
- Body text: BIGKinds provides only part of each article’s body text, so the researchers collected the full text by web scraping.
- Analysis target: The number of articles mentioning “DEF” (요소수) surged from early November 2021. The analysis therefore used 23,925 articles published from October 25 to November 14.
Step 1: Web Data Collection and Preprocessing
- Input: Text of web data created by members of society
- Processing: Keyword extraction, stopword removal
- Output: List of valid keywords
- Collect the text of web data. Web data are generated in real time, so they are updated faster than other public data. Most web platforms provide public APIs, making collection easy, and even without an API the data can be collected by web crawling.
- Extract keywords from the text.
- Filter out keywords that do not fit the analysis or are too general. Remove stopwords, such as emoticons and onomatopoeia, that do not suit social monitoring.
In the case study, the keyword data that BIGKinds provides for each article were used. These data are lists of noun keywords extracted from the article title and body using a morphological analyzer and a Korean dictionary. The researchers removed keywords meeting the following four conditions as stopwords.
- Keywords of length 1: “A”, “끝”
- Keywords starting with a number: “1181조”, “282.8%”
- Foreign-language keywords that are neither Korean nor English: “ざるうどん”, “水火”
- Keywords that are meaningless or too broad: “B씨”, “내년”, “지난달”
As a result, 155,609 valid keywords remained out of 234,975 unique keywords.
Step 2: Building Keyword Co-occurrence Networks by Period
- Input: Valid keywords, publication dates of documents
- Processing: Setting periods, extracting co-occurrence pairs, filtering pairs, building networks
- Output: One keyword co-occurrence network for each period
- Set multiple periods to monitor.
- Find keyword co-occurrence pairs for each period. Rather than building one network from the entire data set, a separate network is built for each period. Because issues change over time, comparison across periods is needed to see when attention converged on what.
- Set the co-occurrence unit according to document length. For short documents such as microblog posts, all keyword combinations within a document are taken as pairs. For long documents such as news articles, only combinations within a paragraph or sentence are taken as pairs. This is because using all combinations in a long document produces too many pairs, making interpretation difficult.
- Build the network with keywords as nodes, co-occurrence as links, and the number of co-occurrences as link weights.
In the case study, the periods were set following the weekday pattern of article counts. Because weekend articles are far fewer than weekday articles, one period was set to 7 days so that every period includes a Saturday and a Sunday. From October 25 to November 14, this produced 8 overlapping periods. The first period is October 25–31, the fifth period is November 2–8, and the last period is November 8–14.
The BIGKinds keyword data provide an average of 99.2 keywords per article. Taking all keywords within an article as pairs would produce too much data. The researchers therefore defined co-occurrence pairs based on the n-gram concept.
- Split the article body into sentences, and split each sentence into space-delimited words.
- Match valid keywords to the split words and assign a position number (index) within the sentence.
- Take two keywords as a pair if the difference in their position numbers is at most N. The researchers set N to 5 through repeated experiments.
For example, if in one sentence “DEF” is at position 2 and “freight truck” (화물차) is at position 6, the difference is 4, so the two keywords form a pair. A keyword at position 9 differs from “DEF” by 7, so they do not form a pair.
The researchers also applied the Pareto principle (the 80-20 rule). Pairs that co-occur too rarely are unlikely to be meaningful. For each period, they therefore kept only the pairs in the top 20% by co-occurrence count, that is, pairs that appeared together 3 or more times. The remaining pairs averaged 26,418.5 per period. There were 27,387 in the first period and 25,564 in the third period.
Step 3: Identifying the Issue Period
- Input: Keyword co-occurrence networks by period
- Processing: Computing centrality of nodes and links, computing structure entropy, monitoring changes between periods
- Output: The period when an issue emerged
- For each network, compute the degree centrality and betweenness centrality of nodes, and the weight and betweenness centrality of links.
- Using these values, compute node structure entropy and link structure entropy, and add the two to obtain network structure entropy.
- Lay out the values by period and find the period when entropy dropped rapidly. That period is regarded as the period when an issue emerged.
This metric comes from the concept of information entropy. When entropy increases, information becomes more diverse and the system becomes disordered; when entropy decreases, information diversity falls and the system takes on order. A rapid drop in the structure entropy of a keyword network in some period is interpreted as follows. Earlier, several topics were mentioned to similar degrees, but in that period particular keyword combinations were mentioned heavily. When links suddenly appear or weights surge, the network structure changes quickly.
The formula divides into a node part and a link part.
- Node part: The share p_i of node i is its degree centrality divided by the sum of all degree centralities. The correction value q_i is 1−(maximum betweenness centrality−betweenness centrality of node i). These two values are computed in the form (p_i^q_i − p_i)/(1 − q_i + ε), summed over all nodes, and the sign is reversed. ε is a small value greater than 0.
- Link part: The share p_j of link j is its weight divided by the sum of all weights. The correction value q_j is obtained in the same way from the link’s betweenness centrality.
- Whole network: Add the node structure entropy and the link structure entropy.
In everyday terms: the more connections concentrate on a few keywords and links, the more the value moves downward, and the more the connections spread across many places, the more the value moves upward. Computing for the fourth period in the case (October 31–November 6), the node structure entropy of −9.0344 plus the link structure entropy of −11.9087 gives −9.0344+(−11.9087)=−20.9431, which is the network structure entropy.
In the fifth period (November 2–8), the node structure entropy was −9.3711 and the link structure entropy was −11.9329, so the network structure entropy was −21.3039. Network structure entropy thus fell from −20.9431 in the fourth period to −21.3039 in the fifth period. The values for all eight periods are in Table 6 of the paper (Lee et al., Applied Sciences 2023, CC BY 4.0).
The researchers visualized these values in time order and judged that entropy dropped rapidly in the fifth period (November 2–8).
Step 4: Defining Issue Candidates and Computing Priorities
- Input: The network of the issue period and the network of the period immediately before it
- Processing: Community detection, computing the impact score of each community
- Output: List of issue candidates with priorities, and the most urgent issue
- Apply a community detection algorithm to the network of the issue period. In a keyword network, a community is a group of keywords that frequently appear together and thus represents a single topic. The extracted communities are therefore defined as issue candidates.
- Compute an impact score for each community. Because degree centrality enters the entropy formula, if the degree centrality of a previously low-centrality keyword suddenly rises, entropy also changes quickly. The score is therefore computed from the change in degree centrality of the keywords that make up the community.
- Detect the candidate with the highest impact score as the most urgent issue. A high-scoring candidate is a topic that members of society suddenly noticed.
In the case study, the Louvain method was used for community detection. This method quickly and accurately finds community structures with high modularity in large networks. Modularity shows how densely links are packed within communities compared with links between communities. The Louvain method repeats two phases.
- Phase 1: Each node initially has its own community. Each node is moved to the neighboring community where modularity increases the most, and this is repeated until there is no further increase.
- Phase 2: Each community is merged into a single node to build a new network. The weight of a new link is the sum of the weights of the original links between communities.
- The two phases are repeated until modularity no longer increases.
The impact score means “the average change in degree centrality of keywords present in both periods.”
- Numerator: Among the keywords in the community, sum the change in degree centrality (current value − previous value) of those present in both the current and the previous period.
- Denominator: The number of keywords present in both periods.
- New keywords that were absent in the previous period are excluded from both the numerator and the denominator.
Let us look at the calculation using keywords of the top-ranked community. The degree centrality of “DEF” rose from 5116 in the fourth period to 11,039 in the fifth period, a change of 5923. “Injection” (주입) rose from 128 to 167, a change of 39. Assuming, only to show the structure of the formula, that the community consists of these two keywords alone, the impact score is (5923+39)÷2=2981. This value is a simple illustration based on the assumption and is not the score of an actual community.
| Keyword | 4th period | 5th period | Change |
|---|---|---|---|
| 요소수 (DEF) | 5116 | 11,039 | 5923 |
| 주입 (injection) | 128 | 167 | 39 |
| 완주 (Wanju) | 42 | 66 | 24 |
| 아톤산업 (Aton Industry) | 3 | 20 | 17 |
| 익산 (Iksan) | 24 | 34 | 10 |
| 화물차 (freight truck) | 11 | 15 | 4 |
| 집중단속 (intensive crackdown) | 30 | 7 | −23 |
Source: Excerpted and reconstructed from part of Table 9 of the paper (Lee et al., Applied Sciences 2023, CC BY 4.0). Values are degree centrality.
Keywords whose degree centrality decreased, such as “intensive crackdown” (집중단속), lower the score. The score of an actual community comes from averaging the changes of far more keywords, so its scale differs from the hypothetical example above.
Results: The DEF Issue Ranked First in the Period When Entropy Dropped Sharply
Comparing Structure Entropy by Period
The researchers compared the network structure entropy of the 8 periods and defined the fifth period (November 2–8) as the period when an issue emerged. The analysis window itself was chosen by looking at when articles mentioning “DEF” surged. In the author’s view, the fifth period, in which entropy fell, overlaps with the early-November surge. However, this overlap is the author’s interpretation, and the paper does not state that it separately confirmed this agreement. The paper also states that it did not validate the results with quantitative metrics.
Community Detection Results
The fifth-period network had 19,152 keywords. The Louvain method found 342 communities in this network. The number of keywords in a community was as follows.
- Mean: 56
- Median: 22
- Maximum: 3341
- Minimum: 10
The researchers defined all 342 of these communities as issue candidates for that period.
Impact Score Ranking of Issue Candidates
A higher impact score means the community contributed more to the sharp change in structure entropy. The top three candidates are as follows.
| Community | Number of keywords | Impact score | Main keywords |
|---|---|---|---|
| #86 | 46 | 194.1 | DEF, express bus, diesel vehicle, freight truck |
| #42 | 13 | 81.1 | resistant bacteria, bacteria, carbapenem, antibiotics |
| #276 | 84 | 22.0 | medical certificate, painkiller, prescription, fentanyl |
Source: Table 8 of the paper
The score of community #86, 194.1, is 194.1−81.1=113.0 higher than that of the runner-up, 81.1. The researchers detected #86 as the most urgent social issue. The community numbers in Table 8 of the paper and in its body commentary do not match, so this post follows the notation in Table 8. Based on Table 8, #42 is the antibiotic resistance candidate and #276 is the fentanyl candidate.
The content of the second- and third-ranked candidates is as follows.
- #42 antibiotic resistance candidate (impact score 81.1): An issue related to the “2nd National Action Plan on Antimicrobial Resistance” that the Ministry of Health and Welfare prepared in that period. The paper states that Korea’s human antibiotic use is the third highest among OECD countries.
- #276 fentanyl candidate (impact score 22.0): The main keywords are medical certificate, painkiller, prescription, therapeutic purpose, and fentanyl. The paper describes how people involved were arrested over excessive prescription and use of fentanyl, an opioid painkiller, which became a topic of conversation at the time.
Interpreting the Keywords of the Top Issue
Community #86 contains “DEF” along with “express bus,” “diesel vehicle,” “freight truck,” “skyrocketing” (천정부지), and “market disruption” (시장교란). The biggest problem of this issue was that vehicles such as express buses and freight trucks stopped running because of the DEF shortage and price surge.
Keywords whose degree centrality rose include “injection,” “Wanju,” “Aton Industry,” and “Iksan.” The paper describes the background of these keywords as follows. A DEF manufacturer in Iksan signed agreements with local governments such as Wanju County and supplied DEF to those regions first. The paper states that this manufacturer limited its daily volume and sold DEF to residents at a reasonable price, and was praised for its efforts to stabilize local industry. The paper states that, as a result, this manufacturer, Iksan City, and Wanju County drew attention in that period. This account is the paper’s description, and the author did not separately verify the facts.
The difference from a frequency-based criterion also shows. The most frequently mentioned keywords in this data were those related to the COVID-19 pandemic, such as COVID-19, vaccine, and vaccination. The paper’s authors judge that the impact score appropriately singled out a different topic, namely one on which attention suddenly converged in that period. This judgment is the authors’ qualitative one, and there is no quantitative validation.
The Unfolding of the Issue Seen Through New Keywords
New keywords, which are excluded from the impact score calculation, can be used to anticipate the near future of the issue. The new keywords in the DEF community appeared in two groups.
- Fire-service keywords: “Ulju Fire Station” and “Gwangyang Fire Station” newly appeared. In fact, as the DEF shortage dragged on, the operation of fire engines and ambulances was restricted. Help also appeared nationwide, such as anonymous donors who secretly left DEF in front of fire stations.
- Freight transport keywords: “Cargo Truckers Solidarity” (화물연대), “freight truck owner,” and “freight transport” newly appeared. The paper states that the Cargo Truckers Solidarity announced a strike, saying the damage from the service halt was being passed on to workers.
The paper states that two strikes by the Cargo Truckers Solidarity in one year caused economic losses of 10.4 trillion won. The researchers judged that monitoring new keywords makes it possible to anticipate that the DEF issue could spread into events related to the Cargo Truckers Solidarity and cause additional losses.
Significance and Limitations
What Changes
- Before looking at the content of issues, it quantitatively finds the period when an issue emerged from changes in network structure.
- It defines an issue as “a topic on which attention suddenly converged” and finds the timing in line with that definition. Its starting point differs from existing studies that rely mainly on the frequency of particular topics.
- It ranks issue candidates by their impact score on structural change, not by frequency.
- It presents the issue’s emergence period and its priority together.
Usefulness for Practitioners
Governments and companies must handle numerous issues in culture, the economy, industry, and technology with limited resources. By presenting a ranking of candidates for the period when an issue emerged, this method helps them respond to urgent issues first, in order or selectively. The researchers expected that, because the method is not biased toward a particular field, it could be used for decision-making in many areas of society.
Monitoring new keywords can be used to speed up responses. The researchers wrote that the method helps anticipate the direction in which an issue will develop, which helps in preparing responses for the future.
Usefulness for Researchers
This is a case in which the detection procedure was designed starting from the definition and characteristics of an issue. Because structure entropy and the impact score are computed from centrality metrics, the same framework can be transferred to other keyword network studies. The paper presents quantitative validation, case studies from other countries, and monitoring of the entire life cycle of issues as future tasks.
Limitations Stated by the Paper
The paper itself states three limitations.
- It did not present a method to validate the results with quantitative metrics. The judgment that the DEF issue and others were detected appropriately is the authors’ qualitative judgment. Quantitative validation is needed to compare detection performance with other studies.
- It dealt only with Korean web news data. Case studies from other countries are needed, and using data from multiple countries together would also make it possible to find common or global issues.
- It did not capture the entire life cycle of an issue. Not extracting candidates in every period is advantageous in terms of cost, but it cannot see an issue from birth to disappearance.
Conditions for Application, According to the Author
The following is not from the paper; it consists of conditions the author adds with practical application in mind.
- The period length and overlap interval should be set by looking at the publication cycle of your own data. If the number of posts fluctuates greatly by day of the week or at month-end, the network size differs by period and the entropy comparison becomes unstable.
- It is better to decide in advance within the organization the criterion for judging that entropy “dropped rapidly.” Deciding how large a change from the previous period should trigger an alert keeps the judgment the same even when the person in charge changes.
- The quality of the results depends heavily on keyword extraction and the stopword list. For data without BIGKinds keyword data, such as internal documents or customer posts, morphological analysis and the stopword list should be checked first.
- Rather than looking only at the top-ranked candidate, it is safer to review the top several candidates together. Because community size ranges from 10 to 3341 keywords, a person should read the keyword composition along with the score.
- New keywords do not enter the score, so they should be managed as a separate list. The direction in which an issue will spread shows up first in this list.
What to Try Right Away
- Collect about 3 weeks of news or posts in your field and divide them into overlapping 7-day periods. Count keyword pairs whose position difference within a sentence is 5 or less, and build a network for each period from the pairs that appeared 3 or more times.
- Compute the centrality of nodes and links in each period’s network and calculate structure entropy. Tabulate the change from the previous period and find the period in which the value fell the most.
- Apply the Louvain method to the network of that period to divide it into communities. For each community, compute the average change in degree centrality of keywords present in both periods and rank the communities, and write down the new keywords absent in the previous period as a separate list.
The Business Intelligence Lab studies methods for reading changes in technology and society using web data and patent data. This blog will continue to introduce other paper reviews.