Product managers read customer posts that pile up on review boards and in communities every day. Once there are tens of thousands of posts, you can count “which features are mentioned often.” But it is hard to pull out a single number for “which features are both mentioned often and a source of strong dissatisfaction.” Looking at the ranking by mention volume and the average sentiment score separately makes it hard to decide which feature to fix first.
Can a single score identify the product features that customers consider important but are not satisfied with, using social media posts?
Let us read a paper that answers this question.
- Title: Identification of time-evolving product opportunities via social media mining
- Journal: Technological Forecasting and Social Change
- Year: 2020
- DOI: 10.1016/j.techfore.2020.120045
The paper proposes a method that combines two components. The first is an Event Detection and Tracking (EDT) algorithm based on aging theory (a theory that views an event as something that is born, grows, withers, and disappears). The second is sentiment analysis together with an opportunity algorithm. In the previous post, we used this paper to look at how opportunities change over time. This post focuses on the score calculation logic and on points to watch in practice.
The terms used throughout this post are as follows.
- Event: A cluster of posts with similar keywords, grouped in chronological order. When a new post arrives, the event’s keywords and weights change.
- Energy: The accumulated support an event has received from similar posts. It decreases over time.
- Importance: Defined from the energy value at the same time point, with a natural logarithm (ln) in the formula. The formula in the original paper is garbled, so its exact form cannot be confirmed.
- Satisfaction: The weighted average, across time points, of the sentiment scores of the posts belonging to an event. It lies between 0 and 1.
- Opportunity score: A value that combines importance and satisfaction into one.
- Opportunity landscape map: A map with importance on the horizontal axis and satisfaction on the vertical axis, on which events are plotted as points.
Motivation: Customer Voices Are Abundant, but There Is No Number to Set Priorities
Why Social Media Customer Voices Matter
The paper holds that social media has become a main channel for consumer opinions, and offers the following evidence.
- Reports say that 86% of adults in the United States and 79% of adults in Europe use social media services.
- 77% of customers read other people’s product reviews on social media before buying.
- 75% of customers trust reviews more than direct personal recommendations.
- Companies open social channels such as “My Starbucks Idea” and “Samsung Global” to hear the uncensored Voice of Customer (VOC).
- Companies that identify customer needs faster than competitors build customer relationships that competitors cannot easily imitate.
What Existing Research Has Not Covered
Existing product planning research collects large volumes of commercial reviews, microblogs, and online discussion posts. Researchers apply topic analysis and sentiment analysis to them. Some studies also combine these with theoretical models such as chance discovery, network analysis, morphological analysis, and the Kano model. In these ways, studies have found product features and issues, identified influential users, evaluated feature satisfaction, and created new product concepts.
The paper points out three gaps.
- A life cycle perspective is missing. Social media customer voices arise and change in real time, so their creation, growth, and extinction must be tracked.
- The opportunity level of individual customer voices has not been analyzed. Most studies expressed importance and satisfaction separately, using the number of documents per topic or the average sentiment score.
- No concrete direction has been given for product planning. Satisfaction was quantified, but the results did not lead to a direction for product development and improvement.
The paper also notes that existing research remains static. Finding new opportunities required rerunning complex experiments such as topic modeling every time.
Method: Group Similar Posts into Events, and Compute Importance, Satisfaction, and Opportunity Score for Each Event at Each Time Point
The paper describes its method as collection and preprocessing plus two stages. The two stages are 1) detecting and tracking events that customers talked about and 2) identifying and evaluating product opportunities. In this post, collection and preprocessing is set apart as Step 1. The paper’s second stage is split into two parts: Step 3 covers the satisfaction and importance calculation, and Step 4 covers the opportunity score and map. Thus Step 2 of this post is the paper’s first stage (Section 3.1), and Steps 3 and 4 of this post are the paper’s second stage (Sections 3.2.1 and 3.2.2).
Case and Data
The paper chose smart speakers as its case. There are three reasons.
- The market grew quickly after the Amazon Echo launched in 2014. The paper cites estimates from two reports showing large numbers of users and devices in the United States.
- It is a technology-intensive product connected to smartphones, tablets, music streaming, and smart homes.
- Because the technology is mature, an analysis from the customer side, from a market pull perspective, is needed.
The data were collected from Reddit. Reddit is organized into boards (subreddits) by brand and product name, and its posts are longer than those on Twitter or Instagram. The analysis period is two months, from April 1 to May 31, 2018. A total of 27,013 posts were collected: 9,437 from the Google Home board and 17,576 from the Amazon Echo board. Comments and posts were treated as documents at the same level.
Step 1: Collection and Preprocessing
- Input: Social media posts and their timestamps
- Processing: Filter documents by length and keywords, and convert them into word bundles
- Output: Bags of words (bag-of-words) sorted by timestamp
- Remove short posts under 200 characters. Posts such as “[deleted]”, “Wow”, and “Yup” cannot form an interpretable event.
- Remove all links to other web addresses, such as Amazon, iTunes, and YouTube.
- Extract keywords from each document with the RAKE (Rapid Automatic Keyword Extraction) algorithm.
- Remove meaningless keywords using the NLTK stop word list. However, “don”, “no”, “not”, and “nor” are kept. The paper does not state why it kept these four tokens.
- Remove documents with fewer than 4 valid keywords.
- Represent the remaining documents as bundles of keywords and weights.
The paper gives an example in which a short complaint about the inconvenience of voice commands on a mobile phone is converted into a bundle in which every keyword has a weight of 1.0. This preprocessing happens in real time whenever a new post arrives.
Step 2: Event Detection and Tracking
- Input: Word bundles arriving in chronological order
- Processing: Apply three policies (energy decay, event creation and support distribution, event update)
- Output: A list of living events, with energy and keywords at each time point
- Compute the similarity between the new document and every living event.
- If the highest similarity is below the threshold, the new document becomes a new event (“birth”).
- If it is at or above the threshold, the document joins the single most similar event and gives that event support (“growth”). The size of the support is the similarity value.
- The energy of every event decreases at each time unit (“decline”). An event whose energy falls below a threshold disappears (“extinction”).
A new document is placed in only the single most similar event so that the content of the event stays clear.
Energy is computed by multiplying the previous time point’s energy by the aging coefficient α (between 0 and 1) and adding the support received at that time point. The original aging theory converts accumulated support into a value between 0 and 1 using a sigmoid function, but the paper does not use that function and uses the accumulated support directly as energy. The smaller α is, the faster energy decreases and the more often events disappear.
Document similarity is computed with pretrained word vectors. In the case, fastText was used. The calculation proceeds as follows.
- The similarity between a word w and a document D is the largest of the similarities between w and all keywords in D.
- Only the keywords of document D1 whose similarity to D2 exceeds the threshold are put into the Relevant Word Set (RWS).
- The similarities of the keywords in the RWS are averaged, weighted by frequency, to get “the similarity of D2 as seen from D1.”
- The value in the opposite direction is also computed, and the two values are averaged.
When a new document joins an event, the event’s keyword weights change by the following rules. The explanation below summarizes only the rules, with reference to the example in Figure 2 of the paper.
- Keywords of the new document whose similarity to the event falls short of the RWS threshold are not reflected.
- Existing keywords similar to the new document receive a reward: their weight is multiplied by 1.2.
- Existing keywords far from the new document receive a penalty: their weight is multiplied by 0.8.
- Keywords present on both sides are multiplied by 1.2, and then the frequency in the new document is added.
- Keywords new to the event are added, receiving their frequency in the new document as their weight.
- Keywords whose weights drop through these updates are removed from the event’s keyword set.
In the case, α was set to 0.9 and the document similarity threshold to 0.7.
Step 3: Computing Satisfaction and Importance
- Input: Energy by event and time point, and the documents belonging to each event
- Processing: Compute a sentiment score for each document and take a weighted average, and compute importance from energy
- Output: Satisfaction and importance by event and time point
- When a new document joins an event or becomes a new event, measure the sentiment polarity of that document.
- At each time point, compute the average sentiment score of the documents that entered that event.
- Take a weighted average of the per-time-point average sentiments to get satisfaction. The older the time point, the smaller the weight.
- Based on the energy value at the same time point, compute importance with a formula that includes the natural logarithm (ln).
The paper notes that Python libraries such as NLTK, TextBlob, gensim, and VADER can be used as sentiment analysis tools.
The reason for reducing the satisfaction weights over time points is to put more weight on recent customer reactions. The weight is set by a real number α between 0 and 1. The paper only states that the weight decreases over time. It does not state whether this α is the same value as the α in the energy calculation. If satisfaction is close to 1, customer reaction to the event is positive, and if it is close to 0, it is negative.
Importance is larger for events that many customers mention often, because an event that receives many similar documents has high energy. In the case’s April 1 importance list, the two lowest events, “Feature comparison” and “Google Assistant,” are at 0.693. Because the importance formula in the original paper is garbled, this value cannot be converted back exactly to energy.
In the case, the May 31 satisfaction values were 0.152 for “Malfunction,” 0.136 for “Reminder,” and 0.129 for “Routines.”
Step 4: Opportunity Score and Opportunity Landscape Map
- Input: Importance and satisfaction by event and time point
- Processing: Compute scores with the opportunity algorithm and place events in the three regions of the map
- Output: Opportunity score ranking by time point, and lists of improvement and stabilization opportunities
- The opportunity score is computed as “importance + max(importance − satisfaction, 0).”
- When drawing the map, importance is min-max normalized (a transformation that sets the smallest value to 0 and the largest to 1) at each time point. This is because the maximum importance grows over time.
- The map is divided into three regions: appropriately served, over-served, and under-served.
- Events with high importance and low satisfaction are seen as “Improvement opportunities.” Events with low importance and high satisfaction are seen as “Stabilization opportunities.”
This formula adapts to events the opportunity algorithm that Ulwick (2009) proposed in outcome-driven innovation theory. The meaning of the formula is simple. If importance is greater than satisfaction, the score gains that difference on top. If satisfaction is greater than importance, the second term becomes 0 and the score equals importance.
Take the May 31 “Malfunction” event as a worked example. Importance is 2.321 from Table 4, and satisfaction is 0.152. The opportunity score is 2.321+(2.321−0.152)=2.321+2.169=4.490. The value the paper reports is 4.491, which differs slightly in the third decimal place.
Using the same formula in reverse, we can check the calculation for an event for which the paper gives importance, satisfaction, and opportunity score. “Reminder” has an importance of 1.998 and an opportunity score of 3.859. Recovering satisfaction from the formula gives 2×1.998−3.859=0.137, which is almost the same as the paper’s 0.136. For the two events (“Malfunction” and “Reminder”), the paper’s opportunity scores are reproduced from the importance values in the table and the formula. Whether this importance is the value before or after normalization is not stated in the paper, so this calculation alone cannot confirm it.
The paper regards improvement opportunities as important features that attract customers in a competitive market and as the ones to address first. However, as the results below show, the score ranking and the map region can differ, so it is safer to check both before setting priorities. The paper sees stabilization opportunities as usable for keeping customers with high satisfaction and loyalty at low investment. It also suggests that stabilization opportunities may point to areas of disruptive innovation.
What Changes with the Settings
The results depend on several values set by the analyst.
- Aging coefficient α: Lowering the value makes events disappear faster. The paper says to set it according to the occurrence frequency of the input data and the purpose of the analysis.
- Similarity threshold: It decides whether to create a new event or add the document to an existing one. Set it according to the content characteristics of the data.
- Time unit: This is the unit at which energy decreases. To observe the life stages of events, an appropriate unit must be chosen.
- Normalization: The position on the map is determined by the per-time-point normalization result.
Results: As of May 31, the “Malfunction” Event Was the First to Fix, with an Opportunity Score of 4.491
Event Detection Results: Only 355 of 5,327 Events Lasted Beyond One Day
Over the whole analysis period, 5,327 events were detected. Of these, 93.3% disappeared within a day. This resulted from similar documents not following, and was interpreted as small issues or unpopular opinions. For example, a YouTube media access event created from two documents on April 25 disappeared the next day.
The paper analyzed the 355 events that lasted more than a day. These events occurred about 7.11 times per day on average, and their average lifespan was 5.42 days. Events that stayed alive through all 60 days of the analysis period include “Bluetooth connection” (event 4) and “Music play” (event 9). “Spotify” (event 931) lasted 50 days. Event names are assigned by people who look at the main keywords and representative documents. The researchers carried out this interpretation process with 9 data analysis experts (3 professors and 6 graduate students).
Top Events by Importance at Each Time Point
On April 1, 62 events were created, and 6 of them lasted beyond one day. Ranked first in importance was “Bluetooth connection” (1.594), and second was “Music play” (1.211). By May 1, 186 events that lasted beyond one day had been created, and 151 of them had disappeared. That day, first place was “Music play” (2.517), second was “Spotify” (2.244), and third was “Bluetooth connection” (2.228).
The top events on May 31 are as follows.
- “Malfunction”: 2.321
- “Bluetooth connection”: 2.090
- “Reminder”: 1.998
- “Spotify”: 1.899
The paper read the appearance of “Spotify,” “Bedroom light,” and “Smart plug” on May 1 as a signal that the scope of use had widened to music streaming and lighting control.
The May 31 Opportunity Landscape Map
The under-served region contained 6 events with high opportunity scores.
- Malfunction
- Bluetooth connection
- Reminder
- Feedback submission
- Music play
- Routines
“Malfunction” had the highest importance but a very low satisfaction of 0.152, so it had the highest opportunity score at 4.491. Customers left complaints about the sound quality of voice clips, language support, message sending, and command recognition. “Routines” is a new feature that runs several tasks with a single command. This event’s opportunity score was 3.020. Customers tried it with high expectations, but satisfaction was only 0.129. One customer complained that basic routines work but custom routines do not.
The appropriately served region had 27 events, and the over-served region had 18 events. The over-served region included “YouTube Music” (2.533), “Bart” (1.890), and “Google Play Music” (1.649). Google Play Music and YouTube Music users were fewer than Spotify users, but their average satisfaction was higher. The appropriately served region included “Spotify” (3.299), “Bedroom light 2” (3.200), and “Harmony” (2.702).
The opportunity score ranking and the map region do not always agree. “Spotify” (3.299) in the appropriately served region has a higher score than “Routines” (3.020) in the under-served region. The paper does not explain why such discrepancies arise. The author guesses that it may be because the map is drawn with normalized importance, but this is not something confirmed as grounds. Because the ranking and the region can differ, check both in practice.
Change in a Single Event’s Opportunity Score
The paper tracked changes in the scores and positions of the “Music play” (event 9) and “Audio device connection” events.
“Music play” was in the over-served region at its birth on April 1, and its opportunity score of 1.600 was higher than the event average of 1.191 at that time. Afterward, interest grew and satisfaction fell. On May 1, the score rose to 4.447 and the event moved to the under-served region. On May 31, satisfaction rose and importance fell slightly, so the score dropped to 3.305.
“Audio device connection” is an event about connecting a smart speaker to soundbars, Bluetooth speakers, and TVs. At its birth on May 9, it was in the appropriately served region. Its score of 1.751 at that time was lower than the average of 1.983. On May 20 it moved to the under-served region, and its score rose to 4.517. After that, importance fell from 0.994 to 0.266 and satisfaction rose from 0.483 to 0.590, so it moved to the over-served region.
Significance and Limitations
What Changes
- Customer voices are tracked by life stage (birth, growth, decline, extinction). Changes that static analysis could not show come to light.
- Importance and satisfaction are combined into one opportunity score per event to set priorities.
- How an event’s position on the map changes at each time point is expressed numerically.
- You can see how an event’s keywords change over time. For example, the top keyword of “Music play” changed from “country music” (11.246) on its birth day to “play music” (1199) on May 31.
Value for Practitioners
Events in the under-served region directly point to places to fix, such as unsatisfactory features or designs. Events in the over-served region show features that can keep loyal customers at low investment. The opportunities derived cover not only the company’s own products but also competing products, related services, software, and the purchase experience. The paper sees this method as something that can be developed into real-time customer monitoring software.
Value for Researchers
The method is not tied to a specific field and can be reproduced. It can also be applied to text on other subjects, such as news. The word vectors, sentiment analysis tool, and normalization method can each be swapped. This makes follow-up research possible that compares how much each component changes the score.
Limitations Stated in the Paper
The paper lists the following as main limitations.
- The document similarity calculation can be improved. Newer models such as BERT or ELMo could be used instead of fastText.
- The phenomenon of an event similar to a previously extinct one reappearing later was not handled. The approach of putting a new document into only the single most similar event could also be changed to assign it to all similar events.
- Only past changes in opportunity scores and map positions were analyzed. Future positions and scores cannot be predicted.
The paper lists applying the method to patents, news, scientific papers, and companies’ own VOC data as future work.
Conditions for Application, in the Author’s View
The following is not content from the paper but conditions the author adds with practical application in mind.
- Check the ranges of the two indicators. Satisfaction lies between 0 and 1, but importance is a log value derived from energy, and in Table 4 its smallest value is 0.693 and its largest is 2.517. If you add and subtract them as they are, the opportunity score is affected more by importance.
- Look at the score and the map region together. In the case, some events had different ranks and regions, and the paper does not explain why.
- Importance is mention volume. In the case, the Amazon Echo board had more posts (17,576) than the Google Home board (9,437). If you mix data from communities of different sizes, events from the larger community tend to come out as important.
- Once you choose a sentiment analysis tool, keep it. If the tool changes, satisfaction values change and comparisons across time points break.
- Re-set the criterion for discarding events that disappear within a day according to data size. In the case, 93.3% of detected events were dropped by this criterion.
What to Try Right Away
- Collect posts from your own product’s community or reviews along with their timestamps. Remove posts under 200 characters and posts with fewer than 4 valid keywords, and convert the rest into keyword bundles. Start by checking how many posts remain.
- If you already have data grouped by topic, run a simplified trial procedure of the paper’s method. The procedure below is not the paper’s method as is.
- Energy: Treat the number of mentions per day in each bundle as support, multiply the previous energy by α, and add. Start with α at 0.9 from the case. The paper accumulates the similarity between a document and an event as support, not the mention count, so this value differs from the energy of the original method.
- Satisfaction: For each post in a bundle, compute a sentiment score between 0 and 1 and average by date. Take a weighted average with smaller weights for older dates.
- Opportunity score: Compute “importance + max(importance − satisfaction, 0)” and pick the top five. Compute the result with importance as is and the result with importance normalized to 0–1, and compare how much the ranking differs.
- Read the representative posts of the selected top bundles yourself and name them. In the paper too, event names were assigned by people looking at keywords and representative documents. At this first step, check whether the score ranking matches the actual content of the complaints.
Business Intelligence Lab continues to introduce methods for analyzing customer data and technology data together.