In my view, product planners often use surveys to gather customer opinions. Survey questions are built from attributes the planner has chosen in advance. So if an attribute that customers actually care about is not in the questions, that opinion does not show up in the results. The survey discussion in this paragraph is my own general observation, not part of the paper discussed below. Meanwhile, customers leave comments about products on social media such as blogs, Twitter, Facebook, and Reddit, and these posts pile up quickly.
So can we extract the product attributes customers actually mention from social media posts alone, and even decide which of those attributes to improve first?
Let us read a paper that answers this question.
- Title: Social media mining for product planning: A product opportunity mining approach based on topic modeling and sentiment analysis
- Journal: International Journal of Information Management
- Year: 2019
- DOI: 10.1016/j.ijinfomgt.2017.09.009
The paper calls this method an opportunity mining approach. The method finds the product topics that customers talk about in social media posts. Then, for each topic, it computes two numbers: how often the topic is mentioned (importance) and how favorably customers speak about it (satisfaction). Finally, it combines the two numbers to rank topics, starting with those that offer the greatest opportunity for improvement.
The key terms used in this post are as follows.
- Product topic: A theme of the product that customers discuss on social media. In this post, it means the same thing as “the product attributes customers see.”
- Topic modeling: A technique that infers the themes hidden in a collection of documents. It treats each document as a mixture of several topics and each topic as a mixture of several words.
- Latent Dirichlet Allocation (LDA): The topic modeling algorithm used in this paper.
- Sentiment analysis: The task of determining whether a writer expressed a positive or a negative opinion.
- Opportunity algorithm: A method that computes an opportunity from two values, importance and satisfaction, to prioritize unmet needs.
Motivation: There Were Few Ways to Turn Accumulating Customer Opinions into an Improvement Order
Why Finding Product Opportunities Early Matters
The paper gives the following reasons.
- In customer analysis, identifying new product opportunities early is regarded as the first requirement of product development and improvement.
- Firms that find opportunities early build customer relationships that competitors cannot easily copy, and secure a competitive position in the value chain.
- As product life cycles shorten and the business environment globalizes, customer needs change faster and have become more complex than before.
- Social media is a source that holds large amounts of customer opinion in real time. In the survey the paper cites, 86% of American adults and 79% of European adults use social media.
What Earlier Research Did Not Cover
Sentiment analysis has already been used in product planning. Studies have identified product weaknesses from online reviews and measured customer satisfaction with mobile services. Other studies have assessed service quality from user posts or defined brand image on social media.
The gaps the paper points out are as follows.
- Analysis stayed at the word level, making it hard to find the latent product attributes customers talk about.
- Studies did not analyze how valuable the attributes they found are as improvement opportunities.
- They focused on classifying positive and negative opinions and did not address latent opportunities.
- Most gave no guidance on the direction of product development.
Method: Find Attributes with Topics, Score Satisfaction with Sentiment, and Then Compute Opportunity Scores
The paper describes the method in four steps. In Step 1, social media posts about the target product are collected, and keywords are extracted and cleaned. In Step 2, LDA finds the product topics and the importance of each topic is computed. In Step 3, keyword sentiment scores are used to compute the satisfaction of each topic, and in Step 4, the opportunity algorithm yields opportunity scores that determine the direction of improvement.
Case and Data
The case product is the Samsung Galaxy Note 5 (SGN5). This smartphone has general features such as calling, wireless internet, and photography. It also has features unique to the device, such as the S Pen, wireless charging, and Samsung Pay. The researchers judged that its many features and large user base make it suitable for explaining the analysis process.
The data source is Reddit. Twitter or Facebook data can be collected only with keywords or hashtags chosen by the analyst. In contrast, a subreddit (a discussion board for a particular area of interest) also accumulates indirect opinions that do not contain product keywords.
The researchers collected 23,614 documents from the SGN5 subreddit: 2,255 posts and 21,359 comments. The writing period runs from November 7, 2014 to January 31, 2016. Before the release date (August 21, 2015), an average of 4.41 documents per day were posted. After release, the average was 136.25 per day, 31 times the pre-release rate.
Step 1: Data Collection and Preprocessing
- Input: Customers’ social media posts about the target product
- Process: Keyword extraction, removal of meaningless keywords and documents
- Output: An array for each document listing keywords and their occurrence frequencies
- Collect customer posts about the target product. Use web crawling or the public Application Programming Interface (API) that the social media platform provides.
- Extract keywords or key phrases from each post.
- Remove emoticons, onomatopoeia, and words that are irrelevant to the analysis or too general.
- Organize each post as an array of keywords and their frequencies.
From the data organized this way, the paper builds a document-keyword frequency matrix and uses this matrix as the input to LDA.
In the case study, the researchers used Rapid Automatic Keyword Extraction (RAKE). RAKE needs no training data, can be used regardless of domain and language, and also extracts compound terms. They first obtained 32,637 keywords, but these included web page addresses and colloquialisms such as “WoW.”
The researchers filtered keywords by their occurrence frequency within documents and by Document Frequency (DF), and then manually deleted those among the remaining keywords that were unrelated to SGN5. As a result, 3,539 keywords remained. The 12,491 documents containing none of these keywords were regarded as noise documents, such as advertisements, and excluded. The final analysis set is 11,123 documents (23,614 − 12,491 = 11,123).
Step 2: Finding Product Topics and Importance
- Input: Document-keyword frequency matrix, keyword list, number of topics
- Process: Run LDA, name the topics, sum and normalize contributions by topic
- Output: List of product topics and an importance value between 0 and 10 for each topic
- Build a matrix with documents as rows and keywords as columns, where each element holds the keyword’s occurrence frequency.
- Decide the number of topics. Vary the number of topics and compute the average cosine similarity over all topic pairs. Then choose the point where this value becomes lowest.
- Run LDA. LDA decomposes the input matrix into a document-topic matrix and a topic-keyword matrix.
- Read the keywords and documents with large contributions, and have a person name each topic.
- For each topic, add up its contribution probabilities over all documents to obtain importance, and normalize it to 0–10.
The reason for choosing the number of topics this way is the view that the number at which topics are most distinct from one another is the appropriate one. Cosine similarity is the cosine of the angle between two vectors. It equals 1 when the two vectors are identical. Therefore, the lower the average similarity of topic pairs, the less the topics overlap.
The starting point of importance is the contribution stock. It is the value obtained by summing a topic’s column in the document-topic matrix over all documents, and it indicates how much that topic was mentioned by users. The normalization formula is “10 × (the topic’s contribution stock − minimum) ÷ (maximum − minimum).” The most mentioned topic receives 10 points and the least mentioned topic receives 0 points.
Let us calculate with the case numbers. The maximum contribution stock is 202.2198 for Samsung Pay, and the minimum is 165.9974 for Accessory. The contribution stock of Fast charge is 191.5701. Importance is 10 × (191.5701 − 165.9974) ÷ (202.2198 − 165.9974) = 10 × 25.5727 ÷ 36.2224 = 7.0599, which matches the value in Table 6 of the paper. The body of the paper gives only the maximum and minimum and the contribution stocks of some topics (Samsung Pay, Fast charge, Software update); the full set of normalized importance values is in Table 6.
In the case study, the inflection point of the average cosine similarity appeared at 65 topics. The average similarity at that point was 0.04554. The researchers ran LDA with Netminer. For example, Topic-10 was named “Expandability” because its main keywords were SD (0.130), SD-card (0.094), removable battery (0.049), MicroSD (0.036), and IR blaster (0.016).
Step 3: Computing the Satisfaction of Product Topics
- Input: Sentences in which keywords appear, the topic-keyword matrix from Step 2
- Process: Average keyword sentiment scores, build a sentiment-weighted matrix, sum and normalize by topic
- Output: A satisfaction value between 0 and 10 for each topic, and a sentiment-weighted topic-keyword matrix
- Compute a sentiment score each time a keyword appears, and average by keyword to build a keyword sentiment vector.
- Multiply each row of the topic-keyword matrix by the keyword sentiment vector to build a sentiment-weighted topic-keyword matrix.
- Sum all the values in each row to obtain the topic’s sentiment stock.
- Scale the sentiment stock to 0–10 in the same way as importance.
Sentiment analysis approaches are divided into dictionary-based, text-classification-based, and deep-learning-based. The paper chose the deep-learning-based approach, because it performed well in sentiment analysis of short texts such as Twitter posts. For the actual analysis, it used AlchemyAPI on the IBM Watson platform. The purpose was to reduce the time spent implementing and testing a classifier and to take advantage of the accuracy of a classifier trained on large data.
Even the same keyword has different sentiment in different sentences. “Battery” received −0.9013 in a sentence blaming an app for draining the battery. Conversely, it received 0.8790 in a sentence saying the battery is fine. So the researchers used the per-keyword average as the sentiment weight.
The per-keyword sentiment weights split into positive and negative values. AMOLED, the screen material, received a positive value of 0.5689, and the fingerprint scanner (finger print scanner) received a negative value of −0.5141. Battery and SD card were also negative. Some of the weights for all keywords are listed in Table 5 of the paper.
In the case study, the 3,539 keywords appeared a total of 105,483 times in the 11,123 documents. The topic with the highest sentiment stock is Touch wiz (0.1093), and the lowest is Detect pen1 (−0.2918). The mean was −0.0999 and the standard deviation was 0.0791. Of the 65 topics, 36 were above the mean. After normalization, the satisfaction of Touch wiz is 10.0000 and that of Detect pen1 is 0.0000.
Step 4: Finding Product Opportunities with the Opportunity Algorithm
- Input: Importance and satisfaction by topic, sentiment-weighted topic-keyword matrix
- Process: Compute opportunity scores, draw an opportunity landscape map, interpret sentiment keywords of top topics
- Output: Opportunity score by topic and direction of improvement
- Compute an opportunity score for each topic.
- Draw an opportunity landscape map with importance and satisfaction as the two axes. This map is divided into three regions: served-right, over-served, and under-served.
- For topics with high opportunity scores, find the keywords with large sentiment-weighted values, and set the direction of improvement toward increasing satisfaction or reducing dissatisfaction.
The opportunity algorithm was proposed by Ulwick together with Outcome-Driven Innovation (ODI). ODI holds that an innovation opportunity exists when a need is important but not sufficiently satisfied. The formula is “Opportunity = Importance + Max(Importance − Satisfaction, 0).” If importance is higher than satisfaction, the difference is added once more to importance. If satisfaction is higher than or equal to importance, only importance remains.
Let us calculate three cases with the case values.
- Fast charge: importance 7.0599, satisfaction 5.4720. 7.0599 + (7.0599 − 5.4720) = 7.0599 + 1.5879 = 8.6478.
- Software update: importance 4.6186, satisfaction 4.9924. The difference is negative, so it becomes 0, and the opportunity score equals importance, 4.6186.
- Detect pen1: importance 2.2850, satisfaction 0.0000. 2.2850 + (2.2850 − 0) = 4.5700.
Needs in the under-served region are needs that are less satisfied relative to their importance, and the opportunity algorithm regards these as innovation opportunities.
Results: Samsung Pay, Fast Charge, and the S Pen and Battery Topics Ranked Highest as Improvement Opportunities
Composition of the 65 Topics Customers Raised
Of the 65 topics, 6 deal with battery and charging: Battery life, Battery usage, Battery drain, Fast charge, Charger, and Wireless charge. Features that are new or changed compared with the previous model, the Galaxy Note 4, were also captured in 6 topics: Samsung Pay, Pen out, Expandability, Wireless charge, Screen off memo, and Samsung Pay on ATM.
Software-related topics were the most numerous, at 13. Software update, Music App, Multi-window, Touch wiz, and others belong here. Based on this result, the researchers interpreted software as being as important as hardware in the smartphone market.
Distribution of Importance and Satisfaction
The mean contribution stock was 171.1228 and the standard deviation was 6.58. Only 20 of the 65 topics were above the mean. The researchers singled out Samsung Pay (202.2198), Fast charge (191.5701), and Software update (182.7272) as values far from the mean.
After normalization, the average importance across all topics was 1.4150 and the average satisfaction was 4.7854.
Topics with the Highest Opportunity Scores
On the opportunity landscape map, 7 topics are in the under-served region and have high opportunity scores: Samsung Pay, Fast charge, Battery life, Expandability, Pen out, Detect pen1, and Detect pen2. All of these topics have satisfaction lower than importance. The average opportunity score across all topics is 1.7098. The full values by topic are in Table 6 (importance and satisfaction) and Table 7 (opportunity scores) of the original paper. Here we look at only two representative cases.
Samsung Pay has importance 10.0000, satisfaction 7.8185, and opportunity score 12.1816, ranking first in opportunity score. Its satisfaction is higher than the average of 4.7854 but lower than its importance of 10.0000. Because satisfaction is lower than importance, it is classified in the under-served region. The researchers interpreted this as the feature being new in SGN5, which drew many mentions and expectations.
Detect pen2 has importance of 4.3805, lower than Pen out’s 4.6130. However, its satisfaction of 0.1384 is lower than Pen out’s 0.6877. So the opportunity score of Detect pen2, 8.6225, surpassed Pen out’s 8.5384. The paper explains that the S Pen is a core feature of the Galaxy Note series that other smartphones do not have.
Fast charge and Battery life ultimately relate to usage time. Removable battery, a main keyword of Expandability, is also tied to usage time. Even though solutions such as high-voltage charging, wireless charging, and new materials had appeared, customers’ needs remained.
Improvement Directions Read from Sentiment Keywords
The researchers read the main positive and negative keywords from the 6 topics with high opportunity scores and summarized the improvement directions.
- Samsung Pay: Negative keywords were concentrated on Near Field Communication (NFC) and competing services such as “Apple pay” and “Android pay.” Positive experiences and reactions to promotions were good. The researchers judged that NFC support and recognition rate need improvement.
- Fast wireless charging: This is a feature that raised charging power by 1.5 times, from 5.0V 2.0A (10Wh) to 9.0V 1.67A (15Wh). Customers were satisfied with the charging speed but unhappy about charging interruptions. A charger released later (EP-NG930) added one more coil to widen the recognition area.
- Detect pen2: In the push-to-eject design, a defect in which inserting the S Pen backwards breaks the device appeared as negative keywords. “S-Pen sensor,” “S-Pen Backwards,” and “spring mechanism” are those keywords. This defect was resolved in a later revision.
- Battery life: Software tips for extending usage time were captured, such as “Package Disable,” “background service,” and “unnecessary process.” The researchers judged that both hardware and software need improvement.
- Expandability: Many customers opposed the removal of the microSD slot. However, customers found alternatives such as cloud storage, “battery-pack,” and “portable power banks.”
- Software update: “optimization” is a positive keyword and appeared 9 times across all documents. In contrast, “factory reset” and post-update heat and battery drain emerged as complaints.
Significance and Limitations
What Changes
- By analyzing at the topic level rather than the word level, it defines the latent product attributes customers talk about.
- It does not stop at classifying sentiment; it computes an opportunity score for each topic and ranks improvement priorities.
- With the sentiment-weighted topic-keyword matrix, it gives guidance on the direction of improvement for each topic.
- Because it relies only on text data, it can be applied to services and product-service systems as well as products.
Usefulness for Practitioners
Product planners obtain both a list of attributes and an improvement order from customer posts without predefined questions. The negative keywords of the top-scoring topics become candidate improvement tasks as they are. The NFC issue in Samsung Pay and the interruptions of the wireless charger are such examples.
The researchers considered that this procedure can be implemented as a software system and hoped it would be used as a real-time customer monitoring tool. If the time at which each review was written is analyzed together, trends in customer needs can also be tracked.
Usefulness for Researchers
This paper combined topic modeling, sentiment analysis, and the opportunity algorithm into a single procedure. The paper states that various techniques, such as dictionary-based, text-classification-based, and deep-learning-based ones, can be used for sentiment analysis. From here on is my own opinion. A research design that swaps different techniques into each step and compares the results is also worth trying.
Limitations the Paper States
The paper states three limitations of its own.
- Importance and satisfaction were used as relative values normalized by the maximum and minimum. This is because it is hard to define absolute values that fit the 0–10 scale, and setting a standard criterion remains a task for future work.
- Needs in the over-served region need further examination. This region can be a clue to disruptive innovation.
- It was applied to only one product. Research applying it to other areas, such as services and product-service systems, is needed.
Conditions for Application, in My View
The following is not in the paper; it is conditions I add with practical application in mind.
- When differences in contribution stock are small, normalization stretches those differences greatly. In the case, the difference between the maximum and minimum contribution stock is 202.2198 − 165.9974 = 36.2224, and this range is spread over the whole 0–10 scale. Therefore, the ordering of mid-ranked topics should be read with care.
- If the number of topics or the keyword list changes, the maximum and minimum change and all scores change together. To compare across periods, the same keyword list and number of topics must be kept.
- Community users may differ from customers as a whole. It is safer to read social media results together with other customer data.
- Topic names are assigned by people. Having two or more people name topics separately and compare them lets you check the consistency of interpretation.
What to Try Right Away
- Collect community posts about your own product and extract keywords. After removing emoticons, onomatopoeia, and irrelevant words, exclude documents with no remaining keywords and check the number of documents and keywords.
- Run LDA with several different numbers of topics. Choose the number at which the average cosine similarity of topic pairs is lowest, then read the top-contributing keywords and documents and name each topic.
- Build a table with importance and satisfaction by topic scaled to 0–10. Compute opportunity scores with “Importance + Max(Importance − Satisfaction, 0),” and read the negative keywords of the top few topics to build a list of improvement tasks.