The same review sentence can mean different things depending on who wrote it. Take the review “The battery lasts long, and the waterproofing is good.” If a customer who bought the product a week ago wrote it, it means roughly that the new product works at a high level. If a customer who has used it for a year wrote the same sentence, it reads as saying that quality and performance hold up over long use. Yet most prior studies analyzed reviews from the entire period all at once.
A feature that is intriguing early in use can become a source of dissatisfaction the longer it is used. Conversely, a feature that gets little attention at first can become a source of satisfaction over time. So, if we split reviews by how long customers have used the product, can we find what needs to be fixed in the early stage of use and in the long-term stage?
Here we read a paper that answers this question.
- Title: Identification of Product Improvement Directions Through Customer Experience Analysis Considering Product Usage Period
- Journal: Journal of the Korean Institute of Industrial Engineers (Vol. 50, No. 4, pp. 251-265)
- Year: 2024
- DOI: 10.7232/JKIIE.2024.50.4.251
The method proposed in the paper divides reviews by the customer’s product usage period. It then calculates the importance and satisfaction of product features for each period and converts them into ranks. Finally, it places those ranks on an importance-performance analysis map for each period and tracks how the positions of the features move. The terms used throughout this post are as follows.
- Product usage period: The period from the time the customer purchased the product to the time the review was written. The paper defined the usage period as the ownership period displayed by the shopping mall.
- Product feature: An element of the product that customers mention in reviews. Each topic extracted by topic modeling (a technique for finding latent themes in a collection of documents) becomes one product feature.
- Importance: The proportion of documents addressing the feature among the reviews in a usage period.
- Performance (satisfaction): The sum of the sentiment scores of the words that make up the feature, weighted by word weights.
- Dynamic IPA Map: A map produced by repeating Importance-Performance Analysis (IPA) for each usage period.
Motivation: Look at How Long Customers Have Used the Product, Not When the Review Was Written
Why Customer Experience by Usage Period Matters
The paper defines customer experience as the sum of all experiences a customer has while maintaining a relationship with a product or service provider. Customer experience arises from the interaction between expectations and actual experience, and it includes emotion, cognition, and behavior. The paper considers this issue important for the following reasons.
- As competition intensifies, product life cycles are getting shorter. To maintain a competitive advantage, companies must continue to innovate and develop new products.
- To develop a new product, a company must understand why customers use the product and what they require. This requires analyzing and reflecting customer experience.
- A positive customer experience strengthens the relationship between customers and the company and builds loyalty. Ultimately, it leads to higher profitability for the company.
- Online reviews reflect the overall customer experience of using a product. In one survey, about 73% of users said online reviews considerably influence their purchase decisions.
- Companies must allocate limited human and material resources to the appropriate areas. They therefore need to distinguish and rank the features that require additional resources and the features that are already working well enough.
What Prior Studies Did Not Address
Prior studies that try to improve products using online Voice of Customer (VoC) fall into two broad types. The first aims at improving quality and customer satisfaction. This group includes a study that classified product attributes using the Kano model (Qi et al., 2016), a study that extracted attributes with LDA and then performed IPA (Bi et al., 2019), and a study that identified satisfaction factors with LASSO and Decision Trees (Chen et al., 2020).
The second type aims at technological innovation and the discovery of market opportunities. One study used VoC and patent data and also considered processes and supply chains (Purnama and Masruroh, 2023). Another linked reviews with patent portfolios to extend Quality Function Deployment (QFD) (Trappey et al., 2018). The paper points out three gaps.
- Most studies analyzed reviews from the entire period all at once. They therefore could not reflect the experience that changes as customers use the product.
- Most studies did not prioritize the improvement elements they extracted.
- Most studies did not use product ownership period information when analyzing reviews. The time a review was written does not represent the usage period. The review may be a mature one written after sufficient use, or it may be a light one.
In its conclusion, the paper restates that existing dynamic analyses used the product release date or the review writing date as the reference point.
Method: Split Reviews by Usage Period and Build a Dynamic IPA Map from Ranks
The paper describes the method in four steps. In Step 1, reviews are collected and preprocessed. In Step 2, product features are extracted by topic modeling. In Step 3, importance and satisfaction are calculated and ranked for each usage period, and in Step 4, a dynamic IPA map is built from those ranks to identify improvement directions. This post follows the same four steps.
Case and Data
The case product is smartwatches released since 2019. Smartwatches have features such as smartphone connectivity, exercise intensity measurement, and GPS. The paper also judged that their short development and release cycles make them suitable for applying this method. The targets are the following eight products.
- Samsung Galaxy Watch4 Aluminum Smartwatch 40mm, 44mm (released August 2021)
- Fitbit Versa 2 Health & Fitness Smartwatch (August 2019), Sense Advanced Health Smartwatch (August 2022)
- Apple Watch Series 6 40mm, 44mm Aluminum Case (September 2020)
- Apple Watch SE 1st Generation 40mm, 44mm Aluminum Case (September 2020)
Reviews were collected from the shopping mall Best Buy. Best Buy shows how long the customer has owned the product along with the time the review was written. The researchers collected 25,655 reviews written between January 1, 2020 and December 31, 2022.
The paper by Ku, Lee, and Yoon (2024) divided the usage period into four intervals, and the number of reviews assigned to each interval is as follows.
- Reviews from customers who used the product for less than 1 month (Period 1): 10,615.
- Reviews from customers who used it for at least 1 month and less than 2 months (Period 2): 9,362.
- Reviews from customers who used it for at least 2 months and less than 6 months (Period 3): 3,584.
- Reviews from customers who used it for at least 6 months and less than 1 year (Period 4): 1,483.
The longer the usage period, the more rapidly the number of reviews falls. Period 4 has about 14% of the reviews of Period 1 (1,483÷10,615=14%).
Step 1: Review Collection and Preprocessing
- Input: Online product reviews whose usage period can be determined
- Processing: Collection, removal of short reviews and stop words
- Output: Reviews cleaned up to suit the identification of product features
The processing sequence is as follows.
- Collect reviews of the target products by web crawling or an Open API.
- Remove manufacturer names, product names, and shopping mall names, which are unrelated to product features.
- Exclude reviews shorter than 70 characters, counting spaces and special characters, on the grounds that they contain no meaningful information.
- Remove stop words such as emoticons, special characters, and onomatopoeia.
- Check the usage period of each review and assign it to one of the four intervals.
Because the goal is to find product features, characters and reviews from which no feature can be identified are filtered out first.
Step 2: Extracting Product Features with BERTopic
- Input: Preprocessed reviews
- Processing: BERTopic topic modeling, hierarchical topic consolidation, selection of product feature topics
- Output: 11 product features and the key words of each feature
The paper chose BERTopic over the long-established Latent Dirichlet Allocation (LDA). The paper explains that LDA has high computational complexity and is limited in capturing complex semantic relationships between words. In studies comparing it with LDA, BERTopic is known to be generally more stable and superior in topic coherence and topic diversity. The BERTopic processing sequence is as follows.
- Convert each document into a vector using the pretrained language model BERT (Bidirectional Encoder Representations from Transformer).
- Reduce the dimensionality of the vectors with Uniform Manifold Approximation Projection. This method preserves both the local and global features of the data.
- Group similar documents with Hierarchical Density-Based Spatial Clustering of Applications with Noise.
- Extract the key words of each topic with class-based Term Frequency-Inverse Document Frequency (c-TF-IDF).
c-TF-IDF is a value indicating how important a word is within a topic. It is calculated by concatenating all the documents belonging to a topic and treating them as a single document. BERTopic also supports dynamic topic modeling, which first finds global topics over the entire period and then rebuilds the topic representation at each time step based on them. The paper applies this function to the usage period intervals and tracks how the key words change from period to period, even within the same “battery” topic.
In the case, up to two consecutive words were used as keywords. This produced 55 topics, which hierarchical topic modeling grouped into 19 major topics. Of the 19, 8 dealing with service, price, and emotion were removed, leaving 11 product features. The 11 features are as follows.
- Battery, Health care, Screen, Waterproof
- Color, Connection stability, Protection, Setup
- Software, Fall detection, Hand washing
For example, the top keywords of the waterproof feature are water, waterproof, pool, water proof, and proof.
Step 3: Calculating Importance and Satisfaction by Usage Period
- Input: Reviews by usage period, documents assigned to each topic, topic words by period and their c-TF-IDF values
- Processing: Importance ratio calculation, sentence-level sentiment analysis, satisfaction calculation as a weighted sum, ranking
- Output: Importance and satisfaction values and ranks for 11 features in Periods 1 to 3 and for 9 features in Period 4
The processing sequence is as follows.
- For each period, count the documents assigned to each feature to calculate importance.
- Split the reviews into sentences, and assign each sentence a sentiment score using VADER, a lexicon- and rule-based sentiment analysis tool.
- For each word that makes up a topic, calculate the average sentiment score of the sentences containing that word.
- Multiply each word’s average sentiment score by its c-TF-IDF value and add them all up to obtain the satisfaction for that period.
- For each period, rank importance and satisfaction.
The basis for importance is that the features customers mention in reviews are the elements they considered important during the usage period. Some features lose attention over long use and are no longer used, so the paper takes the volume of mentions to represent the degree of interest. Expressed as a formula, importance is “the proportion of documents in period p that are assigned to topic t.” For example, the importance of waterproof in Period 4 is 2.750809%. This value represents the proportion of Period 4 documents assigned to the waterproof topic.
Satisfaction is calculated at the sentence level because one review has multiple sentences, and each sentence can talk about a different feature. In addition, even for the same topic, the constituent words and weights differ from period to period, so the word composition of each period must be reflected. Taking waterproof as an example, the c-TF-IDF value of water changes from 0.268 in Period 1 to 0.252 in Period 2, 0.218 in Period 3, and 0.268 in Period 4.
Let us work through a calculation for waterproof in Period 1. The value sought is the satisfaction of the waterproof feature in Period 1. The numbers brought in are the weights of the constituent words: water 0.268, waterproof 0.225, water proof 0.109, and so on. Multiplying each weight by the average sentiment score of the sentences containing that word and adding them up gives 1.093373. The fact that this value ranks 5th among the 11 features in Period 1 becomes the input to the next step.
In Period 4, new words such as pool and shower appeared. Depending on the period, keywords that did not exist before can emerge.
Step 4: Building the Dynamic IPA Map from Ranks
- Input: Importance ranks and satisfaction ranks by period
- Processing: Map the ranks onto the IPA matrix and interpret changes in key words
- Output: Features to focus on in each period and improvement directions
IPA is a method for deciding which areas to focus on first, and in what order, to satisfy customers (Martilla and James, 1977). The four areas described in the paper are as follows.
- A1: A major strength that has achieved a standard level of performance. The status quo can be maintained without additional resources.
- A2: A major weakness. If it is not fixed with focused effort, the company may fall behind in competition, so additional resources are needed.
- A3: Performance is low, but importance is also low, so it is not a threat for now. However, it may become a minor weakness in the future.
- A4: A minor, unimportant strength. It is recommended to cut unnecessary resources and focus on other areas.
There are two existing ways to set the axes of the matrix. The scale-centered method uses the midpoint of the scale as the axis. With this method, features cluster in one quadrant, making it hard to set priorities. The data-centered method uses the mean or median of the scores as the axis, but the mean and median change from period to period, so the axes cannot be fixed.
The paper therefore performs IPA with ranks instead of scores. Using ranks fixes the axes, and the features are distributed across the quadrants rather than clustering in one. The researchers built the dynamic IPA map by mapping the importance ranks and satisfaction ranks that change with the usage period. Among these, they selected four features whose changes were steady or sharp (battery, waterproof, software, protection) to interpret, and used the changes in key keywords to propose specific improvement directions.
Results: Battery Is Always Important, and Waterproof Satisfaction Drops After 6 Months
Importance: Battery and Health Care Rank 1st and 2nd Throughout
The first comparison is the importance of the 11 features across the four usage period intervals. The higher the importance, the more interested customers are in that feature. In Table 4 of the paper, the importance of battery was 91.28641 in Period 1 and 87.05502 in Period 4, ranking 1st in all four intervals. Health care ranked 2nd in all four intervals.
The document proportion of battery is far higher than that of other topics. Battery and health care are the features customers mention with the most interest throughout the usage period. In Table 4, fall detection and hand washing ranked 10th and 11th from Period 1 to Period 3, and have no value in Period 4. The paper interpreted these two features as ones that customers mention less and lose interest in as the usage period grows.
Satisfaction: Features Whose Ranks Change Greatly by Period
The second comparison is the satisfaction in the same intervals. Satisfaction is the sum of word-level sentiment scores weighted by c-TF-IDF. In the ranks in Table 5 of the paper, three features move in different directions.
- Battery: Rises steadily, 10th, 8th, 7th, and 5th.
- Waterproof: 5th, 5th, and 3rd, then drops to 8th in Period 4.
- Software: 8th, 9th, and 6th, then becomes 1st in Period 4 with 1.655252.
In the author’s view, if the entire period is analyzed at once, such movements in different directions can be buried in a single value.
Battery: Relative Satisfaction Rises the Longer It Is Used
Battery ranks 1st in importance in all four intervals. This means battery is the element customers consider most important in a smartwatch. However, its satisfaction rank was relatively low through Period 3 (10th, 8th, 7th) and rose to 5th in Period 4.
The paper interprets this as follows. In the early stage of use, features other than the battery are perceived as attractive elements. As the usage period grows, interest in and satisfaction with the other features decline, and battery performance gives relatively higher satisfaction. It should also be noted that ranks are relative values. The satisfaction score of the battery itself fell from 0.931766 to 0.785576. Over the same period, the scores of waterproof, setup, and protection fell more than that of battery. Also, in Period 4 there are no values for fall detection and hand washing, so the number of ranked features dropped to 9. Both factors worked together in raising the battery’s rank.
The paper’s improvement direction is that companies should conduct R&D so that battery performance does not fall below that of existing products.
Waterproof: Satisfaction Drops After 6 Months, and Shower-Related Problems Are Presumed to Be the Cause
Waterproof stayed within the top 5 in importance throughout, and maintained high satisfaction for under 6 months. From Period 1 to Period 3, satisfaction was 1 or higher, but in Period 4 it was 0.4276, less than half of the previous value. Its satisfaction rank fell from 3rd in Period 3 to 8th in Period 4, while its importance rank rose from 5th to 3rd.
The cause was found in the word composition. In the Period 4 waterproof topic, words such as notice (0.076) and problem (0.074) newly appeared after pool (0.302), water (0.268), and shower (0.206). From these words, the paper inferred that problems arose with the waterproof function when customers showered while wearing the smartwatch.
There are two improvement directions.
- Establish an after-sales service (A/S) system for the waterproof function that takes the customer’s usage period into account.
- When developing new products, consider R&D measures to improve long-term waterproof performance.
Software and Protection: Features Long-Term Users Are Satisfied With
Software had low importance throughout (9th, 9th, 9th, 8th). However, in Period 4 its satisfaction jumped to 1st. The top words in Period 4 are goal (0.437), competition (0.366), neat (0.359), regret (0.293), and others. The paper interpreted that satisfaction with the function of setting and achieving exercise goals, and with the function of sharing and competing on activity levels with acquaintances, pushed up the rank.
Because interest is low relative to satisfaction, the paper suggests establishing a marketing strategy so that customers know and use this function better. It also proposed that, since software is something customers become more satisfied with the longer they use it, updates should continue to be provided.
Protection (written as Scratch in the appendix) was a feature with both low importance and low satisfaction in the early stage of use. Its importance ratio rose from 0.631068 in Period 1 to 1.618123 in Period 4, and its satisfaction rank went from 9th to 6th. The paper saw that mentions of scratches and satisfaction increased together after 6 months. On this basis, it interpreted durability related to scratches as a feature that satisfies long-term customers.
Features with Little Change: Health Care and Color
Health care ranked 2nd in importance throughout, and its satisfaction also stayed in the upper ranks without much change (4th, 3rd, 5th, 4th). The paper judged that this feature maintains decent performance. Color did not have high importance, but its satisfaction ranks were 2nd, 1st, 1st, and 2nd. The paper interpreted that the ability to change the wrist band gives steady satisfaction.
Significance and Limitations
What Changes
What the paper newly did is as follows.
- It divided reviews by the customer’s product usage period, rather than by review writing date or release date, and measured changes in importance and satisfaction by feature.
- As a result, it distinguished features that attracted attention in the early stage of use, core features customers use consistently, and features with low satisfaction in the later stage of use.
- It converted the indicators into ranks to fix the axes of the IPA matrix. In existing dynamic IPA studies, the axes changed with the data distribution in each period.
- Using BERTopic’s dynamic topic modeling, it compared the constituent words by period and found the causes of the changes in importance and satisfaction.
Usefulness for Practitioners
The paper considers that customer experience information by usage period can be used to set A/S policies and marketing strategies. In the waterproof case it proposed an A/S system that takes the usage period into account, and in the software case it proposed marketing that publicizes the function. Because features needing improvement and features that already work well enough are separated on fixed axes, it also helps decide where to spend limited resources. Because it uses a large volume of reviews, it can find customer feedback and requirements relatively faster than survey-based methods.
Usefulness for Researchers
For researchers, this is a case of using the ownership period information provided by shopping malls as a new analytical axis. The paper regards this information as an indicator representing the maturity of the customer’s product use. The paper’s design includes a dynamic IPA that fixes the axes with ranks. It also includes a satisfaction indicator that combines c-TF-IDF weights with sentence-level sentiment scores. The paper dealt with only one product category, smartwatches. The author believes that applying this design to studies of other product categories can also be considered.
Limitations Stated in the Paper
The paper states four limitations of its own.
- The step of selecting product feature topics among the topics may involve the analyst’s subjectivity. Follow-up research on an objective selection method is needed.
- The longer the usage period, the fewer the reviews, so the indicators for the later stage of use may be skewed toward a small number of reviews.
- Reviews of multiple products belonging to one product category were combined for analysis. Because feature names, materials, and sizes differ by product, analyzing individual products would be more accurate.
- Customer experience was analyzed indirectly through reviews by usage period. To see changes in experience directly, the customer’s entire product usage cycle must be tracked.
Conditions for Application, in the Author’s View
The following are not in the paper. They are conditions the author adds with practical application in mind.
- The purchase date or ownership period must be available for each review. In channels without this information, intervals cannot be divided in the same way.
- Check the number of reviews in the last interval first. In the case as well, Period 4 had the fewest reviews at 1,483. If your product has far fewer reviews, it is better to use wider intervals.
- Ranks are relative values. As with battery, the rank can rise even when the score falls, so ranks and the original scores should be viewed together.
- Features whose satisfaction rank stays low throughout, such as connection stability (11th, 11th, 11th, 9th), were not separately interpreted in the paper. In your own analysis, such features are also worth considering as improvement candidates.
What to Try Right Away
- Download your own product reviews and check whether each review has a purchase date or ownership period. If so, divide them into four intervals as in the paper (under 1 month, 1 to 2 months, 2 to 6 months, 6 months to 1 year) and count the reviews in each interval.
- After removing reviews under 70 characters and product and shopping mall names, extract topics from all reviews and keep only the product feature topics. Calculating the document proportion of each feature in each interval gives you the importance table.
- Pick one feature whose importance rank moved a lot, and compare its top words side by side across intervals. Words that newly appear in later intervals, like shower and problem in the waterproof case, become clues to the improvement direction.
The Business Intelligence Lab introduces research that finds product and technology opportunities from large volumes of data such as reviews and patents. In upcoming posts, we will break down one paper at a time in the same way.