Social Media Analytics Research: Where Does It Stand Now?

Which platforms and analytical methods are suitable for using customer opinions on social media in business decision-making? This systematic review of 57 papers examines research trends and gaps.

Translated from the Korean original. Korean original

Product planners and marketing managers watch reviews and posts pile up every day. But it is not clear which platform to collect from, what to collect, or which analysis method yields information usable for decision-making. There are many individual application papers, but it is hard to find a single source that shows where research is concentrated and where gaps remain.

Can customer opinions accumulated on social media serve as material for business decisions? If so, what have studies obtained so far, and with what data and methods?

Let us read one paper that answers these questions.

  • Title: Social media analytics and business intelligence research: A systematic review
  • Journal: Information Processing & Management
  • Year: 2020
  • DOI: 10.1016/j.ipm.2020.102279

The paper does two things. First, it divides open data into six categories and compares them on four dimensions using evaluations by 10 experts. Second, it collects 633 papers from Web of Science (WoS), narrows them to 57, and analyzes them with three research questions.

The terms used throughout this post are as follows.

  • Business Intelligence (BI): The whole process of obtaining, interpreting, and analyzing information in a business context to draw out useful knowledge.
  • Voice of Customer (VoC): Opinions that customers leave about a product or service. This post deals with online VoC left on social media.
  • Open data: Information sources that are not classified as confidential and are accessible to anyone. This is a different concept from open-source software.
  • Systematic review: A review method in which the search string and exclusion criteria are set in advance to collect the literature, which is then analyzed according to predetermined questions. The paper followed the guidelines of Kitchenham (2004).
  • Application Programming Interface (API): The official access method through which a platform provides data to outside parties.
  • Sentiment analysis: A method that computes, as a score, the polarity (positive or negative) of the sentiment expressed in a text.

Motivation: No Review Had Brought Social Media-Based BI Research Together

Why Customer Opinions on Social Media Matter

The paper holds that social media has become the easiest place for customers to voice opinions, and it offers several pieces of evidence.

  • Scale of use: About 79% of the U.S. population had social media accounts, compared with 70% in East Asia and 67% in Northern Europe.
  • Influence of reviews: About 77% of customers read other customers’ online reviews, and 75% trust reviews more than personal recommendations. More than 80% of customers get useful information from social media channels.
  • Uncensored opinions: Customers post opinions on channels such as Twitter, Facebook, and Amazon with little censorship. These opinions influence other customers’ purchase decisions.
  • Corporate use: Samsung mobile opened the “Samsung Global” Facebook page in 2010 and gathered about fifty million followers. Starbucks opened “My Starbucks Idea” in 2008 to collect customer ideas.
  • Need for market-driven innovation: For products such as smartphones, whose underlying technology is already mature, it is hard to innovate through in-house R&D alone. An approach that finds development directions in customer needs is therefore required.

For example, Jeong et al. (2018) collected Samsung Galaxy Note 5 reviews from Reddit. Using topic modeling, the researchers found complaints about battery life, touch recognition, and the camera. They then used sentiment analysis and an opportunity algorithm to pinpoint features with high importance and low satisfaction.

What Previous Research Did Not Cover

Early BI research mainly used traditional data such as patents, research papers, and commercial materials. Park and Yoon (2017) identified firms’ technological capabilities from their patents and recommended technology opportunities using collaborative filtering. Wang et al. (2018) grouped patents by classification codes and computed an emergence indicator to find newly rising topics.

Attention then shifted to customer-generated data. Chung and Tseng (2012) emphasized that online product reviews are a major source for understanding customers and markets. There was agreement that social media data suit BI. But which open data are best suited to BI research remained a matter of debate.

The paper points to two gaps.

  • Few studies covered the overall stream of social media-based BI research. Reviews related to social media existed, but they took different perspectives.
  • Studies comparing multiple open data sources, such as social media and intellectual property, from a BI perspective were lacking.

Separately, the paper claims one contribution: for new researchers and practitioners, it provides a guide that organizes the procedure for collecting, preprocessing, and analyzing social media data.

Method: Evaluating Open Data, and Narrowing 633 Papers to 57 for Analysis with Three Questions

The paper proceeds in two broad parts. The first part classifies and evaluates open data, and the second is a systematic review. In this post, the whole is divided into four steps.

Figure 1. A four-step procedure that classifies and evaluates open data, then narrows 633 papers to 57 and analyzes them with three questions.
  • Step 1: Classify open data.
  • Step 2: Compare the data through expert evaluation.
  • Step 3: Search for and select the literature.
  • Step 4: Analyze the literature with three research questions and an integrated analysis framework.

The paper also has two reasons for choosing a systematic review over a concept-centered review. First, social media-based BI research started from real business decision problems rather than from theory. Second, the authors wanted to see recent trends in methods, algorithms, and data, rather than trace the development path of theory through citation relationships.

Data and Experts

There are two objects of analysis. One is the experts who took part in the evaluation, and the other is the papers collected from WoS.

  • Experts: Ten people engaged in BI research. The 4 professors had conducted BI research for about 10 years using open data such as social media, research proposals, and patents. The 6 graduate students had studied social media mining, customer understanding, and data-driven business planning for about 4 years.
  • Literature: The search was run on WoS on July 4, 2018. The period was limited to the most recent 5 years, and the document type to peer-reviewed journal articles.
  • Scale: The search returned 633 papers, and the final analysis set was 57.

Step 1: Divide Open Data into Six Categories

  • Input: Prior literature on open information sources
  • Process: Define categories and attach examples
  • Output: Open data classification table (six categories)
  1. Starting from the definition in the field of open-source intelligence, set the scope of open data.
  2. Divide categories based on related literature, and attach real examples to each category.

The reason for dividing this way is to compare social media data with other open data on the same standard. The categories and examples are as follows.

  • Mass media: Printed newspapers, magazines, radio, television
  • Commercial data: Commercial imagery, financial and industrial assessment reports
  • Grey literature (open materials not commercially published): Intellectual property, newsletters, technical reports, bibliographies
  • Public government data: Government reports, government open data
  • Professional and academic publications: Information from journals, conferences, and symposia
  • Web-based data: Social media, online discussion groups, online publications

Patents and trademarks fall under grey literature. They are often used in BI research because they describe technology and business fields.

Step 2: Evaluate the Data on Four Dimensions

  • Input: Six categories of open data
  • Process: Ten experts rate each on four dimensions as high, medium, or low
  • Output: Evaluation table by category
  1. The experts decide on four evaluation dimensions. These reflect attributes common to the process of collecting and analyzing data and the process of continuously operating a BI system.
  2. The experts assign high, medium, or low to each category on each dimension.
  3. The researchers gather the experts’ opinions and organize them into a single evaluation table.

The four dimensions mean the following.

  • Data content: How much the subject matter, the expertise, and the users’ level of knowledge help business decision-making
  • Data collection: Whether there is a public API that allows large volumes to be collected
  • Data updating: Whether up-to-date data accumulate and the database is continuously updated
  • Data structure: Whether the data can be processed into a structured format that is easy to analyze

The collection and updating dimensions are closely related to the systematization of BI research. The evaluation results are summarized in the results section.

Step 3: Search for and Select the Literature

  • Input: WoS database
  • Process: Obtain a first list with the search string and narrow it with exclusion criteria
  • Output: First list of 633 papers (Group 1), final list of 57 papers (Group 2)
  1. Set the search string. Because different researchers write the same topic with different words, the authors joined key terms with ‘or’ rather than ‘and’ to find as many papers as possible.
  2. Review papers with “systematic review” in the title are removed, and “sentiment”, a frequently used method, is included in the search terms.
  3. Narrow the search to the most recent 5 years and to journal articles.
  4. Read the titles and abstracts of the first list and apply the exclusion criteria.

The search string is “TITLE: ((social and media or ((product or service) near (review))) and (mining or analy* or sentiment or business)) NOT TITLE: (systematic and review)”.

There are nine exclusion criteria.

  • EC1: Include only peer-reviewed journal articles.
  • EC2: Exclude book reviews, conference abstracts, and editorials.
  • EC3: Exclude papers whose title or abstract does not sufficiently describe the research content.
  • EC4: Exclude papers that cannot be accessed.
  • EC5: Include only papers in English.
  • EC6: Exclude papers with duplicated content.
  • EC7: Include only research papers and exclude literature survey papers.
  • EC8: Exclude papers that did not analyze social media for BI purposes. Papers that used social media only as input to show the performance of a new algorithm fall here.
  • EC9: Exclude papers that did not draw out information for business decision-making.

The selection results are as follows. Of the 633 papers on the first list, 10 were duplicates, and 566 were judged unrelated after reading the abstracts. Thus 633−10−566=57 papers remained as the final analysis set. Conference papers were excluded because the authors considered the corresponding journal articles to represent the latest results.

Step 4: Analyze the Literature with Three Questions and an Integrated Analysis Framework

  • Input: 57 final papers
  • Process: Organize data, methods, and results according to three research questions
  • Output: Platform distribution, classification of analysis approaches, classification of derived information

The three research questions are as follows.

  1. RQ1: What are the promising social media platforms used in this field?
  2. RQ2: What are the main methods and algorithms that researchers apply to social media data?
  3. RQ3: What information does social media-based BI research draw out?

The researchers organized the analysis procedures of the 57 papers into a four-step integrated analysis framework.

  1. Data collection: Collect text, images, and video through public APIs or web crawling to build a database. Considering security and privacy, the paper recommends that using a public API is best.
  2. Preprocessing: Extract keywords with natural language processing and remove stop words. Emoticons (”^^”, ”:-D”), web links (“www”), onomatopoeia (“haha”), generic words (“product”, “device”), and unrelated words (“day”, “month”) are removed. The text is then converted into a structured format such as a document-term matrix.
  3. Analysis: Apply various algorithms.
  4. Validation and interpretation: Validate the results and interpret their meaning.

In preprocessing, some researchers selected important keywords using Term Frequency (TF, the number of occurrences within a document) and Document Frequency (DF, the number of documents in which the word appears). For short texts, some used word embedding results directly without complex preprocessing.

Analysis methods were classified into five approaches according to depth of use and intent: sentiment analysis, topic modeling, Machine Learning (ML)-based approaches, network-based approaches, and theory-based approaches.

Results: Studies Analyzing Commerce Reviews and SNS with Sentiment Analysis and Machine Learning Are the Core

Figure 2. The 57 papers concentrate on commerce review and SNS data and on sentiment analysis and ML approaches.

Open Data Evaluation: Grey Literature and Web-Based Data Rated High

The key points of the evaluation results are as follows. The full values by category can be found in Table 3 of the original paper.

  • Grey literature: High on all four dimensions.
  • Web-based data: High on content, collection, and updating, and medium on structure.
  • Mass media: Low on all four dimensions. This is because it consists of physical materials such as print, and converting it into an analyzable format takes a lot of work.
  • Public government data: Rated lower than grey literature, and especially low on updating.

Grey literature was rated high because official institutions publish and archive it regularly, as with patents. A database can therefore be built systematically.

The strength of web-based data is that users create it in real time, so updating is fast. Most platforms have public APIs, and where there is no API, data can be collected by web crawling. Structure is medium because the data are unstructured, mixing typos, emoticons, and abbreviations, and so require preprocessing.

Public government data were rated low. Data.gov provides more than 300,000 datasets, but few of them are directly related to business opportunities or customer opinions. Moreover, publication timing and formats vary, making it hard to collect them continuously.

Based on the high scores for collection, updating, and structure, the paper concluded that social media is more suitable for BI analysis than other open data. In the values of the evaluation table itself, grey literature is high on all four dimensions, and web-based data is medium on structure. The author’s opinion on how to read this difference is written separately in the application conditions near the end of this post.

Research Trend: The Number of Papers Increased Every Year

The yearly counts for the 633 papers on the first list rose from 2014 as 83, 118, 150, and 180. In 2018, nearly 100 papers had already appeared by July. A similar increasing trend appeared in the final list of 57.

The journals that published the papers were mainly in information processing, intelligent systems, and expert decision support. Some papers also appeared in journals on marketing, service management, e-commerce, tobacco control, and industrial safety. This means social media analytics is being used for different purposes in various business environments.

RQ1: Commerce Reviews and SNS Account for Most

The researchers divided the data used by the 57 papers into four types: 29 papers on commerce reviews, 24 on SNS, 2 on online discussion groups, and 2 on combined data. The total is 29+24+2+2=57 papers.

Among commerce reviews, Amazon (6 papers), Yelp (5), and TripAdvisor (4) were widely used. Among SNS, Twitter (11) was the most common. Both papers on online discussion groups used Reddit. The full distribution by platform can be found in Table 4 of the original paper.

Among individual platforms, Twitter, Amazon, Yelp, and TripAdvisor were frequently used. The reasons for use and the difficulties of each type are as follows.

  • Commerce reviews: Reviews accumulate by product or service, so topics are unified. Therefore, little effort is needed to filter the data after collection. Of the 29 papers, 14 are product reviews and 15 are service reviews (14+15=29). However, abbreviations, emoticons, and autocompleted sentences are mixed in, so cleaning is needed.
  • SNS: Many users create posts in real time, and most platforms have public APIs. On Twitter, more than five hundred million tweets are posted per day on average. On the other hand, because of the 140-character limit there are many abbreviations and emoticons, so systematic preprocessing is needed. A study analyzing Flickr photos collected more than one hundred thousand items.
  • Online discussion groups: Reddit has more than 120 ten-thousands (one million two hundred thousand) subreddits (topic-specific boards). Subreddit users have much knowledge of or interest in the topic, so posts are specific and rich in content.
  • Combined data: Studies analyzed social media and patents together to see a firm’s technological capabilities and customer opinions at once. There is one study each combining Amazon with patents and Twitter with patents.

The combined-data studies also had points to improve. One study did not consider the gap between patent publication dates and product release dates. According to empirical studies, there is a time lag of 3 months to 3 years between patent publication and market release. Another study directly linked patents and reviews through an ontology (a knowledge system defining concepts and relationships). However, this approach requires much expert effort in building the ontology.

RQ2: Sentiment Analysis and Machine Learning Were Used Most

The number of papers and the purpose of each of the five analysis approaches are as follows.

  • Sentiment analysis, 27 papers: Determines the polarity of opinions about specific products, services, or brands. There are lexicon-based, text classification-based, and deep learning-based variants.
  • Topic modeling, 13 papers: Finds topics mentioned by customers and the public. Latent Dirichlet Allocation (LDA) and Latent Semantic Analysis (LSA) were used.
  • ML-based approaches, 27 papers: Classify sentiment polarity or predict stock prices, product sales, and sentiment. Naive Bayes, random forest, Support Vector Machine (SVM), hidden Markov models, and others were used.
  • Network-based approaches, 7 papers: Find customer interaction networks and influential users.
  • Theory-based approaches, 6 papers: Express customer satisfaction, service life cycle, and service quality numerically. Fuzzy theory, game theory, the opportunity algorithm, and the Kano model were used.

One paper can belong to several approaches at the same time. The sum across the five approaches is therefore 27+13+27+7+6=80 papers, more than the total of 57.

There are two reasons sentiment analysis was widely used. Results come out as scores, so customer reactions can be handled numerically, and it captures polarity relatively accurately even in short, low-quality text. However, the paper holds that to obtain better results, ML algorithms, network analysis, and theory should be used together.

Among ML classifiers, SVM was used most consistently over the analysis period of 2013 to 2017. This is because it performs well on high-dimensional, sparse data such as social media posts in which each customer writes with different words. Newer algorithms such as Convolutional Neural Network (CNN) and Long Short-Term Memory (LSTM) were also used. One study found product safety problems using bidirectional LSTM.

Studies that predicted stock prices or sales share a common feature: a tendency to use the sentiment scores of social media posts as independent variables. Most studies compared predictive performance with existing models using precision, recall, accuracy, and F-measure.

RQ3: Information Was Drawn Out from Four Kinds of Objects

The researchers divided the objects analyzed in the 57 papers into four kinds.

  • Customer-generated content: Discovering product and business opportunities, analyzing competitive position in the market, monitoring quality and detecting safety risks, identifying customer needs and satisfaction, investigating service quality and gaps
  • Relationships among customers: Detecting potential customer communities, identifying influential users in a network, extracting information on product adopters, analyzing tourist behavior
  • Customer reactions: Predicting stock price movements, monitoring customer reactions, detecting abnormal sentiment and adverse effects, predicting online sales, monitoring emerging technology development trends
  • Characteristics of social media data: Examining the theoretical basis of review helpfulness, investigating predictors of review reading and helpfulness

Some studies proposed systematic methods, such as recommender systems or expert support systems. But most stopped at one-off analyses that grasp customer satisfaction, loyalty, and perception.

Significance and Limitations

What Changes

  • Six categories of open data were compared on the four dimensions of content, collection, updating, and structure.
  • Fifty-seven social media-based BI papers were classified with the three questions of data, methods, and results.
  • A four-step integrated analysis framework of collection, preprocessing, analysis, and validation and interpretation was organized.
  • A list of public APIs by platform was compiled as an appendix. It includes Twitter, Facebook, Amazon, Reddit, Yelp, TripAdvisor, and others.

Value for Practitioners

Practitioners can refer to this classification when choosing platforms and methods suited to their own business environment. For example, reviews have unified topics and so are well suited to direct use in product improvement. To find influential users, studies that analyzed user relationships on SNS as networks are a useful reference.

The integrated analysis framework can be used as a checklist when designing an in-house analysis procedure. In particular, the paper stresses that to run analysis as a system rather than a one-time exercise, a systematic collection method such as a public API is needed.

Value for Researchers

The paper presents three directions for future research.

  • Beyond sentiment analysis: Researchers should not stop at whether customers say something is good or bad. Quantitative approaches are needed that trace the specific reasons for negative reviews and discover new features.
  • From one-off analysis to systematic frameworks: Research should develop into methodologies reproducible on other data and into expert support systems.
  • Combining heterogeneous data: Social media and patents should be analyzed together. The paper proposes using the product names and applicant information in trademarks as links, or comparing customer terms with the technical terms of patents using word embeddings.

Limitations Stated by the Paper

The paper states five limitations itself.

  • Only one major academic database (WoS) was searched. Relevant papers not in WoS may therefore have been left out of the analysis.
  • Conference papers were excluded. Because some conference papers are highly influential, later studies could include them.
  • A concept-centered approach that analyzes citation relationships could be applied over a longer period.
  • To analyze from more diverse perspectives, more research questions should be added.
  • The paper did not address the language of social media or differences across business sectors such as tourism, healthcare, hotels, and electronics. Most of the collected papers used English or Chinese data.

Conditions for Application as Seen by the Author

What follows is not content from the paper but conditions the author adds with practical application in mind.

  • Conclusion on data comparison: The paper concluded that social media is more suitable for BI than other open data. But in the evaluation table, grey literature is high on all four dimensions, and web-based data is medium on structure. The author thinks it is right to read this conclusion as limited to the perspective of collection and updating. For tasks where structured data matter, it is safer to examine grey literature such as patents together.
  • Search date: The literature search was conducted on July 4, 2018. Changes in methods and platforms after that are not included in this review.
  • API status: The public API list in the appendix reflects information at the time the paper was written. Platform access policies can change, so check each platform’s current terms before use.
  • Evaluation method: The open data evaluation is a qualitative assessment by 10 experts. In your own business field, it is safer to re-evaluate your own data on the same four dimensions.
  • Language: The reviewed papers were limited to English, and the analyzed data were mainly in English and Chinese. When handling Korean reviews, you must separately prepare morphological analysis and a stop-word list.

What to Try Right Away

  1. Process your own product’s reviews in the order of the integrated analysis framework. Collect reviews through a public API or crawling, remove emoticons and generic words, and build a document-term matrix. Apply LDA to this matrix to extract the topics customers talk about most.
  2. Attach a sentiment score to each topic. As in Jeong et al. (2018), topics with high importance but low satisfaction become the first candidates for improvement priorities.
  3. Match the topics found in the reviews against your own patents. Here, set the comparison period by taking into account the lag between patent publication and product release (3 months to 3 years).

Keywords

Related posts