Planners who want to know which products and services a competitor will launch next usually depend on expert opinion. As the relationships among products, companies, and customers have grown more complex and product lifecycles have shortened, expert-centered analysis has demanded more time and effort. Patents and social media can also be analyzed. However, the results come out as bundles of keywords, so interpreting those bundles as actual businesses again takes considerable effort.
So, if we convert the names of goods and services written in trademarks into vectors with a language model, and use outlier detection to find points that stand apart from other businesses, can we spot new business opportunities in advance?
Here we read one paper that answers this question.
- Title: Identification of emerging business areas for business opportunity analysis: An approach based on language model and local outlier factor
- Journal: Computers in Industry
- Year: 2022
- DOI: 10.1016/j.compind.2022.103677
The researchers proposed a method that combines a language model with the Local Outlier Factor (LOF). The language model groups goods and services with similar meanings to define business areas. LOF measures numerically how novel each business area is. In the final step, the method builds a business opportunity map that places the novel areas by recency and visibility.
The terms used throughout this post are as follows.
- Designated goods and services (G&S) statement: the sentence that lists, when a trademark is filed, the goods and services on which the trademark will be used.
- Business item: a single name of a good or service, obtained by splitting a G&S statement at delimiters.
- Business area: a cluster of business items with similar meanings.
- Novelty: how different a business item is from existing goods and services in the way it is expressed.
- Recency and visibility: recency is how recently a business area appeared, and visibility is how often a business area is seen.
- Business landscape: the scope of data in which business opportunities are sought.
Motivation: Few studies have measured numerically how novel a business area is
Why emerging business areas should be found early
The reasons the paper gives are as follows.
- As uncertainty and volatility in the business environment increase, discovering business opportunities has become more important.
- Emerging business areas are early indicators of potential business opportunities. They are key to building new business strategies and anticipating the near-future business environment.
- Existing analyses that rely only on expert opinion take a great deal of time and labor. For that reason, managers and practitioners increasingly use data-driven approaches.
- Successful innovations are closely related to high novelty. It is therefore reasonable to look first for areas made up of new goods and services.
This is also why the researchers chose trademarks as their data. Companies file trademarks to indicate the source of ongoing businesses. They also file them to announce and claim new business areas they are about to start. Because trademarks describe goods and services in a standardized format, they are also well suited to computational methods.
What existing studies have not addressed
Existing studies have used four types of data to find business opportunities.
- Patents: show opportunities at the stage of technology development or technology commercialization.
- Web news and social media: deal with content about the future or with customer reactions after launch.
- Business plans of startups: provide clues to early business ideas.
Most studies follow three steps: defining the business landscape, dividing business areas, and evaluating opportunities. The paper points out the following gaps in each step.
- Business landscape: Because researchers select and collect data on a specific technology or product, the boundaries of the scope are unclear. Even when search queries are carefully constructed, it is hard to reproduce the same results in practice.
- Business area definition: There is no set criterion for how many areas to divide into or how. Even when the same scope is analyzed, areas are defined differently depending on the researcher’s method or judgment. Topic modeling such as Latent Dirichlet Allocation (LDA) produces bundles of keywords, and these bundles are hard to interpret as actual businesses.
- Opportunity evaluation: Growth, popularity, and visibility have been analyzed. In contrast, methods for expressing the novelty of a business area numerically have rarely been addressed.
- Timing of the data: Existing data are close to the technology development stage or the post-launch market stage. Trademark data, which are closer to the business operation stage, have not been used.
Method: Convert the names of goods in trademarks into vectors, and find areas that depart from dense regions with LOF
The method consists of four steps. In Step 1, the business landscape is defined and trademark data are collected and organized. In Step 2, G&S statements are split into business items, vectors are created with a language model, and the vectors are grouped into clusters to define business areas. In Step 3, novelty is computed for each business area with LOF, and in Step 4, recency and visibility are added to build the business opportunity map. This post explains the method in the same four steps as the paper.
Case and data
The data source is the United States Patent and Trademark Office (USPTO) database. The USPTO has administered trademarks since 1870. Detailed information such as legal events, G&S statements, and classification codes can be searched in electronic form. The researchers used three kinds of information.
- Nice Classification (NCL) code: divides goods and services into 45 classes, and was used to narrow the business landscape.
- Filing date: used to determine the timing of business operations.
- G&S statement: used to extract specific business items and define business areas.
The case is NCL class 09, which covers scientific and research apparatus and devices. This class contains many technology-intensive products. Trademarks related to virtual reality devices, lithium-ion batteries, Bluetooth smart devices, artificial intelligence applications, and cryptocurrency are steadily increasing. A 2019 trademark statistics report also showed NCL 09 as a scope in which filings keep rising.
The analysis covers 134,898 trademarks filed in NCL class 09 and their G&S statements. The analysis period runs from early 2017 to the end of 2018, the most recent data available at the time of analysis.
Step 1: Data collection and preprocessing
- Input: raw USPTO trademark data
- Process: define the business landscape, collect trademarks, then parse and store them
- Output: a trademark database containing both structured and unstructured fields
- Define the business landscape. It can be defined narrowly, as trademarks containing specific keywords, or broadly, as all trademarks falling under a single NCL code.
- Collect trademarks from the USPTO according to the chosen criteria.
- Parse the raw data and store it in a database. The NCL code and application number are stored as structured fields, and the trademark name and G&S statement as unstructured fields.
There is a reason to define the scope first. Novelty and emergence are relative concepts that are determined only by comparison with other areas. In the case, the whole of NCL class 09 was set as the scope, and 134,898 trademarks were collected as a result.
Step 2: Business item extraction and business area definition
- Input: the G&S statement written for each trademark
- Process: sentence cleaning, item extraction, language model vectorization, clustering
- Output: business areas made up of business items with similar meanings
The USPTO provides the Trademark Identification Manual, which contains common names of goods and services. However, many applicants use new expressions to differentiate their products. Because there are too many business items with differing expressions for a person to analyze one by one, the following four substeps are applied.
- Sentence cleaning: Remove legal procedure expressions mixed into the G&S statement. For example, in “(Based on 44(e) Priority Application) Smart phones; Television receivers; Monitors for computers; Commercial monitors”, the priority information in parentheses is removed.
- Item extraction: Split the statement at delimiters. The statement above becomes four items: smart phones, television receiver, monitors for computers, and commercial monitors.
- Vector representation: Use a pretrained language model to convert each item into an n-dimensional vector. Items with similar meanings or expressions are placed close together in the vector space.
- Clustering: Group similar items with an unsupervised clustering method such as k-means. Each cluster becomes a business area. The meaning of an area is interpreted by looking at the items closest to the cluster centroid.
There is also a reason the Bag-of-Words (BoW) approach was not used. BoW treats each word as one dimension, so when unique words exceed several thousand, it suffers from the curse of dimensionality (the problem that computation becomes difficult when there are too many dimensions). It also discards word order and context, so it has difficulty distinguishing expressions in which the same words are arranged differently.
In the case, 254,063 business items were extracted from the G&S statements. The researchers filtered the items with two criteria.
- Items with a Document Frequency (DF) of 1, that is, items appearing in only one trademark, were removed.
- Only items not in the Trademark Identification Manual were kept, because items absent from the manual are more likely to be differentiated goods and services.
73,048 items remained. The researchers converted these items into 1,024-dimensional vectors with a RoBERTa model and grouped them with k-means. The number of areas was set by experts’ judgment, with reference to inertia values, because the number of areas changes how specific the business opportunities are.
Step 3: Evaluating the novelty of business areas
- Input: the centroid vector computed for each business area
- Process: compute LOF on the centroid vectors
- Output: a novelty value for each business area
The unit of analysis is the business area. The centroid computed from the item vectors that make up each area is fed into LOF. LOF calculates how isolated a target is by comparing it with its surrounding neighbors. The calculation proceeds as follows.
- k-distance: Find the Euclidean distance from the target p to its k-th nearest neighbor.
- Reachability distance: The reachability distance between p and a neighbor o is the larger of “the k-distance of o” and “the actual distance between p and o.”
- Local reachability density (lrd): It becomes larger the shorter the average reachability distance from p to its neighbors.
- LOF: The average density of the neighbors is divided by the density of p.
If the LOF value is close to 1, the density of p is similar to that of its neighbors, so it is a normal point. If it is greater than 1, p lies in a sparser place than its neighbors, so it is likely to be an outlier. The paper explains this with three cases.
- If p is in the middle of a dense region, both p and its neighbors have high density, so LOF is small.
- If p is in the middle of a sparse region, both have low density, so LOF is small.
- If p is in a sparse place while its neighbors form a dense cluster, LOF is large. This case is the outlier.
Consider one example. In Table 2 of the paper, the novelty value of the “monopod” area is 1.366. The paper explains only that the novelty value is based on LOF. Values for other areas can be found in Table 2 of the original paper.
There is also a reason to look at density rather than distance-based clustering. Even when product names are close in distance in the embedding space, changing a single word can indicate a major technological advance. Distance alone therefore makes it hard to judge novelty.
The number of neighbors k is set by experts. In the case, it was set to 20 on the basis of expert opinion. As a result, 103 business areas emerged as novel areas.
Step 4: Building the business opportunity map
- Input: novel business areas and the filing dates and frequencies of the trademarks belonging to them
- Process: compute recency and visibility and place the areas on a two-dimensional plane
- Output: a business opportunity map divided into four regions
Even a novel area offers a greater chance of opportunity when it is also growing quickly. The researchers therefore added recency and visibility.
- Compute recency. Recency is how recently a business area appeared in trademarks, and it is measured by the time of appearance.
- Compute visibility. Visibility is how prominent a business area is within the business landscape, and it is measured by frequency.
- Place the areas on the map. Each novel area is one point. Point size is novelty, the horizontal coordinate is recency, and the vertical coordinate is visibility.
- Use the mean values of normalized recency and visibility as boundaries to divide the map into four regions.
The four regions mean the following.
- Emerging region (high recency, high visibility): An area that appeared recently but already leads the trend. It is the most valuable opportunity.
- Post-emerging region (low recency, high visibility): Its influence is large, but it may be saturated. It must be investigated in detail to find niche market opportunities.
- Marginal region (low recency, low visibility): Observed for a long time and small in share. Unless there is a major innovation, it can be dropped from the opportunities.
- Pre-emerging region (high recency, low visibility): Newly appeared but still small in count, so it needs continued observation.
In the case, the “smart watch (745)” area fell in the emerging region. The novelty, recency, and visibility values for each area are summarized in Table 3 of the paper, so refer to the original. The paper does not report the mean boundary values themselves. However, because the area was classified as emerging, the author infers that both of its coordinates were above the mean boundaries.
Variations and scenarios
Results vary with the settings and the purpose of the analysis.
- Number of neighbors k: A company exploring new areas broadly sets k large. This takes more adjacent areas into account and yields a macro-level result. A company interested in incremental innovation narrows the scope to adjacent areas.
- Number of business areas: The number of areas directly affects how specific the resulting business opportunities are.
- Excluding the company’s own areas: Areas the company already operates are removed from the map. If experts exclude a specific area from the analysis, LOF and the indicators must be recalculated and the areas placed again.
- Timing: The map is a snapshot at the time of analysis. For practical use, the data and results must be updated continuously.
- Distance and outlier detection methods: Depending on the data, other distance measures or outlier detection methods can be substituted.
Results: Of the 103 novel areas, 16 were emerging areas, and these areas had distinctly more applicants
103 novel business areas
The LOF calculation produced 103 novel areas. The areas the paper gives as examples are as follows.
- Monopod
- Bluetooth communication devices
- Robot exoskeleton suits
- Lithium-ion battery
Looking closely at two of the areas: the monopod area consists of selfie sticks for smartphones and cameras, and handheld monopods. The lithium-ion battery area consists of lithium-ion polymer cells. The novelty values of the two areas are in Table 2 of the paper.
Calling an area novel means that, in the embedding space, its expressions lie apart from those of other areas. There is a reason that products that look common now, such as selfie sticks and Bluetooth speakers, are included. Novelty is a relative value compared with other areas within the same business landscape. Moreover, the analyzed data are trademarks from early 2017 to the end of 2018. This result should therefore be read as meaning that the areas were relatively new within NCL class 09 at that time.
The four regions of the business opportunity map
The researchers found the trademarks for each of the 103 novel areas and computed recency and visibility. The number of areas in each of the four regions is below, and the total is 103 (16+33+40+14=103).
| Map region | Number of areas | Next action |
|---|---|---|
| Emerging | 16 | Investigate early movers and competitors |
| Pre-emerging | 33 | Keep observing |
| Marginal | 40 | Mostly exclude from analysis |
| Post-emerging | 14 | Check for saturation |
Source: Section 4.2 of the paper
In the emerging region, the following areas were named as novel areas with high opportunity potential.
- Encryption software (687)
- Smart watch (745)
- Artificial intelligence software (662)
The pre-emerging region included the following areas. These have weak visibility signals and need continued observation.
- Lithium-ion battery (719)
- Robot exoskeleton suits (948)
- Voice-controlled smart home devices (991)
- Self-checkout (988)
The lithium-ion battery was named a novel area but was placed in the pre-emerging region. This is because on the map novelty is used only as point size, not as a coordinate.
Quantitative validation: Comparing the number of applicants in emerging and pre-emerging regions
The researchers compared the number of applicants in the emerging and pre-emerging regions. This follows prior research holding that the more actors (communities) are active in an area, the higher its emergence. The pre-emerging region was used as the comparison group because it is the most likely to grow into an emerging region but has not yet acquired emergence. The test is a two-tailed t-test for two groups with different sample sizes and variances. The null hypothesis is that the means of the two groups are equal.
The results are in Table 4 of the paper. The 16 emerging areas had a much higher average number of applicants than the 33 pre-emerging areas. Even within the emerging areas, the number of applicants varied widely from area to area. The mean and standard deviation for each group can be found in Table 4 of the original.
The test statistic is t=5.11. The t value is the difference between the two means divided by the standard error. The paper reports p as 0.000, which means the value is so small that it is displayed as 0 after rounding.
The paper’s interpretation and the author’s assessment should be read separately. The researchers interpreted this result as roughly supporting the claim that the proposed method finds emerging areas. However, this comparison is not an independent validation. The two groups were split from the start by visibility, that is, trademark frequency. Areas with many trademarks tend to have many applicants. This comparison is therefore closer to a re-confirmation of the difference between groups divided by visibility, and it is appropriate to read it only as showing a trend consistent with the method’s results.
Qualitative validation: Tracking later corporate activity
The researchers held that numerical comparison alone cannot prove the quality of the method. They therefore tracked how major companies in the emerging regions acted afterward. This tracking is not evidence that the method predicted corporate activity. It is a post hoc check that later corporate activity does not contradict the results.
- The solar power panel area also appeared as an emerging area on all three indicators: novelty, recency, and visibility. The three areas above are examples the paper picked as representative from the map results, while solar panels are an example of an emerging area cited separately in the validation section. The paper describes how Hanwha Q CELLS, which is closely related to this area, kept growing its solar business through mergers and acquisitions and released technology with higher efficiency than existing panels.
- In the encryption software area, HashiCorp became a major company in encryption solutions that protect sensitive data through central key management. The paper describes how Tencent and Tessian, which were involved in this area early on, now provide data security and encryption businesses to customers.
- The paper describes how lithium-ion batteries, smart home devices, and self-checkout in the pre-emerging region later became essential parts of many products and services. According to the paper, the lithium-ion battery market grew through the spread of electric vehicles and demand for smartphones and wearables. The smart home device market grew rapidly as Internet of Things devices increased. Self-checkout grew in kiosk form alongside demand for contactless consumption.
Because emergence is usually judged only in hindsight, deciding how emergent something is is difficult and prone to subjectivity.
Significance and limitations
What changes
- Data: This is the first attempt in the field of business opportunity discovery to apply text mining to trademark G&S statements.
- Unit of results: The analysis results became more concrete, moving from the keyword level to the business item level, that is, the level of goods and services.
- Indicator: The novelty of business areas, which existing studies rarely addressed, was quantified with LOF.
- Timing: Analysis that had stayed at the technology development or market forecasting stage was moved to the early stage of business operation, that is, close to the time of trademark filing.
- Structure: The one-shot existing methods were turned into an interactive structure in which experts adjust and interpret the results at each step.
- Model: It is among the first studies to apply a state-of-the-art pretrained language model to business opportunity identification.
Usefulness for practitioners
The paper expects that business opportunities can be explored over a wide scope, greatly reducing the time and effort spent on analysis. However, the actual savings were not measured. Experts still intervene in deciding the number of areas and neighbors, selecting business items, and interpreting areas.
Because the results come out as lists of goods and services rather than bundles of keywords, they are easy to use directly in decision-making. The procedure is systematic, so it can be built into automated software. This means that field experts who lack data analysis skills can also use it.
Once the business landscape is defined, adding new trademark data is also simple. One only needs to find the G&S statements of the new trademarks and redraw the business landscape. The method can also be used together with existing methods. For example, after finding emerging areas, examining products, markets, and competitors through social media analysis can make the opportunities more concrete. Even so, judging whether the results are real opportunities remains the job of experts and managers.
Usefulness for researchers
- The method can also be used for technology forecasting, future signal detection, and competitive analysis.
- With a few changes in technique, it can be applied to other data sources such as patents, news, and scientific papers.
- Using a database that links patents and trademarks makes it possible to identify both the underlying technologies of emerging areas and competitors’ technologies.
- The paper states that distance measures and outlier detection methods should be changed depending on the data.
Limitations stated by the paper
The paper itself states five limitations.
- It only finds areas with the potential to emerge, and cannot give a definitive answer on whether those areas actually emerged. Its value increases when it is used together with lagging indicators that show actual emergence.
- The opportunities found must be made more concrete and refined with market, technology, and customer data.
- Computational complexity grows as the number of data records or embedding dimensions increases. There is also room for improvement by introducing state-of-the-art techniques.
- Tasks such as selecting business items and interpreting areas are still left to expert judgment and are not fully automated.
- There is only one case study, and the opportunities found have not yet been proven, so performance and usefulness could not be sufficiently confirmed.
Conditions for application in the author’s view
The following is not stated in the paper. It consists of conditions the author adds with practical application in mind.
- The range of products the company can actually enter must match the business landscape. If the scope is broad, most of the results will be areas unrelated to the company.
- Run the analysis with several different numbers of business areas and neighbors, and check how well the list of emerging areas holds up. If the list changes greatly, it is hard to base a decision on the result.
- In industries that file few trademarks, the number of trademarks per area is small, so visibility values may be unstable.
- Decide in advance who will verify the results identified as emerging areas with market or customer data, and by what procedure.
What to do right away
- Choose one NCL class that corresponds to your company’s industry. Collect the recent trademark G&S statements in that class, remove the procedural expressions in parentheses, and split the statements at semicolons. Then remove the items with a document frequency of 1 and the items in the Trademark Identification Manual, and you can check how many differentiated product expressions remain.
- Convert the remaining items into vectors with a pretrained language model and group them with k-means. Compute LOF on the cluster centroids and look first at the areas with values distinctly greater than 1.
- For each novel area, compute recency and visibility from the filing dates and frequencies of its trademarks. Divide the areas into four regions using the mean of the two values as the boundary, and review the items in the emerging region together with your company’s planners.