R&D managers want to know whether a research organization only refines technologies it already knows or draws knowledge from unfamiliar ones. Until now, this judgment has mostly been made on large groups of patents, such as those of a firm or a region. As a result, it has been hard to put a number on how far from existing knowledge a single patent explored.
This article addresses one question. Can we measure, as a figure for a single patent, how far from existing knowledge an individual patent explored?
We read one paper that answers this question.
- Title: Measuring knowledge exploration distance at the patent level: Application of network embedding and citation analysis
- Journal: Journal of Informetrics
- Year: 2022
- DOI: 10.1016/j.joi.2022.101286
The paper proposes a method for calculating the Knowledge Exploration Distance (KED) of a single patent. The method consists of four tasks.
- Build a network from the co-occurrence of patent classification codes.
- Convert the codes into vectors with network embedding (a technique that turns each node of a network into a numeric vector that preserves structural information).
- Represent patents with these vectors.
- Calculate how far apart the citing patent and the cited patents are in the vector space.
The terms used throughout this article are as follows. The knowledge base is the set of existing knowledge elements on which a new invention builds. The paper treats the prior patents cited by a patent as that patent’s knowledge base. A technical element refers to each individual International Patent Classification (IPC) code attached to a patent. Knowledge exploration distance indicates how novel the relationship is between the technical elements of the citing patent and those of the cited patents.
Motivation: Knowledge Exploration Has Been Measured Only at the Organizational Level
Why Knowledge Exploration Distance Matters
The paper raises its problem on the basis of prior research linking knowledge exploration to innovation performance. The reasons the paper gives are as follows.
- Knowledge exploration is the activity of drawing in external knowledge from other technologies or regions to raise the quality of new knowledge. Knowledge created this way is hard to imitate or substitute, which makes it a valuable asset.
- Exploring knowledge from many domains widens the scope of information search. Firms obtain more new information and can use it to identify and solve existing problems.
- Prior research reported that knowledge exploration distance has a significant relationship with an organization’s innovation activity, the quality of innovation, and regional innovation actors.
- In one firm-level study, technological diversity of exploration showed a linear positive relationship with innovation quality. Geographic diversity showed a curvilinear relationship.
What Existing Research Has Not Covered
Existing research defined the knowledge base by the boundaries of a firm, organization, or region. It measured the degree of exploration by comparing patents inside those boundaries with patents outside them. For example, one study defined exploration as the share of citations to technology domains unfamiliar to team members. Another study calculated the technological distance between firms from the number of patents each firm held in each classification.
The gaps the paper points out are as follows.
- Existing methods can be applied only to patents held by a particular entity or filed in a particular region.
- For patents outside the given boundary, the knowledge base is hard to measure. Those patents also explored existing knowledge during the filing process.
- Citation-based novelty indicators already existed. However, text similarity only looks at how different the technical words or expressions are. The indicator of Verhoeven et al. (2016) only counts the number of new IPC pairs.
- Even if one can tell whether new IPC pairs exist, there is no information on the distance between technical elements, so the degree of association is not reflected.
The paper argues that defining the knowledge base through citation relationships frees the scope from owners or regions.
Method: Convert IPC Codes into Vectors and Average the Distances to Cited Patents
The paper describes the method in three stages.
- Data collection and preprocessing
- A stage that converts patents into vectors using classification information
- A stage that calculates knowledge exploration distance using citation relationships
In this article, the three tasks of the second stage are separated out and the method is explained in five steps. Those three tasks are network construction, network embedding, and patent representation. Step 1 of this article corresponds to the paper’s first stage, Steps 2 to 4 to the paper’s second stage, and Step 5 to the paper’s third stage.
Case and Data
The data are patents registered at the Korea Intellectual Property Office (KIPO). The researchers gave as their reason that KIPO provides high-quality citation, classification, and bibliographic information. The case technology is artificial intelligence (AI). Three sets of patents were used in the analysis.
- All patents: 2,438,139 patents registered at KIPO from 1984 to 2019. They are used to build the technology ecology network.
- Base patents: 28,569 registered patents related to AI technology. They are the targets for which knowledge exploration distance is calculated.
- Prior patents: 36,090 patents cited by the base patents. This count excludes duplicates, and these patents form the knowledge base of the base patents.
The base patents were searched using the AI-related classification codes defined by the World Intellectual Property Organization (WIPO). A base patent cited 1.72 patents on average. Because the oldest prior patent was registered in 1984, the network construction period was also set to start in 1984.
Step 1: Data Collection and Defining the Knowledge Base
- Input: A patent database containing citations, classification codes, and bibliographic information
- Process: For each patent to be analyzed, collect the prior patents it cites
- Output: A set of prior patents [c0, c1, c2, …] for each patent
- Choose the patent a to be analyzed.
- Collect all patents cited by a to form a set of prior patents.
- Take this set as the knowledge base of a.
The knowledge base can also be defined by the applicant’s existing patents or by the scientific literature cited. The researchers chose citation information. This is because, during examination, prior patents with similar technology or that serve as the technological basis are likely to be cited. Citation relationships also have a clear direction, so they provide the directional information needed to compare technological knowledge.
In the case study, this step yields 28,569 base patents and 36,090 prior patents.
Step 2: Building the Technology Ecology Network
- Input: The IPC codes of all 2,438,139 patents
- Process: Link pairs of IPC codes attached together to a patent and record their frequency
- Output: A network in which nodes are IPC codes and links are co-occurrence frequencies
- Set the IPC codes as nodes.
- Link two codes that appear together within a patent.
- Set the number of times they appear together as the weight of the link.
Co-occurrence was used because it represents the natural association between technical elements. Citation relationships represent dependence arising from technology development. A co-occurrence network reveals the hidden position of a technical element through indirect links as well as direct links.
The code level must also be decided. The IPC is a hierarchy running from section to subgroup. For example, G06F 30/27 passes through section G (Physics), class G06 (Computing), subclass G06F, and main group G06F 30 (Computer-aided design) to reach a subgroup that uses machine learning. Many studies analyze at the subclass or class level because of computational cost. This paper used IPC codes in full, down to the last digit.
In the case study, the network consisted of 65,556 nodes and 2,314,188 links.
Step 3: Network Embedding
- Input: The technology ecology network from Step 2
- Process: Learn each IPC code with an embedding technique called LINE
- Output: A 128-dimensional vector for each IPC code
- Check the nature of the network. This network does not change over time, has a single node type, and has a very large number of nodes and links.
- Given this nature, choose the LINE technique proposed by Tang et al. in 2015.
- Learn first-order and second-order proximity together to obtain a vector for each code.
When there are too many nodes and links, computation becomes difficult. So the information of each code is reduced to a fixed-length vector. The goal of embedding is to place nodes that are similar in the network close together in the vector space as well.
The two proximities mean the following.
- First-order proximity: The degree to which two codes are directly linked. The vectors are trained to become more similar the more frequently two codes appear together.
- Second-order proximity: The degree to which two codes share common neighbors. Even if they do not appear together directly, they become close if they are often grouped with the same codes.
The settings are as follows. The proximity order is 2, the negative sampling ratio is 5, the initial learning rate is 0.025, and the embedding dimension is 128. In the case study, every IPC code in the network obtained a unique 128-dimensional vector.
Step 4: Creating Patent Vectors
- Input: The IPC codes attached to each patent and the code vectors from Step 3
- Process: Multiply the code vectors by weights and average them
- Output: A 128-dimensional vector for each patent
- Sort the classification codes of the patent by frequency.
- Multiply each code vector by a weight.
- Average the weighted vectors to obtain the patent vector.
Because this paper used codes down to the last digit, no codes overlapped within a patent. Therefore each code had the same weight. In effect, this is a simple average. For example, if a patent has three IPC codes, the values at the same position in the three code vectors are added and divided by 3.
The researchers give two reasons for choosing classification codes over text.
- Text is tied to the language of the patent office where the patent was registered.
- It is hard to reflect all the specialized terms that differ across technology fields.
Classification codes such as the IPC are internationally common and contain enough information about technical elements. In the case study, each of the patents in the full set was represented as a 128-dimensional vector.
Step 5: Calculating Knowledge Exploration Distance
- Input: The vector of patent a and the vectors of the prior patents cited by a
- Process: Compute the cosine distance for each prior patent and average
- Output: The knowledge exploration distance (KED) of patent a
- Compute the cosine similarity between a and one prior patent c. Multiply the values at the same positions of the two vectors, add them up, and divide by the product of the lengths of the two vectors.
- Subtract the cosine similarity from 1 to convert it to a distance.
- Compute this distance for all prior patents and average by dividing by the number of prior patents |C|.
Written as a formula, KED(a) = Σ(1 − cosine similarity(a, c)) ÷ |C|. Cosine similarity becomes larger the more the directions of the two vectors coincide. Therefore, a large KED means that the technical elements of the cited patents are, on average, far from those of a in the technology ecology network. The researchers interpret a large KED as meaning that exploration of existing knowledge was active during the invention process. This interpretation is the researchers’ own, and it is not something separately demonstrated in the results below.
In the case study, this step produced the KED of all 28,569 base patents.
Results: Patents with Large Distances Largely Agree with Existing Novelty Indicators and Show Small but Significant Relationships with Patent Indicators
A Calculation Example for One Patent
The researchers showed the calculation process with a patent on a video summary playback system and a mobile communication terminal equipped with it (KR1020050090452). This patent has three technical elements.
- G11B 27/10: Indexing, addressing, timing
- G11B 20/10: Digital recording and reproducing
- H04B 1/40: Circuits
This patent cited four prior patents. All five patents were turned into vectors using the IPC embedding values, and the distance was calculated for each prior patent. Among the four cited patents, the distance to KR1020020074328 was the largest at 1.1764. The distances of the other three were 0.3168, 0.3332, and 0.3205. The values for each patent can be found in Table 4 of the paper.
When the author averaged the four distances directly, the result was (1.1764+0.3168+0.3332+0.3205)÷4=0.5367. This value is the same as the KED of this patent listed in Table 4 of the paper. The distance of the first prior patent exceeds 1. Since 1−1.1764=−0.1764, this means the cosine similarity of the two vectors is negative. The technical elements of this prior patent were H04N 21/4402, H04N 5/915, and G11B 27/28. These codes had rarely been used together with the technical elements of the citing patent. The distances of the other three were much smaller than that of the first, so a single distant exploration raised the average.
Distribution of the 28,569 Base Patents
The mean KED of the base patents was about 0.3726, and the standard deviation was 0.0762. About 47% of the base patents had values below the mean. The maximum was 1.2018. The minimum occurs when the technical elements of the citing patent and the cited patent are identical. The researchers explain that a few patents with high KED raise the mean.
Comparison with Existing Novelty Indicators
The researchers compared the top 10 patents by KED with two novelty indicators.
- Novelty in Recombination (NR): It is 1 if any IPC pair of the patent had never appeared before the application year, and 0 otherwise.
- Novelty in Technological knowledge Origins (NTO): It is 1 if any of the pairs formed by matching the IPC codes of the citing patent with those of the cited patent is a new pair, and 0 otherwise. In this study, it was measured using IPC codes down to the last digit.
All of the top 10 patents cited only one prior patent. All of the top 10 had an NTO of 1. For NR, 6 of the 10 were 1. The remaining 4 (ranks 1, 3, 5, and 7) had an NR of 0. The values by rank can be found in Table 5 of the paper.
There are two representative cases. The rank 1 patent (an automatic recovery device and method for trunk interfaces) had the largest KED at 1.2018. All six of its citation pairs were new combinations. The rank 2 patent had the pair A61B-005/053 and A61N-001/05. This pair had never appeared before 2010, the application year.
In sum, the top KED patents all agreed with NTO. For NR, 6 were 1, but 4, including the rank 1 patent with the largest KED, were 0. The researchers state that most of the top patents were rated high on both indicators. The researchers point out that the two indicators only tell whether a new pair exists. KED adds the distance between technical elements, and so also tells the degree to which the relationship is new.
Verification Through Content Analysis
The researchers also read the content of the top patents and their prior patents directly.
- The rank 2 patent (a device for measuring the impedance of stimulation electrodes for a neurostimulator) measures the impedance of each electrode from the voltage difference across the electrode. The prior patent was a device that electrically stimulates muscle tissue. This patent differed in that it analyzes the electrochemical characteristics of stimulation electrodes that touch nerves or muscles.
- The rank 1 patent shortened recovery time and increased connection reliability by providing a spare interface. The prior patent did not solve physical transmission errors in the trunk interface.
They also examined patents with low KED. These patents addressed problems similar or identical to those of the cited patents, and it was hard to find clear differences in technical elements or content.
Relationship with Other Patent Indicators
The researchers ran regression analyses with KED as the dependent variable. There were four models.
- Model 1: Included three variables derived from citation information.
- Model 2: Added three stakeholder variables.
- Model 3: Included all possible variables.
- Model 4: Removed the dependent claim scope, which was not significant, through stepwise selection.
The Variance Inflation Factors (VIF) were all below 5.00, so there was no multicollinearity problem. The variables showing a negative relationship with KED in Model 4 were as follows.
- Number of patent citations
- Number of inventors
- Number of applicants
- Commercial scope
The variables showing a positive relationship were as follows.
- Non-patent literature citations
- Technology cycle
- Number of agents
- Independent claim scope
Most of these variables were significant at the p<0.01 level in Model 4. Non-patent literature citations and number of applicants were significant only at the p<0.05 level. All coefficients were small. The coefficient for each variable can be found in Table 7 of the paper.
The interpretation is as follows.
- The more patent literature a patent cites, the lower its KED. The researchers consider that there are inherently few prior patents available for distant exploration, and that the average distance may fall as citations increase.
- Non-patent literature citations and technology cycle have a positive relationship with KED. Patents with large KED tend to be more connected to scientific literature.
- Technology cycle is the median age of the cited prior patents. The researchers note that the technology cycle becomes shorter as a technology moves from the growth stage to the maturity stage of its life cycle, and consider that patents with large KED may increase at the same stage.
- Patents with large KED tend to have fewer inventors and applicants and more agents.
- Independent claim scope has a positive relationship. Patents with large KED tend to have more independent claims. Citing prior research that patents with many claims have a broad scope of protection, the paper interpreted this to mean a broad scope of protection. The effect size is small. The coefficient for independent claim scope is 0.0019 in Model 3 and 0.0017 in Model 4.
- Commercial scope has a negative relationship. The researchers interpret this to mean that patents with large KED tend to be registered at only one patent office rather than in several countries.
The coefficient for examination period was close to 0. The coefficient of determination (R²) of Model 4 was 0.0124. This means these indicators explain only a very small part of the variation in KED. The relationships should be read as significant but small in magnitude.
Significance and Limitations
What Changes
- Knowledge exploration distance was calculated for a single patent rather than for an organization or region.
- By defining the knowledge base through citation relationships, the method can be applied to all patents without owner or regional boundaries.
- It uses the IPC down to the last digit and expresses the distance between codes numerically through network embedding.
- It goes beyond whether a new combination exists and obtains the degree to which a relationship is new.
- It confirmed significant relationships between KED and indicators related to scope of protection, prior knowledge, and patent value.
Usefulness for Practitioners
The researchers consider this indicator a new measure for reviewing a firm’s patent portfolio and knowledge exploration activity. Because it is based on classification codes, it can be used regardless of the patent office. The procedure is systematic and requires little expert intervention, so it can be recalculated in other analysis environments.
For R&D managers, this provides a means of confirming the degree of exploration numerically. For reference, exploitative innovation improves existing methods, and exploratory innovation seeks new opportunities. Collecting the KED of a firm’s own patents makes it possible to analyze the organization’s exploration tendencies. In addition, the author suggests treating patents with unusually high KED as candidates whose content should be read separately. The paper presents this as an implication that, in academic research, patents with large distances can be investigated in depth.
Usefulness for Researchers
- Patent-level KED can be aggregated to analyze the exploration tendencies of firms, organizations, and regions.
- KED can be included as a variable in analyses of forward citations, patent lifetime, and technology convergence.
- Together with text similarity indicators, it can serve as a novelty indicator from the perspective of technological association.
Limitations Stated in the Paper
The paper states four limitations itself.
- There is no absolute rule for defining the knowledge base of each patent. The scope of prior patents could be expanded using continuation or family information.
- The knowledge base was defined only through citation relationships. It could be broadened to the perspective of applicants or scientific literature.
- The effect of KED on patent value or quality needs post hoc analysis. Because forward citations accumulate over time, the aggregation method needs refinement.
- Analysis from the perspective of novelty and lifetime is needed, comparing rejected and registered patents, and lapsed patents and long-maintained patents.
The researchers name three future tasks: developing an embedding technique suited to the technology ecology network, a patent representation that uses bibliographic and text information together, and a content comparison with non-patent literature.
Conditions for Application as Seen by the Author
The following are not in the paper. They are conditions the author adds with practical application in mind.
- For patents with few citations, KED is largely determined by one or two cases. Since all of the top 10 had a single citation, it is safer to interpret KED with the number of citations displayed alongside.
- Because the explanatory power of the regression is low, do not judge patent value by KED alone. Use it together with content review.
- Embedding values vary with the range of patents learned. When comparing results from different periods or patent offices, use the same embedding.
- The case is AI patents at KIPO. For other technology fields, check the mean and distribution anew before setting a threshold.
What to Try Next
- Choose one technology field among your firm’s registered patents and organize, for each patent, its IPC codes and the list of prior patents it cites. This is the same preparation as Step 1 of the paper.
- Pick one prior patent and compare in a table how much its IPC codes overlap with those of the citing patent. Marking the citation pairs with little code overlap can help narrow down candidates to examine. This method is not the paper’s method but a rough preliminary screening proposed by the author. Because embedding also reflects indirect associations such as common neighbors, KED can be small even when few codes overlap.
- If you have calculated the embedding, pick a few top and bottom KED patents and read their content in comparison with their prior patents. This is the same way the paper verified its results.
Business Intelligence Lab will continue to introduce, on this blog, ways of reading technologies and organizations through patent data.