Looking at the attention mechanism alone, the cumulative count of related papers grew from 458 in the first half of 2019 to 18,870 in the first half of 2024. The paper points out that AI has a high publication volume and that new concepts and terms keep appearing. It therefore holds that expert-centered literature reviews can hardly capture research trends in a timely and exhaustive way.
This raises a question. Can language models automatically group AI papers to find topics, and can they even identify with numbers which topics are emerging?
Let us read one paper that answers this question.
- Title: Harnessing language models for computational literature review of emerging AI topics
- Journal: Information Processing & Management
- Year: 2025
- DOI: 10.1016/j.ipm.2025.104245
The paper proposes a Computational Literature Review (CLR) method. CLR is a review approach that uses computers to systematically collect, analyze, and synthesize literature. The proposed method collects papers from PapersWithCode, an AI paper database, and then does the following three things.
- It groups papers and finds topics using a language model further trained on AI papers at each time point.
- It computes how much each topic is emerging using four indicators.
- It shows, on a map, which topics at the preceding and following time points each topic connects to.
The terms used throughout this post are as follows.
- Pre-trained Language Model (PLM): A model trained in advance on large amounts of text that converts the meaning of a sentence into a numeric vector. The paper further trains SciBERT, a model for scientific papers, on AI papers and calls the result AI-SciBERT.
- Emergingness: The degree to which a topic is emerging at a given time point. The paper computes it with four indicators: novelty, growth, coherence, and uncertainty.
- Topic evolution map: A map that shows a topic together with its predecessor topics from the earlier time point and its successor topics from the next time point.
Motivation: AI Research Trends Are Hard to Follow by Hand Alone
Why Reading AI Research Trends Matters
The paper holds that the ability to read AI research trends is needed not only by researchers but also by industry practitioners and policymakers. The reasons the paper gives are as follows.
- AI has great practical potential. Success stories are increasing in industry, the physical sciences, public administration, education, healthcare, and management.
- The ability to quickly grasp and understand the latest AI research is regarded as a source of competitive advantage and a basis for evidence-based strategic planning.
- Analysis is difficult. The publication volume is large, the field spans multiple disciplines, and new concepts and terms keep appearing.
The paper notes that these difficulties often exceed what conventional expert-centered literature reviews can handle. It therefore holds that CLR, which systematically processes large volumes of text and finds hidden patterns, is becoming an essential tool.
What Prior Studies Did Not Address
Prior studies collected AI literature from academic databases such as Web of Science and Scopus. They then found hidden topics with topic modeling methods such as Latent Dirichlet Allocation (LDA), non-negative matrix factorization, Top2Vec, and BERTopic. For example, Arsenyan and Piepenbrink (2024) analyzed 7,708 studies on AI applications in operations research, management, and economics using LDA. Guler et al. (2024) found topics in 6,974 AI-related studies with a structural topic model and named them with the help of ChatGPT. Agbozo et al. (2024) analyzed 383 studies applying AI to sports analytics, centered on basketball, using BERTopic.
The paper identifies three gaps.
- Data source: Most prior studies focused on application areas where AI is newly used. They paid less attention to which algorithms or models are used. When one tries to analyze a specific technique such as attention or activation functions, existing search tools are inefficient. AI results appear first in conference papers rather than journals, yet existing tools mainly index journal papers.
- Methodology: Probabilistic topic models rely on word co-occurrence patterns, so they do not sufficiently capture the meaning and context of individual words. Some studies used PLMs such as BERT, but using them as is, without adapting them to AI-domain terminology, reduces effectiveness. A fixed model trained once before analysis can hardly reflect rapidly changing AI knowledge.
- Forward-looking analysis: Most prior studies stopped at analyses that look back on past changes, and experts interpreted the results to anticipate the future. A scientific method is needed to quantitatively identify the topics that will lead the field.
Two research questions of the paper arise from these gaps. First, which AI topics are emerging, and what are their characteristics? Second, given one target topic, what evolutionary path did it follow?
Method: Find Topics with Time-Specific Language Models, Score Emergingness, and Link the Paths
The paper describes the method in three stages. First, it builds the AI landscape using an AI-specific PLM and unsupervised clustering. Second, it evaluates the emergingness of topics with quantitative indicators. Third, it creates an interactive topic evolution map. In this post, the language model training and clustering contained in the paper’s first stage are separated and explained as four steps. Steps 1 and 2 of this post are the paper’s first stage, and Steps 3 and 4 are the paper’s second and third stages.
Case and Data
The data source is PapersWithCode. According to the paper, PapersWithCode started in 2018, its technical operation is handled by Meta AI, and its content is built jointly by researchers and developers worldwide. This describes the situation at the time the paper was written, so how the service is provided should be checked again before using it now.
The paper chose this data source for three reasons.
- It links papers to source code and public conference information, which provides reliability, transparency, and reproducibility.
- Community-maintained tags (datasets, methods, tasks) allow precise search and comparison.
- It can be obtained as a bulk dataset or via an API, which suits continuous updating.
Each paper carries method (Methods), dataset, and task tags. For example, the entry for “Attention Is All You Need” has the conference information ‘NeurIPS 2017’, the method tags ‘Linear Layer’ and ‘Absolute Position Encodings’, and the dataset tags ‘Penn Treebank’ and ‘WMT 2014’, among others.
The data scale is as follows.
- After excluding papers without a title or abstract and papers not in English, 506,286 AI papers were obtained.
- Of these, papers from the second half of 2017 to the first half of 2024, 409,021 in total, were organized into the dataset.
- This dataset was divided into 11 subsets accumulated in half-year units.
- The case analysis targets 18,870 papers that have ‘Scaled Dot-Product Attention’ among their method tags.
Scaled dot-product attention was chosen as the case because it plays a fundamental role in modern AI models, is developing rapidly, and has broad influence across many fields.
Step 1: Build Time-Specific AI-Specialized Language Models
- Input: Titles and abstracts of AI papers accumulated in half-year units
- Processing: Continued pre-training of SciBERT on AI papers
- Output: An AI-SciBERT model saved for each time point
The processing sequence is as follows.
- For each paper, concatenate the title and abstract into a single input sequence.
- Starting from the original SciBERT model (allenai/scibert_scivocab_uncased), continue pre-training on AI papers. This process is called domain adaptation.
- Using Masked Language Modeling (MLM), mask 15% of tokens and have the model predict them.
- Next sentence prediction is also trained together. Adjacent sentences are used as positive pairs, and in 50% of pairs the sentence is replaced with one from a different document to form negative pairs.
- The model trained on the papers accumulated up to each half year is saved for each time point.
Separate models are kept for each half year in order to fit the model to the expressions and language trends that change over time. The judgment is that, because AI knowledge changes quickly, a model that keeps learning over time is better than a fixed model.
The training settings are as follows.
- Learning rate 2e-5, 3 epochs, batch size of 128 per device
- Weight decay 0.01, 500 warmup steps, with the learning rate decreasing linearly
- Checkpoints are saved every 10,000 steps, keeping at most 2. Logs are recorded every 100 steps.
Step 2: Convert Documents to Vectors and Find Topics with K-Means
- Input: Text of the attention papers and the AI-SciBERT of the corresponding time point
- Processing: Document embedding, principal component analysis, K-means clustering
- Output: A list of topics for each half year
The processing sequence is as follows.
- Convert the paper text into dense vectors with the AI-SciBERT of the corresponding time point. The [CLS] token vector of the last layer is used as the document embedding.
- For computational efficiency, reduce dimensions with principal component analysis down to the number of dimensions that explains 80% of the total information.
- Group similar papers with K-means clustering. The number of clusters is tested in the range of 10% to 20% of the number of method tags.
- The elbow point, where the curvature of the inertia value is greatest, is chosen as the optimal number of clusters.
- The resulting clusters are defined as the AI topics of that time point.
The paper chose K-means for five reasons.
- It operates on Euclidean distance, so it suits fixed-length dense embeddings. Unlike LDA, it does not assume a generative distribution.
- It is fast to compute and easy to scale to large data. It is lighter than BERTopic and Top2Vec, which go through dimensionality reduction and density-based hierarchical clustering.
- The analyst can set the number of clusters k directly, adjusting the granularity of topics to the purpose.
- Documents and keywords near the centroid can be extracted, making the core meaning of a topic easy to read.
- Embeddings that have undergone domain adaptation have more consistent meaning, so topics are separated better.
In the case, the number of topics per half year grew from 29 to 111. Over the same period, cumulative documents grew from 458 in the first half of 2019 to 18,870 in the first half of 2024, about 41.2 times (18,870÷458≈41.2).
Step 3: Evaluate Emergingness with Four Indicators
- Input: Document embeddings and method tags per topic, and the number of new documents per time point
- Processing: Compute novelty, growth, coherence, and uncertainty, then classify topic types
- Output: Indicator values and emergence type for each topic
The four indicators were defined based on the technology forecasting literature. The premise and computation of each indicator are as follows.
- Novelty: The premise is that an emerging topic will contain many new documents signaling a new paradigm. Local Outlier Factor (LOF) is applied to the document embeddings, and the average is taken for each topic.
- Growth: The premise is that an emerging topic will grow quickly as attention from academia and industry increases. It is the logarithm of the ratio of the number of new documents at the current time point to that at the previous time point.
- Coherence: The premise is that an emerging topic will have a certain consistency and independence in content. It is defined as the average distance between the constituent documents and the topic centroid in the embedding space.
- Uncertainty: The premise is that an emerging topic will lean toward innovation with high potential for change rather than toward settled concepts. Shannon entropy is applied to the distribution of method tags within the topic.
LOF is an unsupervised outlier detection method that compares the density around a point with the density of its neighboring points. A document with a high LOF value sits in a low-density position while its surroundings are enclosed by dense clusters. It is therefore likely to be an outlier, that is, a new document. The number of neighbors k was set to the number of documents in each half year multiplied by 1% and rounded down. Because k is varied in proportion to the number of documents, the detection criterion remains balanced across time points even as documents increase.
The growth indicator is 0 when the numbers of new documents at the two time points are equal, positive when it increases, and negative when it decreases. Uncertainty is higher the more varied the methods used within a topic, and lower the more standardized the methods become. For example, among the emerging topics of the first half of 2024, the uncertainty of ETA-LM is 3.7542 and that of GBL-TM is 3.1622. This means the method tags used in ETA-LM documents are spread more widely.
The paper did not use impact indicators such as citations, because such indicators are usually lagging indicators that appear late.
After computing the indicators, the topics of all time points are clustered again using the four indicators to separate types. K-means was run over a range of 2 to 10 clusters, and the optimal number of clusters was chosen by the elbow point. The interpretation criteria for each type are as follows.
- Emerging: Novelty, growth, coherence, and uncertainty are all high.
- Post-emerging: Novelty is high but growth is slow, coherence is strong, and uncertainty is low.
- Pre-emerging: A cluster with low novelty but high growth or uncertainty may indicate pre-emerging potential.
- Marginal: Not prominent on any indicator.
In the case, 317 marginal, 350 pre-emerging, 32 post-emerging, and 93 emerging topics were found. The total is 317+350+32+93=792.
Step 4: Build the Topic Evolution Map
- Input: Topics per time point and the number of papers shared between topics
- Processing: Link topics at adjacent time points by shared papers and interpret type transitions
- Output: An interactive map showing the predecessor and successor paths of a target topic
The processing sequence is as follows.
- The topics at the immediately preceding time point that share papers with the target topic at time T are set as predecessor topics.
- The topics at time T+1 that share papers with the target topic are set as successor topics.
- To keep the visualization concise, only links with two or more shared papers are shown.
- The emergence types of the connected topics are read together to interpret where the topic stands in the technology life cycle.
Because each topic has a type, the direction of a type change carries meaning. The interpretations the paper presents are as follows.
- Pre-emerging → emerging: The technology may have moved from the emergence stage to the growth stage.
- Emerging → post-emerging: The technology may have passed its growth peak and entered the maturity stage.
- Marginal → pre-emerging or emerging: This may reflect the early potential of a marginal area that had received less attention.
- Pre-emerging → post-emerging: The technology may have passed quickly through the hype cycle.
- Emerging → marginal: This may be a signal that the technology is becoming obsolete.
Results: Graph Transformers, Specific-Language Modeling, and Efficient Transformers Emerged as Emerging Topics
Indicator Averages for the Four Types
To answer the first research question, the paper divided the topics of all time points into four types and computed the indicator averages for each type. The notable values in Table 3 of the paper are as follows.
- The post-emerging type has an average novelty of 1.1346 and an average coherence of 0.8260, both the largest among the four types.
- The average growth is 1.3055 for the emerging type and 0.3927 for the pre-emerging type.
- The average uncertainty of the post-emerging type is 3.1928, the lowest among the four types.
The paper describes the post-emerging type as technology that has solidified into one unified paradigm, citing high coherence as the basis. Connecting the fact that its average uncertainty is the lowest to this description is the author’s interpretation, not the paper’s.
Three Representative Topics of the Emerging Type
The paper holds that the emerging type may have the greatest influence on the AI landscape of the near future, and it summarized the core papers of the emerging topics of the first half of 2024.
- 24H1_25, Graph-Based Learning with Transformers (GBL-TM): Novelty 1.0345, growth 1.0678, coherence 0.7471, uncertainty 3.1622. It covers graph transformers used for tasks such as information networks, graph mining, and molecular representation. It includes a study that used a line graph transformer for molecular property prediction and a study that modeled user–item interactions in recommender systems with a heterogeneous subgraph transformer. Recently its scope has widened to knowledge graphs, time series, and multimodal data.
- 24H1_2, Language Modeling for Low-Resource and Specific Languages (LM-LRSL): Novelty 1.0647, growth 1.3863, coherence 0.7523, uncertainty 3.3756. It applies PLMs to natural language processing tasks in specific languages. It includes studies on Russian text detoxification, validation of a BERT model for Italian poetry, a Minangkabau corpus for sentiment analysis and translation, and an evaluation of multilingual BERT for Estonian.
- 24H1_15, Efficient Transformers for Language Modeling (ETA-LM): Novelty 1.0732, growth 0.9445, coherence 0.6608, uncertainty 3.7542. It revisits optimizers, batch training, and attention computation. LAMB, a large-batch optimization technique, cut BERT training time from three days to 76 minutes. Primer made training up to 4 times faster. Collaborative multi-head attention reduced the key and query projection size to one quarter while maintaining accuracy and speed.
There are also examples of other types. The marginal-type 24H1_7 (feature representation learning for image and video processing) was the largest topic in the first half of 2024. The pre-emerging 24H1_104 applies transformers to robotic systems. The post-emerging 24H1_47 deals with cross-lingual understanding of multilingual PLMs. New tasks keep attention, but momentum slowed as methods became standardized.
Evolution Path: The Knowledge-Intensive NLP Topic
To answer the second research question, the paper set as its target the emerging topic of the first half of 2020, 20H1_21 (‘Transformers for Knowledge-Intensive NLP’). This topic includes studies that drove the development of large language models, such as GPT-3. It also includes knowledge-use and retrieval-based techniques such as Retrieval-Augmented Generation (RAG), ColBERT, and DeBERTa.
There are three predecessor topics.
- 19H1_5 (pre-emerging): Includes papers that introduced PLMs with the transformer architecture, such as GPT-1 and BERT, and Transformer-XL.
- 19H1_12 (emerging): Includes transformer applications such as Universal Transformer, Music Transformer, and HellaSwag, a commonsense reasoning dataset.
- 19H2_28 (pre-emerging): Derived from the two topics above, it covers pre-training research and transformer applications together.
In the second half of 2020, the target topic split into five successor topics.
- 20H2_6 (pre-emerging): Deals with the robustness and transferability of PLMs, and led to 21H1_46 (pre-emerging) in the next half year.
- 20H2_17 (marginal): Uses PLMs for tasks such as entity matching and relation extraction, and split into three topics in the next half year.
- 20H2_26 (pre-emerging): Deals with natural language processing applications and practical and ethical considerations. The domain-adapted PLM topic of the next half year came mainly from this topic.
- 20H3_36 (marginal): Deals with efficient training techniques. In the next half year it split into topics including the efficient attention topic 21H1_58 (pre-emerging).
- 20H2_39 (pre-emerging): Deals with data-specific transformer variants. In the next half year it split into the attention computation efficiency topic 21H1_8 and the domain application topic 21H1_27.
The paper showed the predecessor and successor paths of the target topic with this map. On that basis, it interpreted the development of technologies such as GPT-3, RAG, ColBERT, and DeBERTa. The flow runs from transformer pre-training to domain-specific language modeling and attention efficiency.
Comparison with Existing Topic Models
The paper compared the proposed method with the following three methods.
- LDA
- BERTopic
- A method that attaches K-means to a pre-trained BERT
The evaluation indicators fall into three aspects.
- Topic coherence: C_v, UMass, normalized pointwise mutual information
- Cluster quality: Silhouette score, Calinski-Harabasz index, Davies-Bouldin index
- Computational efficiency: Time per epoch
No single model was ahead on all indicators. The paper stated that the proposed method ranks second in all three aspects: topic coherence, cluster quality, and computational efficiency. The paper’s interpretation is that, although it is not first on any one criterion, it delivers balanced performance across all criteria. However, this ranking is the paper’s claim, and the ranking may differ by indicator when looking at the individual indicators in Table 6 of the paper. For example, on the topic coherence indicator C_v, BERTopic is higher at 0.5560 than the proposed method at 0.4894. The computation time of BERTopic was also shorter than that of the proposed method.
Compared with general BERT using the same K-means, the proposed method has higher C_v and Calinski-Harabasz index. The paper explains that further training through domain adaptation raises the semantic consistency of embeddings, so topics are separated better. The differences in Table 6 are consistent with this explanation. However, the paper did not separately verify that this is the cause of the differences in Table 6.
Significance and Limitations
What Changes
- It showed the possibility of using the text, tags, and metadata of PapersWithCode as a CLR data source instead of journal-centered databases.
- It found topics with an AI-specialized PLM retrained every half year, reflecting changing AI terminology.
- In addition to analysis that looks back on the past, it presented a reproducible framework that computes emergingness with four indicators: novelty, growth, coherence, and uncertainty.
- It presented a method that links topics through shared papers and reads transitions of emergence type together to interpret the technology life cycle.
Usefulness for Practitioners
The paper holds that, because PLMs process specialized literature in bulk, AI research trends can be surveyed broadly with less time and cost. Results on emerging topics can be used in the decision-making of managers and policymakers who determine AI adoption and use. The source code is available at https://github.com/jm-chung/AI-CLR.
The paper also left guidance for practical application. When using a new data source, it recommends evaluating reliability, timeliness, domain relevance, and accessibility. In Table 5, which summarizes the evaluation by 13 AI experts (5 doctorates, 8 graduate students), PapersWithCode received medium reliability, Moderately High timeliness, and high domain relevance and accessibility. However, the body of the paper states that the experts picked PapersWithCode as the most comprehensive and superior source, with high reliability and domain relevance. Because the value for reliability in Table 5 differs from the statement in the body, it is better to check both places when citing this evaluation. If a data source has no method tags, method names can be extracted by named entity recognition to compute uncertainty in a similar way.
Usefulness for Researchers
The paper provides a systematic procedure for examining the rapidly changing AI landscape. The case was attention, but the paper states that by changing only the selection of core papers, it can also be applied to computer vision, speech processing, robotics, reinforcement learning, and generative AI. Because the four emergingness indicators were taken from the technology forecasting literature, there is also room to compare the indicators with studies of emerging technologies in other fields.
Limitations Stated by the Paper
The paper states five limitations itself.
- Language modeling needs further improvement. Models outside the autoencoding approach, large language models such as GPT-4 and Claude, knowledge graphs, and RAG can be considered.
- Ex post validation of the emergingness indicators is needed. The influence of topics should be examined separately over time through news, social media, citation networks, and so on.
- The topic map only links topics built from document similarity through shared documents. Entity-level topics and citation relationships between papers should be reflected.
- Only one attention case was analyzed. To secure external validity, further tests on other AI subtopics are needed.
- PapersWithCode data may have biases, such as a composition centered on conference papers and the exclusion of non-English literature.
Conditions for Application, from the Author
The following is not stated in the paper. It is conditions the author adds with practical application in mind.
- Each document needs method tags or comparable structured fields for the uncertainty indicator to be computed as in the paper. Without tags, the quality of the extraction work determines the quality of the indicator.
- Coherence is defined as the average distance from the centroid. With your own data, it is safer to decide beforehand how to read large and small values.
- The four indicators and the type classification are relative comparisons within the same data. Directly comparing indicator values from other datasets or other topics may lead to misinterpretation.
- Retraining the language model every half year requires GPU resources and time. The monitoring cycle and budget should be set first.
- How the data source is provided may change. Record the reference date each time you collect data so that the results can be reproduced.
What to Try Right Away
- Choose one method tag or keyword in your field of interest and collect the titles and abstracts of the corresponding papers. Then divide them into sets accumulated in half-year units, as the paper does, and first check how the number of documents per half year has grown.
- Before training a language model, build half-year clusters with embeddings you already have and K-means. For each cluster, compute growth (the logarithm of the ratio of new documents) and uncertainty (the Shannon entropy of the method tag distribution). Even these two indicators give a sense of which clusters are growing and where methods are not yet settled.
- Link the clusters of two adjacent half years by the number of shared papers, and keep only links with two or more shared papers. If you pick one cluster of interest and read the representative papers of its predecessor and successor clusters, you can outline where the topic came from and where it split.
Business Intelligence Lab will continue to introduce methods for reading technology trends with literature and patent data. The next post will likewise unpack one paper in the same way.