Javascript must be enabled to continue!
News event
View through CrossRef
When analyzing news media data with automated content analysis techniques, studies often aggregate their measures at the article level (Nicholls & Bright, 2019). However, many stories in the news unfold over at least multiple days, across multiple articles and media outlets. Hence, another possibility for the unit of analysis can be “story chains” (Nicholls & Bright, 2019) or “news events” (Trilling & Van Hoof, 2020). Hereafter, we adopt the term news event. Manual content analysis approaches often rely on predefined news events, however, in large-scale automated content analyses this is often not feasible. Alternatively, unsupervised approaches can be used to detect news events, which is the focus of this contribution. Generally, this involves two main steps. First, pairwise similarity between articles is measured. Second, these articles are represented in a network graph and grouped using clustering algorithms. This unsupervised approach offers the benefit of identifying emergent events across extensive corpora. However, this comes with challenges: setting appropriate timeframes and similarity thresholds is complex, and validating the automatically generated clusters is a time-consuming process.
Field of Application/Theoretical Foundation
A news event is a unit of analysis that can be used to study different phenomena related to news media. Trilling and Van Hoof (2020, p. 1321) describe it as a specific event which leads to news coverage in one or more outlets across one or more articles. Usually, a news event is more than one article but less than a topic or an issue. An example of a news event could be a specific airplane crash, while a topic could be airplane crashes. Hence, news events can be part of topics and/or issues.
Closely related to news events are news story chains (Nicholls & Bright, 2019). The key difference lies in the notion of chains which implies a strict temporal order that might be sometimes difficult to identify for a particular news event (Trilling & Van Hoof, 2020). However, the operationalization of news events and news story chains is largely the same in empirical studies relying on unsupervised approaches.
Using news events as the unit of analysis can be helpful when comparing media coverage across outlets (e.g., Kapellas & Kapidakis, 2022), when analyzing phenomena such as media hype (e.g., Vasterman, 2005) or media storms (e.g., Litterer et al., 2023; Markus et al., 2024) which are a subcategory of news events, or fragmentation in news recommender systems (e.g., Polimeno et al., 2023) to name but a few examples.
References/Combination with other Methods
Different fields in computer science develop approaches to identify concepts similar to news events for automated news media monitoring, recommender systems, and many more applications. For a recent survey of the literature, we refer to Keith-Norambuena et al. (2023).
Moreover, news events can be identified ex ante in studies focusing on specific cases (e.g., Langer & Gruber, 2021) or when databases of potential events exist (e.g., Welbers et al., 2022). This contribution focuses on the unsupervised identification of news events. However, Markus et al. (2026) emphasize that if contextual knowledge is required to classify articles (e.g., media storms), semi-supervised annotation strategies like Expert Initiated Latent Space Sampling (EILSS, Markus et al., 2023) might be more suitable. By integrating domain knowledge with document embeddings, EILSS uses iterative sampling to direct human annotators to the most pertinent articles, thereby optimizing the annotation process and allowing for focused validation.
Example studies
We compare the approaches to identify news events in large corpora of four different papers (i.e., Gedikli et al., 2021; Litterer et al., 2023; Nicholls & Bright, 2019; Trilling & Van Hoof, 2020). The studies analyze news content ranging from international and US online news (Gedikli et al., 2021; Litterer et al., 2023) to UK online news (Nicholls & Bright, 2019), and Dutch online news (Trilling & Van Hoof, 2020). Still, all the approaches share the following structure:
First, pairwise similarity between articles is measured. While some approaches rely on co-occurrences of keywords or named entities to calculate similarity (Gedikli et al., 2021; Nicholls & Bright, 2019), others rely on embedding models, that is, models that represent text as numerical vectors capturing semantic similarity (Litterer et al., 2023; Trilling & Van Hoof, 2020). However, not all article pairs are equally likely to report on a given event. As a result, studies choose a time frame of reference within which all article pairs are compared. For instance, Nicholls and Bright (2019) estimate similarity scores for all articles published within three days. Different similarity scores have been used (see below).
Second, the news articles are grouped based on the similarity scores with a similarity cutoff or a clustering algorithm. The resulting clusters of news articles are then the identified news events.
The papers presented use different approaches for both steps of the news event identification pipeline. Because the two stages are usually independent, these pipelines are modular and can be adjusted. Desai and Nagwanshi (2020) compare many different similarity scoring and clustering approaches. They find that different approaches can yield good results, given that pairwise similarity between articles works well and that the number of clusters is not prespecified (Nicholls & Bright, 2019).
Gedikli et al. (2021) further highlight that validation can be rather time-consuming and that it might be more practical to validate precision (share of detected events that are actual events) as opposed to recall (share of actual events that are detected). Moreover, more classic bag-of-words approaches such as TF-IDF weighting remain remarkably reliable as the underlying logic is conceptually close to how news events manifest in news articles. Because articles about the same events tend to use similar words.
Table 1: Comparison of news event identification pipelines.
Article
Pairwise similarity
Comparison frame/Article selection criteria
Clustering
Validation
Nicholls & Bright (2019)
Information Retrieval approaches: Mean similarity score between two scores.
(1) Keyword scoring by identifying the 100 most distinctive words in each article compared to the whole corpus. And then measuring co-occurrence between articles.
(2) BM25F scoring algorithm
3-day sliding window
Representing pairwise similarity metrics as similarity network and then network partitioning into clusters.
Algorithm:
Infomap method
Output: hierarchical clustering
Two validation steps:
1. Pre-classification validation of pairwise similarity scores; 100% agreement (Krippendorff’s α of 1.0).
2. Post-classification evaluation of 25 randomly sampled generated story chains; 93% agreement, lower Krippendorff’s α of 0.50 due to boundary disagreements.
Trilling & Van Hoof (2020)
(1) TF-IDF weighing and then cosine similarity scores between articles.
(2) Amsterdam Embedding Model (pre-trained embeddings) and then softcosine similarity scores between articles.
3-day sliding window accounted for weekends
Representing pairwise similarity metrics as similarity network and then network partitioning into clusters.
Algorithms:
Leiden algorithm
Robustness check: Infomap method
Randomly annotated 100 events identified by the model. The conservative cosine method achieved 89% precision for events and 94.39% precision for articles. The soft cosine method achieved a precision of 75% for events and 86.92% for articles.
Gedikli et al. (2021)
Co-occurrence of named entities (persons, locations, organizations) with a transformer-based Named Entity Recognition model, calculated via the Named Entities Shared Measure (NESM)
No time window.
Clustering based on similarity measure and similarity cutoff.
Validate pairwise similarity scores; 99.4% agreement, Krippendorff’s α of 0.87.
Litterer et al. (2023)
Fine-tuned transformer embedding model: bi-encoder MPNet. Document-level embeddings are created.
8-day sliding window and named entity in common.
Cosine similarity cutoff (> 0.9). The remaining article pairs are represented as a graph, and story clusters are extracted as the connected components of this graph.
Evaluation of the automated similarity model against human annotations from a SemEval competition dataset, achieving a mean Pearson correlation r of 0.86.
References
Desai, A., & Nagwanshi, P. (2020). Grouping news events using semantic representations of hierarchical elements of articles and named entities. 2020 3rd International Conference on Algorithms, Computing and Artificial Intelligence, 1–6. https://doi.org/10.1145/3446132.3446399
Gedikli, F., Stockem Novo, A., & Jannach, D. (2021). Semi-automated identification of news story chains: A new dataset and entity-based labeling method. Proceedings of the 9th International Workshop on News Recommendation and Analytics (INRA 2021), 29–42. https://ceur-ws.org/Vol-3143/paper3.pdf
Kapellas, N., & Kapidakis, S. (2022). A Text Similarity Study: Understanding How Differently Greek News Media Describe News Events: Proceedings of the 14th International Joint Conference on Knowledge Discovery, Knowledge Engineering and Knowledge Management, 245–252. https://doi.org/10.5220/0011589700003335
Langer, A. I., & Gruber, J. B. (2021). Political Agenda Setting in the Hybrid Media System: Why Legacy Media Still Matter a Great Deal. The International Journal of Press/Politics, 26(2), 313–340. https://doi.org/10.1177/1940161220925023
Litterer, B., Jurgens, D., & Card, D. (2023). When it Rains, it Pours: Modeling Media Storms and the News Ecosystem. Findings of the Association for Computational Linguistics: EMNLP 2023, 6346–6361. https://doi.org/10.18653/v1/2023.findings-emnlp.420
Markus, D. K., Levi, E., Sheafer, T., & Shenhav, S. R. (2024). Reap the Wild Wind: Detecting Media Storms in Large-Scale News Corpora. In Y. Al-Onaizan, M. Bansal, & Y.-N. Chen (Eds.), Findings of the Association for Computational Linguistics: EMNLP 2024 (pp. 4786–4797). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.findings-emnlp.275
Markus, D. K., Levi, E., Sheafer, T., & Shenhav, S. R. (2026). From Showers to Hurricanes: A Typology of Media Storms. Political Communication, 1–19. https://doi.org/10.1080/10584609.2026.2656633
Markus, D. K., Mor-Lan, G., Sheafer, T., & Shenhav, S. R. (2023). Leveraging Researcher Domain Expertise to Annotate Concepts Within Imbalanced Data. Communication Methods and Measures, 17(3), 250–271. https://doi.org/10.1080/19312458.2023.2182278
Nicholls, T., & Bright, J. (2019). Understanding News Story Chains using Information Retrieval and Network Clustering Techniques. Communication Methods and Measures, 13(1), 43–59. https://doi.org/10.1080/19312458.2018.1536972
Keith-Norambuena, B., Mitra, T., & North, C. (2023). A Survey on Event-Based News Narrative Extraction. ACM Computing Surveys, 55(14s), 300:1-300:39. https://doi.org/10.1145/3584741
Polimeno, A., Reuver, M., Vrijenhoek, S., & Fokkens, A. (2023). Improving and Evaluating the Detection of Fragmentation in News Recommendations with the Clustering of News Story Chains (arXiv:2309.06192). arXiv. https://doi.org/10.48550/arXiv.2309.06192
Trilling, D., & Van Hoof, M. (2020). Between Article and Topic: News Events as Level of Analysis and Their Computational Identification. Digital Journalism, 8(10), 1317–1337. https://doi.org/10.1080/21670811.2020.1839352
Vasterman, P. L. M. (2005). Media-Hype: Self-Reinforcing News Waves, Journalistic Standards and the Construction of Social Problems. European Journal of Communication, 20(4), 508–530. https://doi.org/10.1177/0267323105058254
Welbers, K., Van Atteveldt, W., Bajjalieh, J., Shalmon, D., Joshi, P. V., Althaus, S., Chan, C.-H., Wessler, H., & Jungblut, M. (2022). Linking event archives to news: A computational method for analyzing the gatekeeping process. Communication Methods and Measures, 16(1), 59–78. https://doi.org/10.1080/19312458.2021.1953455
Title: News event
Description:
When analyzing news media data with automated content analysis techniques, studies often aggregate their measures at the article level (Nicholls & Bright, 2019).
However, many stories in the news unfold over at least multiple days, across multiple articles and media outlets.
Hence, another possibility for the unit of analysis can be “story chains” (Nicholls & Bright, 2019) or “news events” (Trilling & Van Hoof, 2020).
Hereafter, we adopt the term news event.
Manual content analysis approaches often rely on predefined news events, however, in large-scale automated content analyses this is often not feasible.
Alternatively, unsupervised approaches can be used to detect news events, which is the focus of this contribution.
Generally, this involves two main steps.
First, pairwise similarity between articles is measured.
Second, these articles are represented in a network graph and grouped using clustering algorithms.
This unsupervised approach offers the benefit of identifying emergent events across extensive corpora.
However, this comes with challenges: setting appropriate timeframes and similarity thresholds is complex, and validating the automatically generated clusters is a time-consuming process.
Field of Application/Theoretical Foundation
A news event is a unit of analysis that can be used to study different phenomena related to news media.
Trilling and Van Hoof (2020, p.
1321) describe it as a specific event which leads to news coverage in one or more outlets across one or more articles.
Usually, a news event is more than one article but less than a topic or an issue.
An example of a news event could be a specific airplane crash, while a topic could be airplane crashes.
Hence, news events can be part of topics and/or issues.
Closely related to news events are news story chains (Nicholls & Bright, 2019).
The key difference lies in the notion of chains which implies a strict temporal order that might be sometimes difficult to identify for a particular news event (Trilling & Van Hoof, 2020).
However, the operationalization of news events and news story chains is largely the same in empirical studies relying on unsupervised approaches.
Using news events as the unit of analysis can be helpful when comparing media coverage across outlets (e.
g.
, Kapellas & Kapidakis, 2022), when analyzing phenomena such as media hype (e.
g.
, Vasterman, 2005) or media storms (e.
g.
, Litterer et al.
, 2023; Markus et al.
, 2024) which are a subcategory of news events, or fragmentation in news recommender systems (e.
g.
, Polimeno et al.
, 2023) to name but a few examples.
References/Combination with other Methods
Different fields in computer science develop approaches to identify concepts similar to news events for automated news media monitoring, recommender systems, and many more applications.
For a recent survey of the literature, we refer to Keith-Norambuena et al.
(2023).
Moreover, news events can be identified ex ante in studies focusing on specific cases (e.
g.
, Langer & Gruber, 2021) or when databases of potential events exist (e.
g.
, Welbers et al.
, 2022).
This contribution focuses on the unsupervised identification of news events.
However, Markus et al.
(2026) emphasize that if contextual knowledge is required to classify articles (e.
g.
, media storms), semi-supervised annotation strategies like Expert Initiated Latent Space Sampling (EILSS, Markus et al.
, 2023) might be more suitable.
By integrating domain knowledge with document embeddings, EILSS uses iterative sampling to direct human annotators to the most pertinent articles, thereby optimizing the annotation process and allowing for focused validation.
Example studies
We compare the approaches to identify news events in large corpora of four different papers (i.
e.
, Gedikli et al.
, 2021; Litterer et al.
, 2023; Nicholls & Bright, 2019; Trilling & Van Hoof, 2020).
The studies analyze news content ranging from international and US online news (Gedikli et al.
, 2021; Litterer et al.
, 2023) to UK online news (Nicholls & Bright, 2019), and Dutch online news (Trilling & Van Hoof, 2020).
Still, all the approaches share the following structure:
First, pairwise similarity between articles is measured.
While some approaches rely on co-occurrences of keywords or named entities to calculate similarity (Gedikli et al.
, 2021; Nicholls & Bright, 2019), others rely on embedding models, that is, models that represent text as numerical vectors capturing semantic similarity (Litterer et al.
, 2023; Trilling & Van Hoof, 2020).
However, not all article pairs are equally likely to report on a given event.
As a result, studies choose a time frame of reference within which all article pairs are compared.
For instance, Nicholls and Bright (2019) estimate similarity scores for all articles published within three days.
Different similarity scores have been used (see below).
Second, the news articles are grouped based on the similarity scores with a similarity cutoff or a clustering algorithm.
The resulting clusters of news articles are then the identified news events.
The papers presented use different approaches for both steps of the news event identification pipeline.
Because the two stages are usually independent, these pipelines are modular and can be adjusted.
Desai and Nagwanshi (2020) compare many different similarity scoring and clustering approaches.
They find that different approaches can yield good results, given that pairwise similarity between articles works well and that the number of clusters is not prespecified (Nicholls & Bright, 2019).
Gedikli et al.
(2021) further highlight that validation can be rather time-consuming and that it might be more practical to validate precision (share of detected events that are actual events) as opposed to recall (share of actual events that are detected).
Moreover, more classic bag-of-words approaches such as TF-IDF weighting remain remarkably reliable as the underlying logic is conceptually close to how news events manifest in news articles.
Because articles about the same events tend to use similar words.
Table 1: Comparison of news event identification pipelines.
Article
Pairwise similarity
Comparison frame/Article selection criteria
Clustering
Validation
Nicholls & Bright (2019)
Information Retrieval approaches: Mean similarity score between two scores.
(1) Keyword scoring by identifying the 100 most distinctive words in each article compared to the whole corpus.
And then measuring co-occurrence between articles.
(2) BM25F scoring algorithm
3-day sliding window
Representing pairwise similarity metrics as similarity network and then network partitioning into clusters.
Algorithm:
Infomap method
Output: hierarchical clustering
Two validation steps:
1.
Pre-classification validation of pairwise similarity scores; 100% agreement (Krippendorff’s α of 1.
0).
2.
Post-classification evaluation of 25 randomly sampled generated story chains; 93% agreement, lower Krippendorff’s α of 0.
50 due to boundary disagreements.
Trilling & Van Hoof (2020)
(1) TF-IDF weighing and then cosine similarity scores between articles.
(2) Amsterdam Embedding Model (pre-trained embeddings) and then softcosine similarity scores between articles.
3-day sliding window accounted for weekends
Representing pairwise similarity metrics as similarity network and then network partitioning into clusters.
Algorithms:
Leiden algorithm
Robustness check: Infomap method
Randomly annotated 100 events identified by the model.
The conservative cosine method achieved 89% precision for events and 94.
39% precision for articles.
The soft cosine method achieved a precision of 75% for events and 86.
92% for articles.
Gedikli et al.
(2021)
Co-occurrence of named entities (persons, locations, organizations) with a transformer-based Named Entity Recognition model, calculated via the Named Entities Shared Measure (NESM)
No time window.
Clustering based on similarity measure and similarity cutoff.
Validate pairwise similarity scores; 99.
4% agreement, Krippendorff’s α of 0.
87.
Litterer et al.
(2023)
Fine-tuned transformer embedding model: bi-encoder MPNet.
Document-level embeddings are created.
8-day sliding window and named entity in common.
Cosine similarity cutoff (> 0.
9).
The remaining article pairs are represented as a graph, and story clusters are extracted as the connected components of this graph.
Evaluation of the automated similarity model against human annotations from a SemEval competition dataset, achieving a mean Pearson correlation r of 0.
86.
References
Desai, A.
, & Nagwanshi, P.
(2020).
Grouping news events using semantic representations of hierarchical elements of articles and named entities.
2020 3rd International Conference on Algorithms, Computing and Artificial Intelligence, 1–6.
https://doi.
org/10.
1145/3446132.
3446399
Gedikli, F.
, Stockem Novo, A.
, & Jannach, D.
(2021).
Semi-automated identification of news story chains: A new dataset and entity-based labeling method.
Proceedings of the 9th International Workshop on News Recommendation and Analytics (INRA 2021), 29–42.
https://ceur-ws.
org/Vol-3143/paper3.
pdf
Kapellas, N.
, & Kapidakis, S.
(2022).
A Text Similarity Study: Understanding How Differently Greek News Media Describe News Events: Proceedings of the 14th International Joint Conference on Knowledge Discovery, Knowledge Engineering and Knowledge Management, 245–252.
https://doi.
org/10.
5220/0011589700003335
Langer, A.
I.
, & Gruber, J.
B.
(2021).
Political Agenda Setting in the Hybrid Media System: Why Legacy Media Still Matter a Great Deal.
The International Journal of Press/Politics, 26(2), 313–340.
https://doi.
org/10.
1177/1940161220925023
Litterer, B.
, Jurgens, D.
, & Card, D.
(2023).
When it Rains, it Pours: Modeling Media Storms and the News Ecosystem.
Findings of the Association for Computational Linguistics: EMNLP 2023, 6346–6361.
https://doi.
org/10.
18653/v1/2023.
findings-emnlp.
420
Markus, D.
K.
, Levi, E.
, Sheafer, T.
, & Shenhav, S.
R.
(2024).
Reap the Wild Wind: Detecting Media Storms in Large-Scale News Corpora.
In Y.
Al-Onaizan, M.
Bansal, & Y.
-N.
Chen (Eds.
), Findings of the Association for Computational Linguistics: EMNLP 2024 (pp.
4786–4797).
Association for Computational Linguistics.
https://doi.
org/10.
18653/v1/2024.
findings-emnlp.
275
Markus, D.
K.
, Levi, E.
, Sheafer, T.
, & Shenhav, S.
R.
(2026).
From Showers to Hurricanes: A Typology of Media Storms.
Political Communication, 1–19.
https://doi.
org/10.
1080/10584609.
2026.
2656633
Markus, D.
K.
, Mor-Lan, G.
, Sheafer, T.
, & Shenhav, S.
R.
(2023).
Leveraging Researcher Domain Expertise to Annotate Concepts Within Imbalanced Data.
Communication Methods and Measures, 17(3), 250–271.
https://doi.
org/10.
1080/19312458.
2023.
2182278
Nicholls, T.
, & Bright, J.
(2019).
Understanding News Story Chains using Information Retrieval and Network Clustering Techniques.
Communication Methods and Measures, 13(1), 43–59.
https://doi.
org/10.
1080/19312458.
2018.
1536972
Keith-Norambuena, B.
, Mitra, T.
, & North, C.
(2023).
A Survey on Event-Based News Narrative Extraction.
ACM Computing Surveys, 55(14s), 300:1-300:39.
https://doi.
org/10.
1145/3584741
Polimeno, A.
, Reuver, M.
, Vrijenhoek, S.
, & Fokkens, A.
(2023).
Improving and Evaluating the Detection of Fragmentation in News Recommendations with the Clustering of News Story Chains (arXiv:2309.
06192).
arXiv.
https://doi.
org/10.
48550/arXiv.
2309.
06192
Trilling, D.
, & Van Hoof, M.
(2020).
Between Article and Topic: News Events as Level of Analysis and Their Computational Identification.
Digital Journalism, 8(10), 1317–1337.
https://doi.
org/10.
1080/21670811.
2020.
1839352
Vasterman, P.
L.
M.
(2005).
Media-Hype: Self-Reinforcing News Waves, Journalistic Standards and the Construction of Social Problems.
European Journal of Communication, 20(4), 508–530.
https://doi.
org/10.
1177/0267323105058254
Welbers, K.
, Van Atteveldt, W.
, Bajjalieh, J.
, Shalmon, D.
, Joshi, P.
V.
, Althaus, S.
, Chan, C.
-H.
, Wessler, H.
, & Jungblut, M.
(2022).
Linking event archives to news: A computational method for analyzing the gatekeeping process.
Communication Methods and Measures, 16(1), 59–78.
https://doi.
org/10.
1080/19312458.
2021.
1953455
.
Related Results
Types of Media Outlets (Formats and Genre)
Types of Media Outlets (Formats and Genre)
“Types of media outlets”, often referred to as “media type” or “medium type”, is a variable that is widely used for content analyses of news media. The variable indicates which med...
Makna Voice Over dalam Pemberitaan Feature di Televisi
Makna Voice Over dalam Pemberitaan Feature di Televisi
Abstract. Voice Over or what is known as VO is being discussed a lot, not only about the profession, but also from the industry side and the various voice over techniques used. Due...
Event Management Bandung Sneaker Season
Event Management Bandung Sneaker Season
Abstract. Bandung Sneaker Season is the first sneakers and streetwear event to be held in Bandung, an annual event that was first created in 2018 by Maks.co Event Organizer. At the...
The Canberra Bubble
The Canberra Bubble
According to the ABC television program Four Corners, “Parliament House in Canberra is a hotbed of political intrigue and high tension … . It’s known as the ‘Canberra Bubble’ and i...
Reconstructing the Media Space of Digital News from Visualization to Spatial Immersion in the Case of “Dong News”
Reconstructing the Media Space of Digital News from Visualization to Spatial Immersion in the Case of “Dong News”
The rapid development of motion news highlights the necessity of exploring spatial transformations in news communication, particularly the evolution from two-dimensional news visua...
Understanding the Research Challenges in Low-Resource Language and Linking Bilingual News Articles in Multilingual News Archive
Understanding the Research Challenges in Low-Resource Language and Linking Bilingual News Articles in Multilingual News Archive
The developed world has focused on Web preservation compared to the developing world, especially news preservation for future generations. However, the news published online is vol...
Research on Different Strategies in Producing News Programs of The Two Television Stations: Lao National Television Station and Vientiane Capital Television Station
Research on Different Strategies in Producing News Programs of The Two Television Stations: Lao National Television Station and Vientiane Capital Television Station
The aim of this study was to perform a comparative analysis of the news production and transmission tactics of two television stations: Vientiane Capital Television Station and Lao...
Research on Different Strategies in Producing News Programs of The Two Television Stations: Lao National Television Station and Vientiane Capital Television Station
Research on Different Strategies in Producing News Programs of The Two Television Stations: Lao National Television Station and Vientiane Capital Television Station
The aim of this study was to perform a comparative analysis of the news production and transmission tactics of two television stations: Vientiane Capital Television Station and Lao...

