Javascript must be enabled to continue!
Applying (semi-)automatic metadata to early modern normative texts. Annif and Policeygesetzgebung from the City-State of Bern (1528–1798)
View through CrossRef
Abstract
This study investigates the application of modern digital tools to analyze handwritten normative texts from the City-State of Bern (1528-1798). By leveraging the Simple Knowledge Organization System (SKOS) and automatic metadata generation tools such as Transkribus for Handwritten Text Recognition (HTR) and Annif for automated subject indexing, we aim to enhance the efficiency and accuracy of historical text analysis. Our methodology involves digitizing 4,550 police ordinances, manually segmenting them, and employing a hierarchical SKOS framework with 1,830 distinct labels (Rhonda Organisation 2024). Annif’s backends, including TF-IDF, Umikuji Parabel, Umikuji Bonsai, and Machine Learning Language Model, were evaluated using Normalized Discounted Cumulative Gain (NDCG) metrics to generate relevant metadata. The results demonstrate that the Umikuji backends achieved high NDCG scores, highlighting their effectiveness. This research highlights the transformative potential of integrating digital tools in the humanities, providing a scalable solution for managing large historical datasets and enhancing access to historical texts. The findings advocate for broader adoption of these technologies in historical research, promoting greater collaboration and innovation. Future research will explore the application of these methods to other historical corpora and additional backend systems.
Oxford University Press (OUP)
Title: Applying (semi-)automatic metadata to early modern normative texts. Annif and
Policeygesetzgebung
from the City-State of Bern (1528–1798)
Description:
Abstract
This study investigates the application of modern digital tools to analyze handwritten normative texts from the City-State of Bern (1528-1798).
By leveraging the Simple Knowledge Organization System (SKOS) and automatic metadata generation tools such as Transkribus for Handwritten Text Recognition (HTR) and Annif for automated subject indexing, we aim to enhance the efficiency and accuracy of historical text analysis.
Our methodology involves digitizing 4,550 police ordinances, manually segmenting them, and employing a hierarchical SKOS framework with 1,830 distinct labels (Rhonda Organisation 2024).
Annif’s backends, including TF-IDF, Umikuji Parabel, Umikuji Bonsai, and Machine Learning Language Model, were evaluated using Normalized Discounted Cumulative Gain (NDCG) metrics to generate relevant metadata.
The results demonstrate that the Umikuji backends achieved high NDCG scores, highlighting their effectiveness.
This research highlights the transformative potential of integrating digital tools in the humanities, providing a scalable solution for managing large historical datasets and enhancing access to historical texts.
The findings advocate for broader adoption of these technologies in historical research, promoting greater collaboration and innovation.
Future research will explore the application of these methods to other historical corpora and additional backend systems.
Related Results
Big Metadata, Smart Metadata, and Metadata Capital: Toward Greater Synergy Between Data Science and Metadata
Big Metadata, Smart Metadata, and Metadata Capital: Toward Greater Synergy Between Data Science and Metadata
Abstract
Purpose
The purpose of the paper is to provide a framework for addressing the disconnect between metadata and data scie...
Literature Review on Metadata Governance
Literature Review on Metadata Governance
The framework of metadata governance is a subset of the primary data governance framework implementation within an enterprise. Metadata management helps identify data provenance an...
FAIR Digital Objects in Official Statistics
FAIR Digital Objects in Official Statistics
Introduction*1
Statistical offices on national and international scale provide statistics on demography, labour, income, society, economy, environment and othe...
Automatic Metadata Mining from Multilingual Enterprise Content
Automatic Metadata Mining from Multilingual Enterprise Content
Personalization is increasingly vital especially for enterprises to be able to reach their customers. The key challenge in supporting personalization is the need for rich metadata,...
Globally Findable Planetary Data: The Interdisciplinary TRR170-DB Repository
Globally Findable Planetary Data: The Interdisciplinary TRR170-DB Repository
Introduction: The TRR170-DB data repository (https://planetary-data-portal.org/) manages the research data from the collaborative research center ‘Late Accretion onto Ter...
Metadata quality and interoperability of GLAM digital images
Metadata quality and interoperability of GLAM digital images
PurposeThis study aims to explore how metadata have been applied in GLAM (galleries, libraries, archives and museums) institutions in New Zealand (NZ) and to analyse its overall qu...
Metadata Schema for Managing Digital Data and Images of Thai Human Skulls
Metadata Schema for Managing Digital Data and Images of Thai Human Skulls
This research was aimed at developing metadata that meets international standards for the purpose of managing digital data and images of Thai human skulls for medical studies. The ...

