Search engine for discovering works of Art, research articles, and books related to Art and Culture
ShareThis
Javascript must be enabled to continue!

Automatic Metadata Mining from Multilingual Enterprise Content

View through CrossRef
Personalization is increasingly vital especially for enterprises to be able to reach their customers. The key challenge in supporting personalization is the need for rich metadata, such as metadata about structural relationships, subject/concept relations between documents and cognitive metadata about documents (e.g. difficulty of a document). Manual annotation of large knowledge bases with such rich metadata is not scalable. As well as, automatic mining of cognitive metadata is challenging since it is very difficult to understand underlying intellectual knowledge about document automatically. On the other hand, the Web content is increasing becoming multilingual since growing amount of data generated on the Web is non-English. Current metadata extraction systems are generally based on English content and this requires to be revolutionized in order to adapt to the changing dynamics of the Web. To alleviate these problems, we introduce a novel automatic metadata extraction framework, which is based on a novel fuzzy based method for automatic cognitive metadata generation and uses different document parsing algorithms to extract rich metadata from multilingual enterprise content using the newly developed DocBook, Resource Type and Topic ontologies. Since the metadata generation process is based upon DocBook structured enterprise content, our framework is focused on enterprise documents and content which is loosely based on the DocBook type of formatting. DocBook is a common documentation formatting to formally produce corporate data and it is adopted by many enterprises. The proposed framework is illustrated and evaluated on English, German and French versions of the Symantec Norton 360 knowledge bases. The user study showed that the proposed fuzzy-based method generates reasonably accurate values with an average precision of 89.39% on the metadata values of document difficulty, document interactivity level and document interactivity type. The proposed fuzzy inference system achieves improved results compared to a rulebased reasoner for difficulty metadata extraction (~11% enhancement). In addition, user perceived metadata quality scores (mean of 5.57 out of 6) found to be high and automated metadata analysis showed that the extracted metadata is high quality and can be suitable for personalized information retrieval.
Title: Automatic Metadata Mining from Multilingual Enterprise Content
Description:
Personalization is increasingly vital especially for enterprises to be able to reach their customers.
The key challenge in supporting personalization is the need for rich metadata, such as metadata about structural relationships, subject/concept relations between documents and cognitive metadata about documents (e.
g.
difficulty of a document).
Manual annotation of large knowledge bases with such rich metadata is not scalable.
As well as, automatic mining of cognitive metadata is challenging since it is very difficult to understand underlying intellectual knowledge about document automatically.
On the other hand, the Web content is increasing becoming multilingual since growing amount of data generated on the Web is non-English.
Current metadata extraction systems are generally based on English content and this requires to be revolutionized in order to adapt to the changing dynamics of the Web.
To alleviate these problems, we introduce a novel automatic metadata extraction framework, which is based on a novel fuzzy based method for automatic cognitive metadata generation and uses different document parsing algorithms to extract rich metadata from multilingual enterprise content using the newly developed DocBook, Resource Type and Topic ontologies.
Since the metadata generation process is based upon DocBook structured enterprise content, our framework is focused on enterprise documents and content which is loosely based on the DocBook type of formatting.
DocBook is a common documentation formatting to formally produce corporate data and it is adopted by many enterprises.
The proposed framework is illustrated and evaluated on English, German and French versions of the Symantec Norton 360 knowledge bases.
The user study showed that the proposed fuzzy-based method generates reasonably accurate values with an average precision of 89.
39% on the metadata values of document difficulty, document interactivity level and document interactivity type.
The proposed fuzzy inference system achieves improved results compared to a rulebased reasoner for difficulty metadata extraction (~11% enhancement).
In addition, user perceived metadata quality scores (mean of 5.
57 out of 6) found to be high and automated metadata analysis showed that the extracted metadata is high quality and can be suitable for personalized information retrieval.

Related Results

Literature Review on Metadata Governance
Literature Review on Metadata Governance
The framework of metadata governance is a subset of the primary data governance framework implementation within an enterprise. Metadata management helps identify data provenance an...
FAIR Digital Objects in Official Statistics
FAIR Digital Objects in Official Statistics
Introduction*1 Statistical offices on national and international scale provide statistics on demography, labour, income, society, economy, environment and othe...
Ontomet
Ontomet
Proper description of data, or metadata, is important to facilitate data sharing among Geospatial Information Communities. To avoid the production of arbitrary metadata annotations...
Globally Findable Planetary Data: The Interdisciplinary TRR170-DB Repository
Globally Findable Planetary Data: The Interdisciplinary TRR170-DB Repository
Introduction: The TRR170-DB data repository (https://planetary-data-portal.org/) manages the research data from the collaborative research center ‘Late Accretion onto Ter...
Light at the End of the Tunnel: Mining Justice and Health
Light at the End of the Tunnel: Mining Justice and Health
The mining industry provides valuable mined commodities and financial support for communities worldwide. Mining has become safer for workers. Significant injustices, however, are c...
Metadata quality and interoperability of GLAM digital images
Metadata quality and interoperability of GLAM digital images
PurposeThis study aims to explore how metadata have been applied in GLAM (galleries, libraries, archives and museums) institutions in New Zealand (NZ) and to analyse its overall qu...
Productivity Measure in Using Enterprise Resource Planning System in Selected Companies in Beijing, China
Productivity Measure in Using Enterprise Resource Planning System in Selected Companies in Beijing, China
With the globalization of economic development and social development, the business environment of enterprises has changed. Only by continuously improving the digital level and man...

Back to Top