Search engine for discovering works of Art, research articles, and books related to Art and Culture
ShareThis
Javascript must be enabled to continue!

Exploiting Wikipedia Semantics for Computing Word Associations

View through CrossRef
<p><b>Semantic association computation is the process of automatically quantifying the strength of a semantic connection between two textual units based on various lexical and semantic relations such as hyponymy (car and vehicle) and functional associations (bank and manager). Humans have can infer implicit relationships between two textual units based on their knowledge about the world and their ability to reason about that knowledge. Automatically imitating this behavior is limited by restricted knowledge and poor ability to infer hidden relations.</b></p> <p>Various factors affect the performance of automated approaches to computing semantic association strength. One critical factor is the selection of a suitable knowledge source for extracting knowledge about the implicit semantic relations. In the past few years, semantic association computation approaches have started to exploit web-originated resources as substitutes for conventional lexical semantic resources such as thesauri, machine readable dictionaries and lexical databases. These conventional knowledge sources suffer from limitations such as coverage issues, high construction and maintenance costs and limited availability. To overcome these issues one solution is to use the wisdom of crowds in the form of collaboratively constructed knowledge sources. An excellent example of such knowledge sources is Wikipedia which stores detailed information not only about the concepts themselves but also about various aspects of the relations among concepts.</p> <p>The overall goal of this thesis is to demonstrate that using Wikipedia for computing word association strength yields better estimates of humans' associations than the approaches based on other structured and unstructured knowledge sources. There are two key challenges to achieve this goal: first, to exploit various semantic association models based on different aspects of Wikipedia in developing new measures of semantic associations; and second, to evaluate these measures compared to human performance in a range of tasks. The focus of the thesis is on exploring two aspects of Wikipedia: as a formal knowledge source, and as an informal text corpus.</p> <p>The first contribution of the work included in the thesis is that it effectively exploited the knowledge source aspect of Wikipedia by developing new measures of semantic associations based on Wikipedia hyperlink structure, informative-content of articles and combinations of both elements. It was found that Wikipedia can be effectively used for computing noun-noun similarity. It was also found that a model based on hybrid combinations of Wikipedia structure and informative-content based features performs better than those based on individual features. It was also found that the structure based measures outperformed the informative content based measures on both semantic similarity and semantic relatedness computation tasks.</p> <p>The second contribution of the research work in the thesis is that it effectively exploited the corpus aspect of Wikipedia by developing a new measure of semantic association based on asymmetric word associations. The thesis introduced the concept of asymmetric associations based measure using the idea of directional context inspired by the free word association task. The underlying assumption was that the association strength can change with the changing context. It was found that the asymmetric association based measure performed better than the symmetric measures on semantic association computation, relatedness based word choice and causality detection tasks. However, asymmetric-associations based measures have no advantage for synonymy-based word choice tasks. It was also found that Wikipedia is not a good knowledge source for capturing verb-relations due to its focus on encyclopedic concepts specially nouns.</p> <p>It is hoped that future research will build on the experiments and discussions presented in this thesis to explore new avenues using Wikipedia for finding deeper and semantically more meaningful associations in a wide range of application areas based on humans' estimates of word associations.</p>
Victoria University of Wellington Library
Title: Exploiting Wikipedia Semantics for Computing Word Associations
Description:
<p><b>Semantic association computation is the process of automatically quantifying the strength of a semantic connection between two textual units based on various lexical and semantic relations such as hyponymy (car and vehicle) and functional associations (bank and manager).
Humans have can infer implicit relationships between two textual units based on their knowledge about the world and their ability to reason about that knowledge.
Automatically imitating this behavior is limited by restricted knowledge and poor ability to infer hidden relations.
</b></p> <p>Various factors affect the performance of automated approaches to computing semantic association strength.
One critical factor is the selection of a suitable knowledge source for extracting knowledge about the implicit semantic relations.
In the past few years, semantic association computation approaches have started to exploit web-originated resources as substitutes for conventional lexical semantic resources such as thesauri, machine readable dictionaries and lexical databases.
These conventional knowledge sources suffer from limitations such as coverage issues, high construction and maintenance costs and limited availability.
To overcome these issues one solution is to use the wisdom of crowds in the form of collaboratively constructed knowledge sources.
An excellent example of such knowledge sources is Wikipedia which stores detailed information not only about the concepts themselves but also about various aspects of the relations among concepts.
</p> <p>The overall goal of this thesis is to demonstrate that using Wikipedia for computing word association strength yields better estimates of humans' associations than the approaches based on other structured and unstructured knowledge sources.
There are two key challenges to achieve this goal: first, to exploit various semantic association models based on different aspects of Wikipedia in developing new measures of semantic associations; and second, to evaluate these measures compared to human performance in a range of tasks.
The focus of the thesis is on exploring two aspects of Wikipedia: as a formal knowledge source, and as an informal text corpus.
</p> <p>The first contribution of the work included in the thesis is that it effectively exploited the knowledge source aspect of Wikipedia by developing new measures of semantic associations based on Wikipedia hyperlink structure, informative-content of articles and combinations of both elements.
It was found that Wikipedia can be effectively used for computing noun-noun similarity.
It was also found that a model based on hybrid combinations of Wikipedia structure and informative-content based features performs better than those based on individual features.
It was also found that the structure based measures outperformed the informative content based measures on both semantic similarity and semantic relatedness computation tasks.
</p> <p>The second contribution of the research work in the thesis is that it effectively exploited the corpus aspect of Wikipedia by developing a new measure of semantic association based on asymmetric word associations.
The thesis introduced the concept of asymmetric associations based measure using the idea of directional context inspired by the free word association task.
The underlying assumption was that the association strength can change with the changing context.
It was found that the asymmetric association based measure performed better than the symmetric measures on semantic association computation, relatedness based word choice and causality detection tasks.
However, asymmetric-associations based measures have no advantage for synonymy-based word choice tasks.
It was also found that Wikipedia is not a good knowledge source for capturing verb-relations due to its focus on encyclopedic concepts specially nouns.
</p> <p>It is hoped that future research will build on the experiments and discussions presented in this thesis to explore new avenues using Wikipedia for finding deeper and semantically more meaningful associations in a wide range of application areas based on humans' estimates of word associations.
</p>.

Related Results

Woningcorporaties en Vastgoedontwikkeling
Woningcorporaties en Vastgoedontwikkeling
This summary highlights the findings of the PhD-thesis ‘Woningcorporaties en Vastgoedontwikkeling: Fit for Use’ (‘Housing associations and Real Estate Development: Fit for Use?’). ...
Where are (women) planetary scientists on Wikipedia?
Where are (women) planetary scientists on Wikipedia?
BackgroundWikipedia is an open source, web-based encyclopedia, and allows anonymous and registered users to edit and create articles. This means that anyone can create, edit and im...
An empirical examination of Wikipedia's credibility
An empirical examination of Wikipedia's credibility
Wikipedia is an free, online encyclopaedia which anyone can add content to or edit the existing content of. The idea behind Wikipedia is that members of the general public can add ...
Wikipedia: a tool to monitor seasonal diseases trends?
Wikipedia: a tool to monitor seasonal diseases trends?
ObjectiveTo explore the interest of Wikipedia as a data source to monitorseasonal diseases trends in metropolitan France.IntroductionToday, Internet, especially Wikipedia, is an im...
Wikipedia in Vascular Surgery Medical Education: Comparative Study (Preprint)
Wikipedia in Vascular Surgery Medical Education: Comparative Study (Preprint)
BACKGROUND Medical students commonly refer to Wikipedia as their preferred online resource for medical information. The quality and readability of articles ...
COVID-19 research in Wikipedia
COVID-19 research in Wikipedia
Wikipedia is one of the main sources of free knowledge on the Web. During the first few months of the pandemic, over 5,200 new Wikipedia pages on COVID-19 were created, accumulatin...
COVID-19 research in Wikipedia
COVID-19 research in Wikipedia
Abstract Wikipedia is one of the main sources of free knowledge on the Web. During the first few months of the pandemic, over 5,200 new Wikipedia...
Wikipedia citations: Reproducible citation extraction from multilingual Wikipedia
Wikipedia citations: Reproducible citation extraction from multilingual Wikipedia
Abstract Wikipedia is an essential component of the open science ecosystem, yet it is poorly integrated with academic open science initiatives. Wikipedia Citation...

Back to Top