Search engine for discovering works of Art, research articles, and books related to Art and Culture
ShareThis
Javascript must be enabled to continue!

Rough Set Based Analysis of Document Images

View through CrossRef
Rough set is a well-studied subject with a theoretical foundation and many applications. However, its usage in image processing has been very sparse. Most of the well-known algorithms for document image processing related to character recognition, character spotting, and logo retrieval resort to supervised classification, causing the system to slow down in the speed with increasing diversity in the documents, as well as the need to have a large training dataset. Hence, with an aim to resolve the tediousness and pitfalls of training, but without compromising on the efficiency, we introduce a rough-set-theoretic model. It is designed to perform an unsupervised classification of optical characters and logos with a small subset of attributes, called the semi-reduct. The semi-reduct attributes are mostly geometric and topological in nature, each having a small range of discrete values estimated from different combinatorial characteristics of rough-set approximations. This eventually leads to quick and easy discernibility of almost all the characters and logos. In this thesis, we first explain the basics of rough set theory. Subsequently, we propose various attributes that can be easily computed from the binary representation of the images. In subsequent chapters we show how one can select an appropriate subset of such attributes, known as semi-reduct, to perform a document processing task. We demonstrate in this thesis that using the above attributes one can design a character recognition system that is both computationally and storage efficient. Using a different semi-reduct, we show that one can also solve the very delicate task of character spotting in ancient inscriptions. Additionally, we propose appropriate pre-processing steps to binarize the old and dilapidated inscriptions. Finally, we propose a novel technique for logo retrieval using a suitably prepared semi-reduct. Comparison with other existing techniques substantiates our claim that attributes from the rough set are indeed good candidates for document image processing.
Center for Open Science
Title: Rough Set Based Analysis of Document Images
Description:
Rough set is a well-studied subject with a theoretical foundation and many applications.
However, its usage in image processing has been very sparse.
Most of the well-known algorithms for document image processing related to character recognition, character spotting, and logo retrieval resort to supervised classification, causing the system to slow down in the speed with increasing diversity in the documents, as well as the need to have a large training dataset.
Hence, with an aim to resolve the tediousness and pitfalls of training, but without compromising on the efficiency, we introduce a rough-set-theoretic model.
It is designed to perform an unsupervised classification of optical characters and logos with a small subset of attributes, called the semi-reduct.
The semi-reduct attributes are mostly geometric and topological in nature, each having a small range of discrete values estimated from different combinatorial characteristics of rough-set approximations.
This eventually leads to quick and easy discernibility of almost all the characters and logos.
In this thesis, we first explain the basics of rough set theory.
Subsequently, we propose various attributes that can be easily computed from the binary representation of the images.
In subsequent chapters we show how one can select an appropriate subset of such attributes, known as semi-reduct, to perform a document processing task.
We demonstrate in this thesis that using the above attributes one can design a character recognition system that is both computationally and storage efficient.
Using a different semi-reduct, we show that one can also solve the very delicate task of character spotting in ancient inscriptions.
Additionally, we propose appropriate pre-processing steps to binarize the old and dilapidated inscriptions.
Finally, we propose a novel technique for logo retrieval using a suitably prepared semi-reduct.
Comparison with other existing techniques substantiates our claim that attributes from the rough set are indeed good candidates for document image processing.

Related Results

Spectroscopic and transition properties of SeH<sup>–</sup> anion including spin-orbit coupling
Spectroscopic and transition properties of SeH<sup>–</sup> anion including spin-orbit coupling
<sec>Potential energy curves (PECs), permanent dipole moments (PDMs) and transition dipole moments (TMDs) of five Λ-S states of SeH<sup>−</sup> anion are calculat...
Theoretical study of laser-cooled SH<sup>–</sup> anion
Theoretical study of laser-cooled SH<sup>–</sup> anion
The potential energy curves, dipole moments, and transition dipole moments for the <inline-formula><tex-math id="M13">\begin{document}${{\rm{X}}^1}{\Sigma ^ + }$\end{do...
Ab initio study on the hydrogen desorption from $\rm {MH\text{&#x2013;}NH}_3$MH–NH3 (M = Li, Na, K) hydrogen storage systems
Ab initio study on the hydrogen desorption from $\rm {MH\text{&#x2013;}NH}_3$MH–NH3 (M = Li, Na, K) hydrogen storage systems
The hydrogen storage system LiH + \documentclass[12pt]{minimal}\begin{document}$\rm {NH}_3$\end{document} NH 3 ↔ \documentclass[12pt]{minimal}\begin{document}$\rm {LiNH}_2$\end{doc...
Rough set theory for document clustering: A review
Rough set theory for document clustering: A review
Rough set theory is a mathematical framework that can be visualized as a soft computing tool dealing with the vagueness and uncertainty of data and is applied to pattern recognitio...
Revisiting near-threshold photoelectron interference in argon with a non-adiabatic semiclassical model
Revisiting near-threshold photoelectron interference in argon with a non-adiabatic semiclassical model
<sec> <b>Purpose:</b> The interaction of intense, ultrashort laser pulses with atoms gives rise to rich non-perturbative phenomena, which are encoded within th...
Frequency of Common Chromosomal Abnormalities in Patients with Idiopathic Acquired Aplastic Anemia
Frequency of Common Chromosomal Abnormalities in Patients with Idiopathic Acquired Aplastic Anemia
Objective: To determine the frequency of common chromosomal aberrations in local population idiopathic determine the frequency of common chromosomal aberrations in local population...
Ukrainian Embroidery as a Type of Document
Ukrainian Embroidery as a Type of Document
The purpose of the article is to determine the general and specific features of Ukrainian embroidery as a type of carrier of documented information. The methodology. We chose the ...
Transformation of recording features in an electronic environment
Transformation of recording features in an electronic environment
The article deals with one of the main theoretical problems of document science related to the definition of document features. This problem is also of applied importance, since wh...

Back to Top