Javascript must be enabled to continue!
Automatic Data Extraction Utilizing Structural Similarity From A Set of Portable Document Format (PDF) Files
View through CrossRef
Instead of storing data in databases, common computer-aided office workers often choose to keep data related to their work in the form of document or report files that they can conveniently and comfortably access with popular off-the-shelf softwares, such as in Portable Document Format (PDF) format files. Their workplaces may actually use databases but they usually do not possess the privilege nor the proficiency to fully utilize them. Said workplaces likely have front-end systems such as Management Information System (MIS) from where workers get their data containing reports or documents.These documents are meant for immediate or presentational uses but workers often keep these files for the data inside which may come to be useful later on. This way, they can manipulate and combine data from one or more report files to suit their work needs, on the occasions that their MIS were not able to fulfill such needs. To do this, workers need to extract data from the report files. However, the files also contain formatting and other contents such as organization banners, signature placeholders, and so on. Extracting data from these files is not easy and workers are often forced to use repeated copy and paste actions to get the data they want. This is not only tedious but also time-consuming and prone to errors. Automatic data extraction is not new, many existing solutions are available but they typically require human guidance to help the data extraction before it can become truly automatic. They may also require certain expertise which can make workers hesitant to use them in the first place. A particular function of an MIS can produce many report files, each containing distinct data, but still structurally similar. If we target all PDF files that come from such same source, in this paper we demonstrated that by exploiting the similarity it is possible to create a fully automatic data extraction system that requires no human guidance. First, a model is generated by analyzing a small sample of PDFs and then the model is used to extract data from all PDF files in the set. Our experiments show that the system can quickly achieve 100% accuracy rate with very few sample files. Though there are occasions where data inside all the PDFs are not sufficiently distinct from each other resulting in lower than 100% accuracy, this can be easily detected and fixed with slight human intervention. In these cases, total no human intervention may not be possible but the amount needed can be significantly reduced.
Universitas Sriwijaya - Pusat Inovasi Pembelajaran Unsri
Title: Automatic Data Extraction Utilizing Structural Similarity From A Set of Portable Document Format (PDF) Files
Description:
Instead of storing data in databases, common computer-aided office workers often choose to keep data related to their work in the form of document or report files that they can conveniently and comfortably access with popular off-the-shelf softwares, such as in Portable Document Format (PDF) format files.
Their workplaces may actually use databases but they usually do not possess the privilege nor the proficiency to fully utilize them.
Said workplaces likely have front-end systems such as Management Information System (MIS) from where workers get their data containing reports or documents.
These documents are meant for immediate or presentational uses but workers often keep these files for the data inside which may come to be useful later on.
This way, they can manipulate and combine data from one or more report files to suit their work needs, on the occasions that their MIS were not able to fulfill such needs.
To do this, workers need to extract data from the report files.
However, the files also contain formatting and other contents such as organization banners, signature placeholders, and so on.
Extracting data from these files is not easy and workers are often forced to use repeated copy and paste actions to get the data they want.
This is not only tedious but also time-consuming and prone to errors.
Automatic data extraction is not new, many existing solutions are available but they typically require human guidance to help the data extraction before it can become truly automatic.
They may also require certain expertise which can make workers hesitant to use them in the first place.
A particular function of an MIS can produce many report files, each containing distinct data, but still structurally similar.
If we target all PDF files that come from such same source, in this paper we demonstrated that by exploiting the similarity it is possible to create a fully automatic data extraction system that requires no human guidance.
First, a model is generated by analyzing a small sample of PDFs and then the model is used to extract data from all PDF files in the set.
Our experiments show that the system can quickly achieve 100% accuracy rate with very few sample files.
Though there are occasions where data inside all the PDFs are not sufficiently distinct from each other resulting in lower than 100% accuracy, this can be easily detected and fixed with slight human intervention.
In these cases, total no human intervention may not be possible but the amount needed can be significantly reduced.
.
Related Results
Theoretical study of laser-cooled SH<sup>–</sup> anion
Theoretical study of laser-cooled SH<sup>–</sup> anion
The potential energy curves, dipole moments, and transition dipole moments for the <inline-formula><tex-math id="M13">\begin{document}${{\rm{X}}^1}{\Sigma ^ + }$\end{do...
Ab initio study on the hydrogen desorption from $\rm {MH\text{–}NH}_3$MH–NH3 (M = Li, Na, K) hydrogen storage systems
Ab initio study on the hydrogen desorption from $\rm {MH\text{–}NH}_3$MH–NH3 (M = Li, Na, K) hydrogen storage systems
The hydrogen storage system LiH + \documentclass[12pt]{minimal}\begin{document}$\rm {NH}_3$\end{document} NH 3 ↔ \documentclass[12pt]{minimal}\begin{document}$\rm {LiNH}_2$\end{doc...
Indo-Anglian: Connotations and Denotations
Indo-Anglian: Connotations and Denotations
A different name than English literature, ‘Anglo-Indian Literature’, was given to the body of literature in English that emerged on account of the British interaction with India un...
Revisiting near-threshold photoelectron interference in argon with a non-adiabatic semiclassical model
Revisiting near-threshold photoelectron interference in argon with a non-adiabatic semiclassical model
<sec> <b>Purpose:</b> The interaction of intense, ultrashort laser pulses with atoms gives rise to rich non-perturbative phenomena, which are encoded within th...
Epi Archive: Automated Synthesis of Global Notifiable Disease Data
Epi Archive: Automated Synthesis of Global Notifiable Disease Data
ObjectiveLANL has built software that automatically collects global notifiable disease data, synthesizes the data, and makes it available to humans and computers within the Biosurv...
Form Follows Force: A theoretical framework for Structural Morphology, and Form-Finding research on shell structures
Form Follows Force: A theoretical framework for Structural Morphology, and Form-Finding research on shell structures
The springing up of freeform architecture and structures introduces many challenges to structural engineers. The main challenge is to generate structural forms with high structural...
O cuidado e suas dimensões: uma revisão bibliográfica
O cuidado e suas dimensões: uma revisão bibliográfica
Introdução: a problemática central deste artigo é o cuidado. O cuidado não como expressão única do tecnicismo, mas em suas múltiplas dimensões. Objetivo: discutir o cuidado numa pe...
January 2024 , Volume 22, Issue 1 Full pdf of issue Editorial Plea for Peace - Publisher World Family Medicine Health-Related Quality of Life (HRQoL) in Haemodialysis Patients in Khartoum, Sudan [Abstract] [pdf] Samira Khatir Ali Fadlalla DOI: 10.5742/
January 2024 , Volume 22, Issue 1 Full pdf of issue Editorial Plea for Peace - Publisher World Family Medicine Health-Related Quality of Life (HRQoL) in Haemodialysis Patients in Khartoum, Sudan [Abstract] [pdf] Samira Khatir Ali Fadlalla DOI: 10.5742/
While the bombing of Gaza and the resulting loss of civilians continues, I urge the international community to stop the war now, protect civilians (including health-care workers), ...

