Javascript must be enabled to continue!
Protocol: Investigating DOIs classes of errors v4
View through CrossRef
The purpose of this protocol is to provide an automated process to repair invalid DOIs that have been collected by the OpenCitations Index Of Crossref Open DOI-To-DOI References (COCI) while processing data provided by Crossref. The data needed for this work is provided by COCI as a CSV containing pairs of valid citing DOIs and invalid cited DOIs. With the goal to determine an automated process, we first classified the errors that characterize the wrong DOIs in the list. The starting hypothesis is that there are two main classes of errors: factual errors, such as wrong characters, and DOIs that are not yet valid at the time of processing. The first class can be furtherly divided into three classes: errors due to irrelevant strings added to the beginning (prefix-type errors) or at the end (suffix-type errors) of the correct DOI, and errors due to unwanted characters in the middle (other-type errors). Once the classes of errors are addressed, we propose automatic processes to obtain correct DOIs from wrong ones. These processes involve the use of the information returned from DOI API, the January 2021 Public Data File from Crossref, as well as rule-based methods, including regular expressions to correct invalid DOIs. The application of this methodology produced a CSV dataset containing all the pairs of citing and cited DOIs in the original dataset, each one enriched by 5 fields: "Already_Valid", which tells if the cited DOI was already valid before cleaning, "New_DOI", which contain a clean, valid DOI (if our procedure was able to produce one), and "prefix_error", "suffix_error" and "other-type_error" fields, which contain, for each cleaned DOI the number of errors that were cleaned.
Springer Science and Business Media LLC
Title: Protocol: Investigating DOIs classes of errors v4
Description:
The purpose of this protocol is to provide an automated process to repair invalid DOIs that have been collected by the OpenCitations Index Of Crossref Open DOI-To-DOI References (COCI) while processing data provided by Crossref.
The data needed for this work is provided by COCI as a CSV containing pairs of valid citing DOIs and invalid cited DOIs.
With the goal to determine an automated process, we first classified the errors that characterize the wrong DOIs in the list.
The starting hypothesis is that there are two main classes of errors: factual errors, such as wrong characters, and DOIs that are not yet valid at the time of processing.
The first class can be furtherly divided into three classes: errors due to irrelevant strings added to the beginning (prefix-type errors) or at the end (suffix-type errors) of the correct DOI, and errors due to unwanted characters in the middle (other-type errors).
Once the classes of errors are addressed, we propose automatic processes to obtain correct DOIs from wrong ones.
These processes involve the use of the information returned from DOI API, the January 2021 Public Data File from Crossref, as well as rule-based methods, including regular expressions to correct invalid DOIs.
The application of this methodology produced a CSV dataset containing all the pairs of citing and cited DOIs in the original dataset, each one enriched by 5 fields: "Already_Valid", which tells if the cited DOI was already valid before cleaning, "New_DOI", which contain a clean, valid DOI (if our procedure was able to produce one), and "prefix_error", "suffix_error" and "other-type_error" fields, which contain, for each cleaned DOI the number of errors that were cleaned.
Related Results
NICU Medication Errors: Describing the Cause and Nature of Medication Errors in a NICU in Qatar
NICU Medication Errors: Describing the Cause and Nature of Medication Errors in a NICU in Qatar
IntroductionA medication error can be defined as “any error occurring in the medication use process” and focuses on problems with the delivery of medication to a patient [1]. Medic...
THE EIGHTH GRADE STUDENTS’ ERRORS OF SMP NEGERI 6 LHOKSUKON ACEH UTARA IN WRITING NARRATIVE TEXT
THE EIGHTH GRADE STUDENTS’ ERRORS OF SMP NEGERI 6 LHOKSUKON ACEH UTARA IN WRITING NARRATIVE TEXT
The aims of the study is to find out the kinds of errors are made by the seventh grade students of SMP Negeri 6 Lhoksukon Aceh Utara in writing narrative text and the most domin...
OS SERVIDORES PÚBLICOS MUNICIPAIS
OS SERVIDORES PÚBLICOS MUNICIPAIS
I. Organização do funcionalismo municipal1. A Autonomia dos Municípios e a organização de seu funcionalismo — A Constituição Federal assegura, aos Municípios, a autonomia de autogo...
Comparison of the Influences of Initial Errors and Model Parameter Errors on Predictability of Numerical Forecast
Comparison of the Influences of Initial Errors and Model Parameter Errors on Predictability of Numerical Forecast
AbstractBased on the nonlinear local Lyapunov exponent (NLLE) approach introduced by the authors recently, the influences of initial errors and parameter errors on the predictabili...
Wastewater QC workflow in GalaxyTrakr (SSQuAWK4) v2
Wastewater QC workflow in GalaxyTrakr (SSQuAWK4) v2
PURPOSE: Step-by-step instructions for checking sequence quality for SARS-CoV-2 wastewater samples using SSQuAWK: SARS - CoV - 2 Sequence Quality Assurance Workflow and Kontrapti...
KEEFEKTIFAN KALIMAT DALAM TEKS BERITA SISWA KELAS VIII SMP NEGERI 16 PADANG
KEEFEKTIFAN KALIMAT DALAM TEKS BERITA SISWA KELAS VIII SMP NEGERI 16 PADANG
ABSTRACTThe purpose of this study is to describe the effectiveness of sentences in the student news text in terms of five things. First, the effectiveness of sentences in terms of ...
An Error Analysis of Students in Writing Narrative Text
An Error Analysis of Students in Writing Narrative Text
This research is aimed to find out the typical errors on students in writing narrative text by using simple past tense at grade IX of SMP SwastaTalitakum Medan. The researchers use...
PROFIL KESALAHAN SISWA DALAM MENYELESAIKAN SOAL MATRIKS BERDASARKAN JENIS KELAMIN DI SMA NEGERI 7 PALU
PROFIL KESALAHAN SISWA DALAM MENYELESAIKAN SOAL MATRIKS BERDASARKAN JENIS KELAMIN DI SMA NEGERI 7 PALU
Abstrak: Penelitian ini merupakan penelitian kualitatif yang bertujuan untuk memperoleh profil kesalahan yang dilakukan siswa dalam menyelesaikan soal matriks berdasarkan jenis kel...

