Javascript must be enabled to continue!
Understanding reproducibility of bioinformatics workflows
View through CrossRef
Reproducibility is an essential factor in establishing the reliability of results reported in the scientific literature and ultimately in the knowledge that is generated. This knowledge demands that experiments should be verifiable and potentially extended through open and transparent processes. This need for reproducibility is magnified in the field of genomics, where advances in massively parallel DNA sequencing technologies have resulted in an explosion of data and increased dependence on computational analysis of this data. Workflows are one common way through which such data is analyzed. In this context, there is an absolute need for reliable, reproducible, effective and timely interpretation of genomics data sets especially since they can be used for understanding of diseases and support personalized medicine. Faced with the complexity of ge- nomics datasets and the plethora of bioinformatics tools and workflow systems that have emerged over the last few years, a major challenge is supporting re- producibility of bioinformatics workflows. This is complicated by the current incomplete understanding regarding requirements for assuring reproducibility of workflows and assumptions relating to the many different approaches that can be taken when defining and enacting (executing) workflows. This problem is exacerbated by the lack of clarity in terminologies and concepts associated with reproducibility.
The objective of this project is to empirically study the reproducibility requirements associated with bioinformatics workflows. This research highlights the need for, and critical dimensions of, reproducibility of bioinformatics workflows through a case study exploring the genomic analysis of a rare genetic disorder. In addition, it focuses on the reproducibility challenges that arose during the data analysis by the different (independent) teams. We subsequently categorize widely used workflow definition and implementation approaches for genomic data analysis and offer detailed implementation discussion of a complex but widely adopted genomic data analysis workflow in a selected exemplar from each of the workflow categories. Through empirical analysis, we identify assumptions that were implicit in the exemplar approaches, resulting in insufficient documentation of workflow requirements that impacted directly on the reproducibility of the results.
Building on this, we identify key implicit and explicit requirements that impact directly on the reproducibility of bioinformatics workflows. We characterize these implicit and explicit requirements according to critical dimensions of reproducibility to establish a reproducibility framework that considers different categories of workflow approaches. The framework offers a coherent explanation and exploration of the concepts associated with workflow reproducibility. This research illustrates the benefits of the framework by demonstrating the portability of a fully declarative complex workflow to an independent workflow platform to support reproducible experiments. The evaluation illustrates that the highest degree of reproducibility is supported by workflow definition approaches that make the least assumptions on the software environment.
Title: Understanding reproducibility of bioinformatics workflows
Description:
Reproducibility is an essential factor in establishing the reliability of results reported in the scientific literature and ultimately in the knowledge that is generated.
This knowledge demands that experiments should be verifiable and potentially extended through open and transparent processes.
This need for reproducibility is magnified in the field of genomics, where advances in massively parallel DNA sequencing technologies have resulted in an explosion of data and increased dependence on computational analysis of this data.
Workflows are one common way through which such data is analyzed.
In this context, there is an absolute need for reliable, reproducible, effective and timely interpretation of genomics data sets especially since they can be used for understanding of diseases and support personalized medicine.
Faced with the complexity of ge- nomics datasets and the plethora of bioinformatics tools and workflow systems that have emerged over the last few years, a major challenge is supporting re- producibility of bioinformatics workflows.
This is complicated by the current incomplete understanding regarding requirements for assuring reproducibility of workflows and assumptions relating to the many different approaches that can be taken when defining and enacting (executing) workflows.
This problem is exacerbated by the lack of clarity in terminologies and concepts associated with reproducibility.
The objective of this project is to empirically study the reproducibility requirements associated with bioinformatics workflows.
This research highlights the need for, and critical dimensions of, reproducibility of bioinformatics workflows through a case study exploring the genomic analysis of a rare genetic disorder.
In addition, it focuses on the reproducibility challenges that arose during the data analysis by the different (independent) teams.
We subsequently categorize widely used workflow definition and implementation approaches for genomic data analysis and offer detailed implementation discussion of a complex but widely adopted genomic data analysis workflow in a selected exemplar from each of the workflow categories.
Through empirical analysis, we identify assumptions that were implicit in the exemplar approaches, resulting in insufficient documentation of workflow requirements that impacted directly on the reproducibility of the results.
Building on this, we identify key implicit and explicit requirements that impact directly on the reproducibility of bioinformatics workflows.
We characterize these implicit and explicit requirements according to critical dimensions of reproducibility to establish a reproducibility framework that considers different categories of workflow approaches.
The framework offers a coherent explanation and exploration of the concepts associated with workflow reproducibility.
This research illustrates the benefits of the framework by demonstrating the portability of a fully declarative complex workflow to an independent workflow platform to support reproducible experiments.
The evaluation illustrates that the highest degree of reproducibility is supported by workflow definition approaches that make the least assumptions on the software environment.
Related Results
A large-scale analysis of bioinformatics code on GitHub
A large-scale analysis of bioinformatics code on GitHub
AbstractIn recent years, the explosion of genomic data and bioinformatic tools has been accompanied by a growing conversation around reproducibility of results and usability of sof...
Advancements in Biomedical and Bioinformatics Engineering
Advancements in Biomedical and Bioinformatics Engineering
Abstract: The field of biomedical and bioinformatics engineering is witnessing rapid advancements that are revolutionizing healthcare and medical research. This chapter provides a...
From high school to postdoc: Lessons from a decade of bioinformatics education
From high school to postdoc: Lessons from a decade of bioinformatics education
As a postdoctoral research fellow with both a PhD and a bachelor’s degree in bioinformatics, my scientific background is the product of over a decade of bioinformatics training and...
New classifications for quantum bioinformatics: Q-bioinformatics, QCt-bioinformatics, QCg-bioinformatics, and QCr-bioinformatics
New classifications for quantum bioinformatics: Q-bioinformatics, QCt-bioinformatics, QCg-bioinformatics, and QCr-bioinformatics
Abstract
Bioinformatics has revolutionized biology and medicine by using computational methods to analyze and interpret biological data. Quantum mechanics has recent...
Improving Reproducibility in AI Research: Four Mechanisms Adopted by JAIR
Improving Reproducibility in AI Research: Four Mechanisms Adopted by JAIR
Background: Lately, the reproducibility of scientific results has become an increasing worry in the scientific community. Several studies show that artificial intelligence research...
Improving bioinformatics software quality through incorporation of software engineering practices
Improving bioinformatics software quality through incorporation of software engineering practices
BackgroundBioinformatics software is developed for collecting, analyzing, integrating, and interpreting life science datasets that are often enormous. Bioinformatics engineers ofte...
CREDO: a friendly Customizable, REproducible, DOcker file generator for bioinformatics applications
CREDO: a friendly Customizable, REproducible, DOcker file generator for bioinformatics applications
Abstract
Background
The analysis of large and complex biological datasets in bioinformatics poses a significant challenge to achieving reproducible ...
Chem-bioinformatics: Computational Alternatives to Clinical Diagnosis, Treatment and Preventative Measures
Chem-bioinformatics: Computational Alternatives to Clinical Diagnosis, Treatment and Preventative Measures
Nowadays, chem-bioinformatics tools are widely used for genomic and
proteomic data analysis, gene prediction, genome annotation, expression profiling,
biological network building, ...

