Search engine for discovering works of Art, research articles, and books related to Art and Culture
ShareThis
Javascript must be enabled to continue!

BioWorkbench: a high-performance framework for managing and analyzing bioinformatics experiments

View through CrossRef
Advances in sequencing techniques have led to exponential growth in biological data, demanding the development of large-scale bioinformatics experiments. Because these experiments are computation- and data-intensive, they require high-performance computing techniques and can benefit from specialized technologies such as Scientific Workflow Management Systems and databases. In this work, we present BioWorkbench, a framework for managing and analyzing bioinformatics experiments. This framework automatically collects provenance data, including both performance data from workflow execution and data from the scientific domain of the workflow application. Provenance data can be analyzed through a web application that abstracts a set of queries to the provenance database, simplifying access to provenance information. We evaluate BioWorkbench using three case studies: SwiftPhylo, a phylogenetic tree assembly workflow; SwiftGECKO, a comparative genomics workflow; and RASflow, a RASopathy analysis workflow. We analyze each workflow from both computational and scientific domain perspectives, by using queries to a provenance and annotation database. Some of these queries are available as a pre-built feature of the BioWorkbench web application. Through the provenance data, we show that the framework is scalable and achieves high-performance, reducing up to 98% of the case studies execution time. We also show how the application of machine learning techniques can enrich the analysis process.
Title: BioWorkbench: a high-performance framework for managing and analyzing bioinformatics experiments
Description:
Advances in sequencing techniques have led to exponential growth in biological data, demanding the development of large-scale bioinformatics experiments.
Because these experiments are computation- and data-intensive, they require high-performance computing techniques and can benefit from specialized technologies such as Scientific Workflow Management Systems and databases.
In this work, we present BioWorkbench, a framework for managing and analyzing bioinformatics experiments.
This framework automatically collects provenance data, including both performance data from workflow execution and data from the scientific domain of the workflow application.
Provenance data can be analyzed through a web application that abstracts a set of queries to the provenance database, simplifying access to provenance information.
We evaluate BioWorkbench using three case studies: SwiftPhylo, a phylogenetic tree assembly workflow; SwiftGECKO, a comparative genomics workflow; and RASflow, a RASopathy analysis workflow.
We analyze each workflow from both computational and scientific domain perspectives, by using queries to a provenance and annotation database.
Some of these queries are available as a pre-built feature of the BioWorkbench web application.
Through the provenance data, we show that the framework is scalable and achieves high-performance, reducing up to 98% of the case studies execution time.
We also show how the application of machine learning techniques can enrich the analysis process.

Related Results

Advancements in Biomedical and Bioinformatics Engineering
Advancements in Biomedical and Bioinformatics Engineering
Abstract: The field of biomedical and bioinformatics engineering is witnessing rapid advancements that are revolutionizing healthcare and medical research. This chapter provides a...
A large-scale analysis of bioinformatics code on GitHub
A large-scale analysis of bioinformatics code on GitHub
AbstractIn recent years, the explosion of genomic data and bioinformatic tools has been accompanied by a growing conversation around reproducibility of results and usability of sof...
Improving bioinformatics software quality through incorporation of software engineering practices
Improving bioinformatics software quality through incorporation of software engineering practices
BackgroundBioinformatics software is developed for collecting, analyzing, integrating, and interpreting life science datasets that are often enormous. Bioinformatics engineers ofte...
From high school to postdoc: Lessons from a decade of bioinformatics education
From high school to postdoc: Lessons from a decade of bioinformatics education
As a postdoctoral research fellow with both a PhD and a bachelor’s degree in bioinformatics, my scientific background is the product of over a decade of bioinformatics training and...
Chem-bioinformatics: Computational Alternatives to Clinical Diagnosis, Treatment and Preventative Measures
Chem-bioinformatics: Computational Alternatives to Clinical Diagnosis, Treatment and Preventative Measures
Nowadays, chem-bioinformatics tools are widely used for genomic and proteomic data analysis, gene prediction, genome annotation, expression profiling, biological network building, ...
Baseline Evaluation of Bioinformatics Capacity in Tanzania
Baseline Evaluation of Bioinformatics Capacity in Tanzania
Abstract BackgroundEven though the genomics technologies have grown to a large extent, Sub Saharan Africa countries have not entirely reaped the benefits due to the lack of...
Bioinformatics tool and web server development focusing on structural bioinformatics applications
Bioinformatics tool and web server development focusing on structural bioinformatics applications
This thesis is divided into two main sections: Part 1 describes the design, and evaluation of the accuracy of a new web server – PRotein Interactive MOdeling (PRIMO-Complexes) for ...

Back to Top