Search engine for discovering works of Art, research articles, and books related to Art and Culture
ShareThis
Javascript must be enabled to continue!

Genomic benchmarks: a collection of datasets for genomic sequence classification

View through CrossRef
Abstract Background Recently, deep neural networks have been successfully applied in many biological fields. In 2020, a deep learning model AlphaFold won the protein folding competition with predicted structures within the error tolerance of experimental methods. However, this solution to the most prominent bioinformatic challenge of the past 50 years has been possible only thanks to a carefully curated benchmark of experimentally predicted protein structures. In Genomics, we have similar challenges (annotation of genomes and identification of functional elements) but currently, we lack benchmarks similar to protein folding competition. Results Here we present a collection of curated and easily accessible sequence classification datasets in the field of genomics. The proposed collection is based on a combination of novel datasets constructed from the mining of publicly available databases and existing datasets obtained from published articles. The collection currently contains nine datasets that focus on regulatory elements (promoters, enhancers, open chromatin region) from three model organisms: human, mouse, and roundworm. A simple convolution neural network is also included in a repository and can be used as a baseline model. Benchmarks and the baseline model are distributed as the Python package ‘genomic-benchmarks’, and the code is available at https://github.com/ML-Bioinfo-CEITEC/genomic_benchmarks . Conclusions Deep learning techniques revolutionized many biological fields but mainly thanks to the carefully curated benchmarks. For the field of Genomics, we propose a collection of benchmark datasets for the classification of genomic sequences with an interface for the most commonly used deep learning libraries, implementation of the simple neural network and a training framework that can be used as a starting point for future research. The main aim of this effort is to create a repository for shared datasets that will make machine learning for genomics more comparable and reproducible while reducing the overhead of researchers who want to enter the field, leading to healthy competition and new discoveries.
Title: Genomic benchmarks: a collection of datasets for genomic sequence classification
Description:
Abstract Background Recently, deep neural networks have been successfully applied in many biological fields.
In 2020, a deep learning model AlphaFold won the protein folding competition with predicted structures within the error tolerance of experimental methods.
However, this solution to the most prominent bioinformatic challenge of the past 50 years has been possible only thanks to a carefully curated benchmark of experimentally predicted protein structures.
In Genomics, we have similar challenges (annotation of genomes and identification of functional elements) but currently, we lack benchmarks similar to protein folding competition.
Results Here we present a collection of curated and easily accessible sequence classification datasets in the field of genomics.
The proposed collection is based on a combination of novel datasets constructed from the mining of publicly available databases and existing datasets obtained from published articles.
The collection currently contains nine datasets that focus on regulatory elements (promoters, enhancers, open chromatin region) from three model organisms: human, mouse, and roundworm.
A simple convolution neural network is also included in a repository and can be used as a baseline model.
Benchmarks and the baseline model are distributed as the Python package ‘genomic-benchmarks’, and the code is available at https://github.
com/ML-Bioinfo-CEITEC/genomic_benchmarks .
Conclusions Deep learning techniques revolutionized many biological fields but mainly thanks to the carefully curated benchmarks.
For the field of Genomics, we propose a collection of benchmark datasets for the classification of genomic sequences with an interface for the most commonly used deep learning libraries, implementation of the simple neural network and a training framework that can be used as a starting point for future research.
The main aim of this effort is to create a repository for shared datasets that will make machine learning for genomics more comparable and reproducible while reducing the overhead of researchers who want to enter the field, leading to healthy competition and new discoveries.

Related Results

Effective Use of Benchmarking: The Context of the Centre for Preparatory Studies in Oman
Effective Use of Benchmarking: The Context of the Centre for Preparatory Studies in Oman
This paper examines the importance and the process of developing benchmarks for the courses offered at the Centre for Preparatory Studies, Sultan Qaboos University, Oman. Benchmark...
Review of public motor imagery and execution datasets in brain-computer interfaces
Review of public motor imagery and execution datasets in brain-computer interfaces
The demand for public datasets has increased as data-driven methodologies have been introduced in the field of brain-computer interfaces (BCIs). Indeed, many BCI datasets are avail...
Hist2Vec: Kernel-Based Embeddings for Biological Sequence Classification
Hist2Vec: Kernel-Based Embeddings for Biological Sequence Classification
AbstractBiological sequence classification is vital in various fields, such as genomics and bioinformatics. The advancement and reduced cost of genomic sequencing have brought the ...
Detectability of an intermediate layer by magnetotelluric sounding
Detectability of an intermediate layer by magnetotelluric sounding
Abstract The recent publication by Verma and Mallick (1979) on the detectability of an intermediate layer by time domain EM sounding provides some informative ans...
GENERator: A Long-Context Generative Genomic Foundation Model
GENERator: A Long-Context Generative Genomic Foundation Model
Abstract The rapid advancement of DNA sequencing has produced vast genomic datasets, yet the interpretation and rational engineering of sequence function remain fun...

Back to Top