Javascript must be enabled to continue!
Predicting Genome-Wide Approximate Match Frequencies with Hit Frequency Vectors
View through CrossRef
Abstract
Background: Estimating how many approximate matches a query sequence has in a genome is a fundamental task in bioinformatics, with applications including read mapping or genome mappability analysis. While exact string algorithms and specialized index data structures enable low-error matching, counting matches for short k-mers at high dissimilarity (small k/e ratio) becomes computationally intractable as the number of potential matches grow exponentially for Hamming distance, and even more so for other dissimilarity measures including edit distance.
Results: We use machine learning to solve this algorithmic problem. Specifically, we propose deep learning models which in contrast to learned index approaches do not accelerate exact algorithms, but rather directly predict Hit Frequency Vectors (HFVs), which count approximate matches at increasing distance e = 1, 2, 3, . . .. The models are trained on HFVs which are computed using brute-force (for Hamming distance), respectively an efficient sampling approach followed by a search to account for undersampling close matches and correcting counts. Model training and inference are computationally efficient, and enable rapid genome-scale estimation of approximate match frequencies for short query sequences and small k/e. The models capture genome-wide match frequency distributions with high accuracy and allow fast batch prediction even in regimes where exhaustive computation is computationally infeasible for exact algorithms relying on seed-and-extend due to limited seed length leading to exponential effort. We provide loss functions and evaluation measures appropriate for histograms with bin count varying by several orders of magnitude.
Conclusion: Our approach enables practical approximation of HFVs and can be easily extended to complex dissimilarities which are too computationally expensive to evaluate exhaustively across an entire genome. The combination of approximate HFV generation and predictive modeling extends the scope of sequence uniqueness analyses considerably. Our approach can trivially be extended to continuous dissimilarities like oligonucleotide binding affinity via appropriate binning for applications in DNA storage, target depletion, or oligonucleotide therapeutics. Lastly, we believe that our work can provide a blue-print for further applications in bioinformatics where straight-forward ML prediction—possibly followed by confirmation with exact algorithms—could have great utility.
Title: Predicting Genome-Wide Approximate Match Frequencies with Hit Frequency Vectors
Description:
Abstract
Background: Estimating how many approximate matches a query sequence has in a genome is a fundamental task in bioinformatics, with applications including read mapping or genome mappability analysis.
While exact string algorithms and specialized index data structures enable low-error matching, counting matches for short k-mers at high dissimilarity (small k/e ratio) becomes computationally intractable as the number of potential matches grow exponentially for Hamming distance, and even more so for other dissimilarity measures including edit distance.
Results: We use machine learning to solve this algorithmic problem.
Specifically, we propose deep learning models which in contrast to learned index approaches do not accelerate exact algorithms, but rather directly predict Hit Frequency Vectors (HFVs), which count approximate matches at increasing distance e = 1, 2, 3, .
.
.
The models are trained on HFVs which are computed using brute-force (for Hamming distance), respectively an efficient sampling approach followed by a search to account for undersampling close matches and correcting counts.
Model training and inference are computationally efficient, and enable rapid genome-scale estimation of approximate match frequencies for short query sequences and small k/e.
The models capture genome-wide match frequency distributions with high accuracy and allow fast batch prediction even in regimes where exhaustive computation is computationally infeasible for exact algorithms relying on seed-and-extend due to limited seed length leading to exponential effort.
We provide loss functions and evaluation measures appropriate for histograms with bin count varying by several orders of magnitude.
Conclusion: Our approach enables practical approximation of HFVs and can be easily extended to complex dissimilarities which are too computationally expensive to evaluate exhaustively across an entire genome.
The combination of approximate HFV generation and predictive modeling extends the scope of sequence uniqueness analyses considerably.
Our approach can trivially be extended to continuous dissimilarities like oligonucleotide binding affinity via appropriate binning for applications in DNA storage, target depletion, or oligonucleotide therapeutics.
Lastly, we believe that our work can provide a blue-print for further applications in bioinformatics where straight-forward ML prediction—possibly followed by confirmation with exact algorithms—could have great utility.
Related Results
Frequency of Common Chromosomal Abnormalities in Patients with Idiopathic Acquired Aplastic Anemia
Frequency of Common Chromosomal Abnormalities in Patients with Idiopathic Acquired Aplastic Anemia
Objective: To determine the frequency of common chromosomal aberrations in local population idiopathic determine the frequency of common chromosomal aberrations in local population...
Incidence of Thrombocytopenia and Heparin Induced Thrombocytopenia in Patients Receiving Extracorporeal Membrane Oxygenation (ECMO) Compared to Cardiopulmonary Bypass and the Limited Sensitivity of Pre-Test Probability Score
Incidence of Thrombocytopenia and Heparin Induced Thrombocytopenia in Patients Receiving Extracorporeal Membrane Oxygenation (ECMO) Compared to Cardiopulmonary Bypass and the Limited Sensitivity of Pre-Test Probability Score
Abstract
Introduction
Extracorporeal membrane oxygenation (ECMO) is a life saving measure for severe respiratory (veno-venous ECMO [VV-ECMO]) or cardi...
Identifying the Care Gaps in Heparin-Induced Thrombocytopenia Testing Around the COVID-19 Pandemic
Identifying the Care Gaps in Heparin-Induced Thrombocytopenia Testing Around the COVID-19 Pandemic
Introduction:
Heparin induced-thrombocytopenia (HIT) is a severe autoimmune reaction to heparin that increases patients' risk of developing venous thrombosis, lea...
Blood Cross Matching Without Anti-Human Globulin (AHG) and Bovine Serum: A New Interest for an Old Idea
Blood Cross Matching Without Anti-Human Globulin (AHG) and Bovine Serum: A New Interest for an Old Idea
Abstract
Introduction
Transfusion medicine promotes the safety of blood transfusions by rigorously testing to eliminate risks of infection and hemolytic. The efficacy (to correct ...
Targeted Modifications in Adeno-Associated Virus Serotype (AAV)- 8 Capsid Improves Its Hepatic Gene Transfer Efficiency in Vivo
Targeted Modifications in Adeno-Associated Virus Serotype (AAV)- 8 Capsid Improves Its Hepatic Gene Transfer Efficiency in Vivo
Abstract
Abstract 2045
Recombinant adeno-associated virus vectors based on serotype (AAV)-8 have shown significant promise for liver directed gene the...
[RETRACTED] Keanu Reeves CBD Gummies v1
[RETRACTED] Keanu Reeves CBD Gummies v1
[RETRACTED]Keanu Reeves CBD Gummies ==❱❱ Huge Discounts:[HURRY UP ] Absolute Keanu Reeves CBD Gummies (Available)Order Online Only!! ❰❰= https://www.facebook.com/Keanu-Reeves-CBD-G...
Parameterized Strings: Algorithms and Applications
Parameterized Strings: Algorithms and Applications
The parameterized string (p-string), a generalization of the traditional string, is composed of constant and parameter symbols. A parameterized match (p-match) exists between two p...
Correlation of PF4 ELISA Results and Clinical Likelihood of HIT: Can We Improve on Specificity?.
Correlation of PF4 ELISA Results and Clinical Likelihood of HIT: Can We Improve on Specificity?.
Abstract
The diagnosis of heparin-induced thrombocytopenia (HIT) is made clinically with the support from the laboratory. While the Serotonin Release Assay (SRA) is ...

