Javascript must be enabled to continue!
Context dependency of nucleotide probabilities and variants in human DNA
View through CrossRef
Abstract
Background
Genomic DNA has been shaped by mutational processes through evolution. The cellular machinery for error correction and repair has left its marks in the nucleotide composition along with structural and functional constraints. Therefore, the probability of observing a base in a certain position in the human genome is highly context-dependent.
Results
Here we develop context-dependent nucleotide models. We first investigate models of nucleotides conditioned on sequence context. We develop a bidirectional Markov model that use an average of the probability from a Markov model applied to both strands of the sequence and thus depends on up to 14 bases to each side of the nucleotide. We show how the genome predictability varies across different types of genomic regions. Surprisingly, this model can predict a base from its context with an average of more than 50% accuracy. For somatic variants we show a tendency towards higher probability for the variant base than for the reference base. Inspired by DNA substitution models, we develop a model of mutability that estimates a mutation matrix (called the alpha matrix) on top of the nucleotide distribution. The alpha matrix can be estimated from a much smaller context than the nucleotide model, but the final model will still depend on the full context of the nucleotide model. With the bidirectional Markov model of order 14 and an alpha matrix dependent on just one base to each side, we obtain a model that compares well with a model of mutability that estimates mutation probabilities directly conditioned on three nucleotides to each side. For somatic variants in particular, our model fits better than the simpler model. Interestingly, the model is not very sensitive to the size of the context for the alpha matrix.
Conclusions
Our study found strong context dependencies of nucleotides in the human genome. The best model uses a context of 14 nucleotides to each side. Based on these models, a substitution model was constructed that separates into the context model and a matrix dependent on a small context. The model fit somatic variants particularly well.
Title: Context dependency of nucleotide probabilities and variants in human DNA
Description:
Abstract
Background
Genomic DNA has been shaped by mutational processes through evolution.
The cellular machinery for error correction and repair has left its marks in the nucleotide composition along with structural and functional constraints.
Therefore, the probability of observing a base in a certain position in the human genome is highly context-dependent.
Results
Here we develop context-dependent nucleotide models.
We first investigate models of nucleotides conditioned on sequence context.
We develop a bidirectional Markov model that use an average of the probability from a Markov model applied to both strands of the sequence and thus depends on up to 14 bases to each side of the nucleotide.
We show how the genome predictability varies across different types of genomic regions.
Surprisingly, this model can predict a base from its context with an average of more than 50% accuracy.
For somatic variants we show a tendency towards higher probability for the variant base than for the reference base.
Inspired by DNA substitution models, we develop a model of mutability that estimates a mutation matrix (called the alpha matrix) on top of the nucleotide distribution.
The alpha matrix can be estimated from a much smaller context than the nucleotide model, but the final model will still depend on the full context of the nucleotide model.
With the bidirectional Markov model of order 14 and an alpha matrix dependent on just one base to each side, we obtain a model that compares well with a model of mutability that estimates mutation probabilities directly conditioned on three nucleotides to each side.
For somatic variants in particular, our model fits better than the simpler model.
Interestingly, the model is not very sensitive to the size of the context for the alpha matrix.
Conclusions
Our study found strong context dependencies of nucleotides in the human genome.
The best model uses a context of 14 nucleotides to each side.
Based on these models, a substitution model was constructed that separates into the context model and a matrix dependent on a small context.
The model fit somatic variants particularly well.
Related Results
Genome wide hypomethylation and youth-associated DNA gap reduction promoting DNA damage and senescence-associated pathogenesis
Genome wide hypomethylation and youth-associated DNA gap reduction promoting DNA damage and senescence-associated pathogenesis
Abstract
Background: Age-associated epigenetic alteration is the underlying cause of DNA damage in aging cells. Two types of youth-associated DNA-protection epigenetic mark...
Genome wide hypomethylation and youth-associated DNA gap reduction promoting DNA damage and senescence-associated pathogenesis
Genome wide hypomethylation and youth-associated DNA gap reduction promoting DNA damage and senescence-associated pathogenesis
Introduction: The United States currently faces two opioid crises, an evolved crisis currently manifesting as widespread abuse of illicit opioids, and a crisis in pain management l...
7
th
International Symposium on Enabling Technologies for Life Sciences (ETP)
7
th
International Symposium on Enabling Technologies for Life Sciences (ETP)
The seventh in the series of ETP Symposia (see
Rapid Communications in Mass Spectrometry
2012,
26
, ...
Echinococcus granulosus in Environmental Samples: A Cross-Sectional Molecular Study
Echinococcus granulosus in Environmental Samples: A Cross-Sectional Molecular Study
Abstract
Introduction
Echinococcosis, caused by tapeworms of the Echinococcus genus, remains a significant zoonotic disease globally. The disease is particularly prevalent in areas...
Biophysical studies of RNA:DNA:DNA triplexes and characterization of riboswitches in cell-free transcription-translation systems
Biophysical studies of RNA:DNA:DNA triplexes and characterization of riboswitches in cell-free transcription-translation systems
RNA research is very important since RNA molecules are involved in various gene regulatory mechanisms as well as pathways of cell physiology and disease development.1 RNAs have evo...
DNA replication initiation and fidelity : a nanoscale view of the code of life
DNA replication initiation and fidelity : a nanoscale view of the code of life
<p dir="ltr">Every time a cell divides it needs to copy its entire genome. This is a fragile and challenging task, involving billions of DNA base pairs, tightly bound protein...
DNA replication initiation and fidelity : a nanoscale view of the code of life
DNA replication initiation and fidelity : a nanoscale view of the code of life
<p dir="ltr">Every time a cell divides it needs to copy its entire genome. This is a fragile and challenging task, involving billions of DNA base pairs, tightly bound protein...
DNA replication initiation and fidelity: a nanoscale view of the code of life
DNA replication initiation and fidelity: a nanoscale view of the code of life
<p dir="ltr">Every time a cell divides it needs to copy its entire genome. This is a fragile and challenging task, involving billions of DNA base pairs, tightly bound protein...

