Javascript must be enabled to continue!
Using phylogenetic summary statistics for epidemiological inference
View through CrossRef
Abstract
Since the coining of the term phylodynamics, the use of phylogenies to understand infectious disease dynamics has steadily increased. As methods for phylodynamics and genomic epidemiology have proliferated and grown more computationally expensive, the epidemiological information they extract has also evolved to better complement what can be learned through traditional epidemiological data. However, for genomic epidemiology to continue to grow, and for the accumulating number of pathogen genetic sequences to fulfill their potential widespread utility, the extraction of epidemiological information from phylogenies needs to be simpler and more efficient. Summary statistics provide a straightforward way of extracting information from a phylogenetic tree, but the relationship between these statistics and epidemiological quantities needs to be better understood. In this work we address this need via simulation. Using two different benchmark scenarios, we evaluate 74 tree summary statistics and their relationship to epidemiological quantities. In addition to evaluating the epidemiological information that can be inferred from each summary statistic, we also assess the computational cost of each statistic. This helps us optimize the selection of summary statistics for specific applications. Our study offers guidelines on essential considerations for designing or choosing summary statistics. The evaluated set of summary statistics, along with additional helpful functions for phylogenetic analysis, is accessible through an open-source Python library. Our research not only illuminates the main characteristics of many tree summary statistics but also provides valuable computational tools for real-world epidemiological analyses. These contributions aim to enhance our understanding of disease spread dynamics and advance the broader utilization of genomic epidemiology in public health efforts.
Author Summary
Our study focuses on the use of phylogenetic analysis to get valuable epidemiological insights. We conducted a simulation study to evaluate 74 phylogenetic summary statistics and their relationship to epidemiological quantities, shedding light on the potential of each of these statistics to quantify different characteristics of disease spread dynamics. Additionally, we assessed the computational cost of each statistic. This gives us additional information when selecting a statistic for a particular application. Our research is available through an open-source Python library. This work helps us enhance our understanding of phylogenetic tree structures and contributes to the broader application of genomic epidemiology in public health initiatives.
Title: Using phylogenetic summary statistics for epidemiological inference
Description:
Abstract
Since the coining of the term phylodynamics, the use of phylogenies to understand infectious disease dynamics has steadily increased.
As methods for phylodynamics and genomic epidemiology have proliferated and grown more computationally expensive, the epidemiological information they extract has also evolved to better complement what can be learned through traditional epidemiological data.
However, for genomic epidemiology to continue to grow, and for the accumulating number of pathogen genetic sequences to fulfill their potential widespread utility, the extraction of epidemiological information from phylogenies needs to be simpler and more efficient.
Summary statistics provide a straightforward way of extracting information from a phylogenetic tree, but the relationship between these statistics and epidemiological quantities needs to be better understood.
In this work we address this need via simulation.
Using two different benchmark scenarios, we evaluate 74 tree summary statistics and their relationship to epidemiological quantities.
In addition to evaluating the epidemiological information that can be inferred from each summary statistic, we also assess the computational cost of each statistic.
This helps us optimize the selection of summary statistics for specific applications.
Our study offers guidelines on essential considerations for designing or choosing summary statistics.
The evaluated set of summary statistics, along with additional helpful functions for phylogenetic analysis, is accessible through an open-source Python library.
Our research not only illuminates the main characteristics of many tree summary statistics but also provides valuable computational tools for real-world epidemiological analyses.
These contributions aim to enhance our understanding of disease spread dynamics and advance the broader utilization of genomic epidemiology in public health efforts.
Author Summary
Our study focuses on the use of phylogenetic analysis to get valuable epidemiological insights.
We conducted a simulation study to evaluate 74 phylogenetic summary statistics and their relationship to epidemiological quantities, shedding light on the potential of each of these statistics to quantify different characteristics of disease spread dynamics.
Additionally, we assessed the computational cost of each statistic.
This gives us additional information when selecting a statistic for a particular application.
Our research is available through an open-source Python library.
This work helps us enhance our understanding of phylogenetic tree structures and contributes to the broader application of genomic epidemiology in public health initiatives.
Related Results
Predictors of Statistics Anxiety Among Graduate Students in Saudi Arabia
Predictors of Statistics Anxiety Among Graduate Students in Saudi Arabia
Problem The problem addressed in this study is the anxiety experienced by graduate students toward statistics courses, which often causes students to delay taking statistics cours...
Latency-Critical Inference Serving for Deep Learning
Latency-Critical Inference Serving for Deep Learning
Deep learning (DL) technology has made remarkable strides in terms of accuracy through the advancement of sophisticated and large deep neural networks (DNNs). Yet, its adoption in ...
Empirical Performance of Tree-based Inference of Phylogenetic Networks
Empirical Performance of Tree-based Inference of Phylogenetic Networks
Abstract
Phylogenetic networks extend the phylogenetic tree structure and allow for modeling vertical and horizontal evolution in a single framework. Statistical in...
Species of Fusarium and Neocosmospora associated with citrus branch diseases in China
Species of Fusarium and Neocosmospora associated with citrus branch diseases in China
Fig. S1. Phylogenetic tree generated by Bayesian inference analyses based on the individual CaM, rpb1, rpb2 and tef1 (A–D) for species in Fusarium fujikuroi species complex (FFSC)....
PaNDA: Efficient Optimization of Phylogenetic Diversity in Networks
PaNDA: Efficient Optimization of Phylogenetic Diversity in Networks
Abstract
Phylogenetic diversity plays an important role in biodiversity, conservation, and evolutionary studies by measuring the diversity of a s...
Phylogenetic overdispersion of plant species in southern Brazilian savannas
Phylogenetic overdispersion of plant species in southern Brazilian savannas
Ecological communities are the result of not only present ecological processes, such as competition among species and environmental filtering, but also past and continuing evolutio...
A world without statistics?
A world without statistics?
s a practicing statistician, we frequently are asked questions like: What is the role of statistics in our daily life? Why do we need statistics? What would the world be without st...
Do evidence summaries increase health policy‐makers' use of evidence from systematic reviews? A systematic review
Do evidence summaries increase health policy‐makers' use of evidence from systematic reviews? A systematic review
This review summarizes the evidence from six randomized controlled trials that judged the effectiveness of systematic review summaries on policymakers' decision making, or the most...

