Javascript must be enabled to continue!
Expanding the boundaries of local similarity analysis
View through CrossRef
Abstract
Background
Pairwise comparison of time series data for both local and time-lagged relationships is a computationally challenging problem relevant to many fields of inquiry. The Local Similarity Analysis (LSA) statistic identifies the existence of local and lagged relationships, but determining significance through a p-value has been algorithmically cumbersome due to an intensive permutation test, shuffling rows and columns and repeatedly calculating the statistic. Furthermore, this p-value is calculated with the assumption of normality -- a statistical luxury dissociated from most real world datasets.
Results
To improve the performance of LSA on big datasets, an asymptotic upper bound on the p-value calculation was derived without the assumption of normality. This change in the bound calculation markedly improved computational speed from O(pm
2
n) to O(m
2
n), where p is the number of permutations in a permutation test, m is the number of time series, and n is the length of each time series. The bounding process is implemented as a computationally efficient software package, FAST LSA, written in C and optimized for threading on multi-core computers, improving its practical computation time. We computationally compare our approach to previous implementations of LSA, demonstrate broad applicability by analyzing time series data from public health, microbial ecology, and social media, and visualize resulting networks using the Cytoscape software.
Conclusions
The FAST LSA software package expands the boundaries of LSA allowing analysis on datasets with millions of co-varying time series. Mapping metadata onto force-directed graphs derived from FAST LSA allows investigators to view correlated cliques and explore previously unrecognized network relationships. The software is freely available for download at: http://www.cmde.science.ubc.ca/hallam/fastLSA/.
Springer Science and Business Media LLC
Title: Expanding the boundaries of local similarity analysis
Description:
Abstract
Background
Pairwise comparison of time series data for both local and time-lagged relationships is a computationally challenging problem relevant to many fields of inquiry.
The Local Similarity Analysis (LSA) statistic identifies the existence of local and lagged relationships, but determining significance through a p-value has been algorithmically cumbersome due to an intensive permutation test, shuffling rows and columns and repeatedly calculating the statistic.
Furthermore, this p-value is calculated with the assumption of normality -- a statistical luxury dissociated from most real world datasets.
Results
To improve the performance of LSA on big datasets, an asymptotic upper bound on the p-value calculation was derived without the assumption of normality.
This change in the bound calculation markedly improved computational speed from O(pm
2
n) to O(m
2
n), where p is the number of permutations in a permutation test, m is the number of time series, and n is the length of each time series.
The bounding process is implemented as a computationally efficient software package, FAST LSA, written in C and optimized for threading on multi-core computers, improving its practical computation time.
We computationally compare our approach to previous implementations of LSA, demonstrate broad applicability by analyzing time series data from public health, microbial ecology, and social media, and visualize resulting networks using the Cytoscape software.
Conclusions
The FAST LSA software package expands the boundaries of LSA allowing analysis on datasets with millions of co-varying time series.
Mapping metadata onto force-directed graphs derived from FAST LSA allows investigators to view correlated cliques and explore previously unrecognized network relationships.
The software is freely available for download at: http://www.
cmde.
science.
ubc.
ca/hallam/fastLSA/.
Related Results
Frequency of Common Chromosomal Abnormalities in Patients with Idiopathic Acquired Aplastic Anemia
Frequency of Common Chromosomal Abnormalities in Patients with Idiopathic Acquired Aplastic Anemia
Objective: To determine the frequency of common chromosomal aberrations in local population idiopathic determine the frequency of common chromosomal aberrations in local population...
News event
News event
When analyzing news media data with automated content analysis techniques, studies often aggregate their measures at the article level (Nicholls & Bright, 2019). However, many ...
Similarity Search with Data Missing
Similarity Search with Data Missing
Similarity search is a fundamental research problem with broad applications in various research fields, including data mining, information retrieval, and machine learning. The core...
Spectral Jaccard Similarity for long-read alignment
Spectral Jaccard Similarity for long-read alignment
A key step in genomic analysis pipelines is the identification of regions of similarity between pairs of DNA sequencing reads. This task, known as pairwise sequence alignment, is a...
GEOMORPHIC BOUNDARIES WITHIN RIVER NETWORKS
GEOMORPHIC BOUNDARIES WITHIN RIVER NETWORKS
Author contributions: MWS and MCT contributed equally to all aspects of
this research and manuscript preparation. Key Points 1. The physical
character of different functional proce...
Similarity of Sentences With Contradiction Using Semantic Similarity Measures
Similarity of Sentences With Contradiction Using Semantic Similarity Measures
AbstractShort text or sentence similarity is crucial in various natural language processing activities. Traditional measures for sentence similarity consider word order, semantic f...
Spectral Jaccard Similarity: A new approach to estimating pairwise sequence alignments
Spectral Jaccard Similarity: A new approach to estimating pairwise sequence alignments
Abstract
A key step in many genomic analysis pipelines is the identification of regions of similarity between pairs of DNA sequencing reads. This...
Using covariance weighted euclidean distance to assess the dissimilarity between integral experiments
Using covariance weighted euclidean distance to assess the dissimilarity between integral experiments
Integral experiments especially criticality experiments help a lot in designing either new nuclear reactor or criticality assembly. The calculation uncertainty of the integral para...

