Search engine for discovering works of Art, research articles, and books related to Art and Culture
ShareThis
Javascript must be enabled to continue!

A Comparative Study of Join Algorithms in MapReduce

View through CrossRef
To analyze large volumes of data, a set of techniques is present in the IT community as the MapReduce paradigm, parallel RDBMS, column storage, and combinations of those technics. MapReduce is a parallel programming model introduced by Google, which enables an easy parallelization of tasks while hiding the details and the complexity of parallel computations on very large datasets across a large number of machines. Our study will be concerned about the MapReduce Data Analytics, the most important data analysis treatments in MapReduce is the treatment of the Logs files as in the case of web applications, which can be in the form of selection operation, aggregation, or filtering, the most useful and expensive operation is that of the join, the logs files will often need to be joined with one or more table references, however the MapReduce paradigm is not designed to process multiple inputs. While processing relational data is a common need, this limitation causes difficulties and inefficiency when MapReduce is applied on relational operations like joins. The aim of this paper is to compare a number of wellknown join strategies in MapReduce, analyze the costs associated to a MapReduce program in terms of I/O and CPU used, and present some techniques of optimizations from related works.
Title: A Comparative Study of Join Algorithms in MapReduce
Description:
To analyze large volumes of data, a set of techniques is present in the IT community as the MapReduce paradigm, parallel RDBMS, column storage, and combinations of those technics.
MapReduce is a parallel programming model introduced by Google, which enables an easy parallelization of tasks while hiding the details and the complexity of parallel computations on very large datasets across a large number of machines.
Our study will be concerned about the MapReduce Data Analytics, the most important data analysis treatments in MapReduce is the treatment of the Logs files as in the case of web applications, which can be in the form of selection operation, aggregation, or filtering, the most useful and expensive operation is that of the join, the logs files will often need to be joined with one or more table references, however the MapReduce paradigm is not designed to process multiple inputs.
While processing relational data is a common need, this limitation causes difficulties and inefficiency when MapReduce is applied on relational operations like joins.
The aim of this paper is to compare a number of wellknown join strategies in MapReduce, analyze the costs associated to a MapReduce program in terms of I/O and CPU used, and present some techniques of optimizations from related works.

Related Results

Multi-constraint scheduling of MapReduce workloads
Multi-constraint scheduling of MapReduce workloads
In recent years there has been an extraordinary growth of large-scale data processing and related technologies in both, industry and academic communities. This trend is mostly driv...
Primerjalna književnost na prelomu tisočletja
Primerjalna književnost na prelomu tisočletja
In a comprehensive and at times critical manner, this volume seeks to shed light on the development of events in Western (i.e., European and North American) comparative literature ...
Optimizing data management for MapReduce applications on large-scale distributed infrastructures
Optimizing data management for MapReduce applications on large-scale distributed infrastructures
Optimisation de la gestion des données pour les applications MapReduce sur des infrastructures distribuées à grande échelle Les applications data-intensive sont lar...
Using join.me to help library patrons
Using join.me to help library patrons
PurposeAs the Informatics Librarian at Olivet Nazarene University, my staff and I are often responsible for troubleshooting our patrons' technology issues. My experience with join....
Improving MapReduce Performance on Clusters
Improving MapReduce Performance on Clusters
Amélioration des performances de MapReduce sur grappe de calcul Beaucoup de disciplines scientifiques s'appuient désormais sur l'analyse et la fouille de masses gig...
A Novel Approach to Translate Structural Aggregation Queries to MapReduce Code
A Novel Approach to Translate Structural Aggregation Queries to MapReduce Code
Abstract Data management applications are rapidly growing applications that require more attention, especially in the big data era. Thus, it is critical to support these ap...
TriJoin: A Time-Efficient and Scalable Three-Way Distributed Stream Join System
TriJoin: A Time-Efficient and Scalable Three-Way Distributed Stream Join System
<p>Stream join is one of the most fundamental operations in data stream processing applications. Existing distributed stream join systems can support efficient two-way join, ...
Efficient parallel implementation of the SHRiMP sequence alignment tool using MapReduce
Efficient parallel implementation of the SHRiMP sequence alignment tool using MapReduce
With the advent of ultra high-throughput DNA sequencing technologies used in Next-Generation Sequencing (NGS) machines, we are facing a daunting new era in petabyte scale bioinform...

Back to Top