Javascript must be enabled to continue!
Comparison of fast regression algorithms in large datasets
View through CrossRef
The aim is to compare the performances of fast regression methods, namely dimensional reduction of correlation matrix (DRCM), nonparametric dimensional reduction of correlation matrix (N-DRCM), variance inflation factor (VIF) regression, and robust VIF (R-VIF) regression in the presence of multicollinearity and outliers problems. In all simulation-scenarios, all the target variables were chosen for final models using four methods. The DRCM and N-DRCM are the methods that reach the final model in the shortest time, respectively. The time to reach the final model using R-VIF regression was approximately twice shorter than that of VIF regression. In each method, as the number of variables and the level of outliers increased, the time taken to reach the final model increased. When the level of multicollinearity and the number of variables (p > 500) increased, the times to reach the final models using DRCM in datasets with outliers were slightly shorter than the those of N-DRCM. The largest numbers of noise variables were selected to the model using DRCM and N-DRCM, but the least number of them were selected to the model using the R-VIF regression. The RMSE values obtained using DRCM, N-DRCM and VIF regression were similar in each scenario. As a result of the real dataset, the final model selected using R-VIF regression had the highest R2. It also had the lowest RMSE value among those obtained with other approaches excluding VIF regression. As such, the R-VIF regression method demonstrated a better performance than the others in all datasets.
Title: Comparison of fast regression algorithms in large datasets
Description:
The aim is to compare the performances of fast regression methods, namely dimensional reduction of correlation matrix (DRCM), nonparametric dimensional reduction of correlation matrix (N-DRCM), variance inflation factor (VIF) regression, and robust VIF (R-VIF) regression in the presence of multicollinearity and outliers problems.
In all simulation-scenarios, all the target variables were chosen for final models using four methods.
The DRCM and N-DRCM are the methods that reach the final model in the shortest time, respectively.
The time to reach the final model using R-VIF regression was approximately twice shorter than that of VIF regression.
In each method, as the number of variables and the level of outliers increased, the time taken to reach the final model increased.
When the level of multicollinearity and the number of variables (p > 500) increased, the times to reach the final models using DRCM in datasets with outliers were slightly shorter than the those of N-DRCM.
The largest numbers of noise variables were selected to the model using DRCM and N-DRCM, but the least number of them were selected to the model using the R-VIF regression.
The RMSE values obtained using DRCM, N-DRCM and VIF regression were similar in each scenario.
As a result of the real dataset, the final model selected using R-VIF regression had the highest R2.
It also had the lowest RMSE value among those obtained with other approaches excluding VIF regression.
As such, the R-VIF regression method demonstrated a better performance than the others in all datasets.
Related Results
Selection of Injectable Drug Product Composition using Machine Learning Models (Preprint)
Selection of Injectable Drug Product Composition using Machine Learning Models (Preprint)
BACKGROUND
As of July 2020, a Web of Science search of “machine learning (ML)” nested within the search of “pharmacokinetics or pharmacodynamics” yielded over 100...
Assessment of Chlorophyll-a Algorithms Considering Different Trophic Statuses and Optimal Bands
Assessment of Chlorophyll-a Algorithms Considering Different Trophic Statuses and Optimal Bands
Numerous algorithms have been proposed to retrieve chlorophyll-a concentrations in Case 2 waters; however, the retrieval accuracy is far from satisfactory. In this research, seven ...
Comparative Analysis of Classical and Quantum Machine Learning Algorithms in Breast Cancer Classification
Comparative Analysis of Classical and Quantum Machine Learning Algorithms in Breast Cancer Classification
Abstract
This study presents a comparison between classical machine learning (ML) algorithms and their quantum-enhanced counterparts in classifying scikit’s breast ...
Food Sales Prediction Using MLP, RANSAC, and Bagging
Food Sales Prediction Using MLP, RANSAC, and Bagging
Many datasets about food sales, these datasets contain different features depending on the data present. Also, the way these features are correlated differs from one dataset to ano...
Review of public motor imagery and execution datasets in brain-computer interfaces
Review of public motor imagery and execution datasets in brain-computer interfaces
The demand for public datasets has increased as data-driven methodologies have been introduced in the field of brain-computer interfaces (BCIs). Indeed, many BCI datasets are avail...
Large-Scale Kernel Machines
Large-Scale Kernel Machines
Solutions for learning from large scale datasets, including kernel learning algorithms that scale linearly with the volume of the data and experiments carried out on realistically ...
Integrating quantum neural networks with machine learning algorithms for optimizing healthcare diagnostics and treatment outcomes
Integrating quantum neural networks with machine learning algorithms for optimizing healthcare diagnostics and treatment outcomes
The rapid advancements in artificial intelligence (AI) and quantum computing have catalyzed an unprecedented shift in the methodologies utilized for healthcare diagnostics and trea...
EVOLVING CONSUMER PREFERENCES ON THE FAST FASHION MARKET
EVOLVING CONSUMER PREFERENCES ON THE FAST FASHION MARKET
Eter Kharaishvili
E-mail: eter.kharaishvili@tsu.ge
Professor, Ivane Javakhishvili Tbilisi State University
Tbilisi, Georgia
https://orcid.org/0000-0003-4013-7354
Nino Lobzh...

