Search engine for discovering works of Art, research articles, and books related to Art and Culture
ShareThis
Javascript must be enabled to continue!

Unlocking the power of optimized data balancing ratios: a new frontier in tackling imbalanced datasets

View through CrossRef
Abstract Data balancing methods eliminate the problem of imbalanced class distributions, which often lead to the majority class being well-learned while the minority class remains underrepresented, negatively affecting classification performance. This study applies data balancing to the healthcare domain, a critical field where classification success directly impacts human life. The primary aim is to introduce novel balancing methods while addressing the previously overlooked problem of optimizing data balancing ratios. Six healthcare datasets were used: Wisconsin Diagnostic Breast Cancer (WDBC), Wisconsin Prognostic Breast Cancer (WPBC), Z-Alizadeh Sani, Kidney, Diabetes, and Stroke, all characterized by significant diseases and imbalanced class distributions. Six balancing methods were tested, including synthetic minority oversampling technique (SMOTE), adaptive synthetic sampling (ADASYN), support vector machine-SMOTE (SVM-SMOTE), Borderline-SMOTE, cubic interpolation, and quadratic interpolation, with interpolation-based methods being adapted to this domain for the first time. The critical factor in data balancing is identifying the optimal ratio that maximizes classification performance. In this study, particle swarm optimization (PSO), whale optimization algorithm (WOA), and Optuna optimization methods were used to optimize balancing ratios via a custom-designed fitness function that simultaneously optimizes classification accuracy and resource consumption. Classification was conducted for three scenarios: full balance, optimized balance, and imbalance, using support vector machine (SVM), random forest (RF), and ensemble learning (EL) classifiers, allowing for extensive analysis. Each combination of balancing methods, classifiers, and optimization techniques was separately analyzed using metrics such as accuracy, precision, recall, F1-score, time, central processing unit (CPU) usage, and memory usage. As a result, the combination that optimally balances classification accuracy and resource consumption was determined for each dataset, providing both comprehensive analysis and insights into the impact of balancing ratio optimization on diagnostic success in health care.
Springer Science and Business Media LLC
Title: Unlocking the power of optimized data balancing ratios: a new frontier in tackling imbalanced datasets
Description:
Abstract Data balancing methods eliminate the problem of imbalanced class distributions, which often lead to the majority class being well-learned while the minority class remains underrepresented, negatively affecting classification performance.
This study applies data balancing to the healthcare domain, a critical field where classification success directly impacts human life.
The primary aim is to introduce novel balancing methods while addressing the previously overlooked problem of optimizing data balancing ratios.
Six healthcare datasets were used: Wisconsin Diagnostic Breast Cancer (WDBC), Wisconsin Prognostic Breast Cancer (WPBC), Z-Alizadeh Sani, Kidney, Diabetes, and Stroke, all characterized by significant diseases and imbalanced class distributions.
Six balancing methods were tested, including synthetic minority oversampling technique (SMOTE), adaptive synthetic sampling (ADASYN), support vector machine-SMOTE (SVM-SMOTE), Borderline-SMOTE, cubic interpolation, and quadratic interpolation, with interpolation-based methods being adapted to this domain for the first time.
The critical factor in data balancing is identifying the optimal ratio that maximizes classification performance.
In this study, particle swarm optimization (PSO), whale optimization algorithm (WOA), and Optuna optimization methods were used to optimize balancing ratios via a custom-designed fitness function that simultaneously optimizes classification accuracy and resource consumption.
Classification was conducted for three scenarios: full balance, optimized balance, and imbalance, using support vector machine (SVM), random forest (RF), and ensemble learning (EL) classifiers, allowing for extensive analysis.
Each combination of balancing methods, classifiers, and optimization techniques was separately analyzed using metrics such as accuracy, precision, recall, F1-score, time, central processing unit (CPU) usage, and memory usage.
As a result, the combination that optimally balances classification accuracy and resource consumption was determined for each dataset, providing both comprehensive analysis and insights into the impact of balancing ratio optimization on diagnostic success in health care.

Related Results

Modeling active cell balancing of lithium-ion bat-teries in MATLAB/Simulink
Modeling active cell balancing of lithium-ion bat-teries in MATLAB/Simulink
Problem. The article is devoted to the study of active balancing of lithium-ion battery cells. Active balancing of lithium-ion battery cells is crucial for ensuring high efficiency...
Zhong-Yong as dynamic balancing between Yin-Yang opposites
Zhong-Yong as dynamic balancing between Yin-Yang opposites
Purpose The purpose of this paper is to comment on Peter Ping Li’s understanding of Zhong-Yong balancing, presented in his article titled “Global implications of the indigenous epi...
Machine Learning Algorithms for Health Care Data Analytics Handling Imbalanced Datasets
Machine Learning Algorithms for Health Care Data Analytics Handling Imbalanced Datasets
In Machine Learning, classification is considered a supervised learning technique to predict class samples based on labeled data. Classification techniques have been applied to var...
Appendix A: Frontier Conflicts
Appendix A: Frontier Conflicts
Appendix A provides a conventional, numbered list of the major episodes of confrontation, or "Frontier Wars," that occurred on the eastern frontier of the Cape Colony between 1781 ...
Advanced Re-Sampling Techniques for Multi-Class Imbalanced Classification
Advanced Re-Sampling Techniques for Multi-Class Imbalanced Classification
Imbalanced classification is a common problem in machine learning, where one class significantly outnumbers the others. This imbalance leads to biased model performance, where the ...
Comparative Analysis of Active and Passive Cell Balancing Strategies in Battery Management Systems
Comparative Analysis of Active and Passive Cell Balancing Strategies in Battery Management Systems
Battery management systems (BMS) play a crucial role in ensuring the performance, reliability, and longevity of modern battery systems by employing cell balancing techniques. This ...
POWER-EFFICIENT VLSI DESIGN: STRATEGIES FOR LOW-POWER APPLICATIONS
POWER-EFFICIENT VLSI DESIGN: STRATEGIES FOR LOW-POWER APPLICATIONS
“Power-Efficient VLSI Design: Strategies for Low-Power Applications” is a comprehensive guide that explores the intricacies of designing energy-efficient integrated circuits, addre...
Optimasi Data Tidak Seimbang pada Interaksi Drug Target dengan Sampling dan Ensemble Support Vector Machine
Optimasi Data Tidak Seimbang pada Interaksi Drug Target dengan Sampling dan Ensemble Support Vector Machine
<p>Data tidak seimbang menjadi salah satu masalah yang muncul pada masalah prediksi atau klasifikasi. Penelitian ini memfokuskan untuk mengatasi masalah data tidak seimbang p...

Back to Top