Search engine for discovering works of Art, research articles, and books related to Art and Culture
ShareThis
Javascript must be enabled to continue!

Enhancing ConvNeXt for efficient small-size image classification

View through CrossRef
Abstract Vision Transformers (ViTs) have achieved remarkable success in computer vision, particularly with the advent of the Swin Transformer (Swin-T). Recently, ConvNeXt was proposed to revisit the architecture of convolutional neural networks by incorporating techniques from Swin-T, which achieves competitive performance. However, ConvNeXt exhibited relatively lower accuracy and efficiency on small-size datasets with low-resolution images. To address this issue, based on the structure of ConvNeXt, we propose to employ small kernel sizes to better capture features from low-resolution images. A symmetrical inverted bottleneck structure inspired by MobileNet is used to refine the traditional residual block design. We combine the batch normalization and layer normalization to strengthen the relationships between different samples. We use a twice-patch embedding methodology to dynamically generate adaptive patch sizes according to different input image sizes. Subsequently, we integrate the global response normalization to amplify the contrast and specificity of multiple channels for learning diverse features. Finally, we introduce a novel spatial coordinate attention mechanism that effectively captures global feature information. The proposed method demonstrates superior performance on small-size datasets of CIFAR-10, CIFAR-100, Tiny-ImageNet and Fashion-MNIST with fewer parameters, which confirms the effectiveness and efficiency of our approach.
Springer Science and Business Media LLC
Title: Enhancing ConvNeXt for efficient small-size image classification
Description:
Abstract Vision Transformers (ViTs) have achieved remarkable success in computer vision, particularly with the advent of the Swin Transformer (Swin-T).
Recently, ConvNeXt was proposed to revisit the architecture of convolutional neural networks by incorporating techniques from Swin-T, which achieves competitive performance.
However, ConvNeXt exhibited relatively lower accuracy and efficiency on small-size datasets with low-resolution images.
To address this issue, based on the structure of ConvNeXt, we propose to employ small kernel sizes to better capture features from low-resolution images.
A symmetrical inverted bottleneck structure inspired by MobileNet is used to refine the traditional residual block design.
We combine the batch normalization and layer normalization to strengthen the relationships between different samples.
We use a twice-patch embedding methodology to dynamically generate adaptive patch sizes according to different input image sizes.
Subsequently, we integrate the global response normalization to amplify the contrast and specificity of multiple channels for learning diverse features.
Finally, we introduce a novel spatial coordinate attention mechanism that effectively captures global feature information.
The proposed method demonstrates superior performance on small-size datasets of CIFAR-10, CIFAR-100, Tiny-ImageNet and Fashion-MNIST with fewer parameters, which confirms the effectiveness and efficiency of our approach.

Related Results

On Flores Island, do "ape-men" still exist? https://www.sapiens.org/biology/flores-island-ape-men/
On Flores Island, do "ape-men" still exist? https://www.sapiens.org/biology/flores-island-ape-men/
<span style="font-size:11pt"><span style="background:#f9f9f4"><span style="line-height:normal"><span style="font-family:Calibri,sans-serif"><b><spa...
Hubungan Perilaku Pola Makan dengan Kejadian Anak Obesitas
Hubungan Perilaku Pola Makan dengan Kejadian Anak Obesitas
<p><em><span style="font-size: 11.0pt; font-family: 'Times New Roman',serif; mso-fareast-font-family: 'Times New Roman'; mso-ansi-language: EN-US; mso-fareast-langua...
A Comprehensive Review of ConvNeXt Architecture in Image Classification: Performance, Applications, and Prospects
A Comprehensive Review of ConvNeXt Architecture in Image Classification: Performance, Applications, and Prospects
The convergence of ConvNeXt architecture and classification tasks highlights an auspicious direction in modern computer vision. This review systematically analyzes how ConvNeXt tra...
Program Auto Parkir Untuk Analisis Parkir di Goa Gong
Program Auto Parkir Untuk Analisis Parkir di Goa Gong
<p class="western" style="line-height: 100%;" lang="en-AU" align="justify"><span style="font-family: Times New Roman,serif;"><span style="font-size: medium;"><...
Even Star Decomposition of Complete Bipartite Graphs
Even Star Decomposition of Complete Bipartite Graphs
<p><span lang="EN-US"><span style="font-family: 宋体; font-size: medium;">A decomposition (</span><span><span style="font-family: 宋体; font-size: medi...
Deep Learning in Dermatology: Exploring Convnext Model Hierarchies and Ensembles for Enhanced Diagnostic Precision
Deep Learning in Dermatology: Exploring Convnext Model Hierarchies and Ensembles for Enhanced Diagnostic Precision
Skin cancer is one of the most widespread and life-threatening malignancies in the world that requires early and proper accurate ‎diagnosis to be efficiently solved. This research ...
An Empirical Research on Factors Influencing Purchase Intention of Mobile Shopping
An Empirical Research on Factors Influencing Purchase Intention of Mobile Shopping
<p><span style="font-size: xx-small;">Based on the theory of Flow experience, this thesis combines the theory of perceived value with the theory of customer innovation,...

Back to Top