Javascript must be enabled to continue!
Curvature-Enhanced Leaky AdamW: Adaptive Gradient Optimization Based on Dynamic Loss Surface Correction for Deep Neural Networks
View through CrossRef
Adaptive gradient optimization algorithms represented by AdamW have been widely used in deep learning tasks such as image classification. However, traditional AdamW suffers from fixed update rules, slow convergence speed, severe gradient oscillation in late training stages, insufficient parameter update perception, and poor stability in flat or sharp loss regions, resulting in limited final test accuracy of the model. To solve these problems, this paper proposes a novel adaptive optimization algorithm named Curvature-Enhanced Leaky AdamW (CEL-AdamW). Based on the gradient momentum difference and real-time parameter update variation, the algorithm constructs a dynamic training curvature estimation mechanism, realizing lightweight and fine-grained perception of loss surface geometric features. It designs pixel-level abnormal judgment rules for parameter stagnation and curvature overflow, and replaces invalid curvature features with momentum difference to ensure update effectiveness. A targeted curvature enhancement strategy is innovatively introduced to quantify the geometric characteristics of real-time loss surface, which adaptively corrects the step size scaling factor. Combined with Leaky nonlinear mapping and three-stage training scheduling strategy (dormancy, warm-up, full activation), the proposed algorithm significantly accelerates model convergence speed, suppresses training oscillation, and improves convergence stability and final test accuracy. Comprehensive experiments on CIFAR-10 with ResNet18 show that compared with baseline optimizers including Adam, AdamW, and AdaBelief, the proposed CEL-AdamW achieves faster loss convergence, lower training loss and higher classification accuracy. Notably, CEL-AdamW reaches the optimal test accuracy of 89.34% at approximately the 50th epoch, outperforming Adam (88.17%), AdamW (88.57%), and AdaBelief (88.86%) in both convergence speed and final precision, with the peak accuracy achieved 8 epochs earlier than Adam and AdaBelief. Strict mathematical convergence proof verifies the rationality and superiority of the proposed algorithm. The proposed dynamic curvature enhancement optimization mechanism provides an effective solution for stable, fast and high-precision training of deep neural networks.
Title: Curvature-Enhanced Leaky AdamW: Adaptive Gradient Optimization Based on Dynamic Loss Surface Correction for Deep Neural Networks
Description:
Adaptive gradient optimization algorithms represented by AdamW have been widely used in deep learning tasks such as image classification.
However, traditional AdamW suffers from fixed update rules, slow convergence speed, severe gradient oscillation in late training stages, insufficient parameter update perception, and poor stability in flat or sharp loss regions, resulting in limited final test accuracy of the model.
To solve these problems, this paper proposes a novel adaptive optimization algorithm named Curvature-Enhanced Leaky AdamW (CEL-AdamW).
Based on the gradient momentum difference and real-time parameter update variation, the algorithm constructs a dynamic training curvature estimation mechanism, realizing lightweight and fine-grained perception of loss surface geometric features.
It designs pixel-level abnormal judgment rules for parameter stagnation and curvature overflow, and replaces invalid curvature features with momentum difference to ensure update effectiveness.
A targeted curvature enhancement strategy is innovatively introduced to quantify the geometric characteristics of real-time loss surface, which adaptively corrects the step size scaling factor.
Combined with Leaky nonlinear mapping and three-stage training scheduling strategy (dormancy, warm-up, full activation), the proposed algorithm significantly accelerates model convergence speed, suppresses training oscillation, and improves convergence stability and final test accuracy.
Comprehensive experiments on CIFAR-10 with ResNet18 show that compared with baseline optimizers including Adam, AdamW, and AdaBelief, the proposed CEL-AdamW achieves faster loss convergence, lower training loss and higher classification accuracy.
Notably, CEL-AdamW reaches the optimal test accuracy of 89.
34% at approximately the 50th epoch, outperforming Adam (88.
17%), AdamW (88.
57%), and AdaBelief (88.
86%) in both convergence speed and final precision, with the peak accuracy achieved 8 epochs earlier than Adam and AdaBelief.
Strict mathematical convergence proof verifies the rationality and superiority of the proposed algorithm.
The proposed dynamic curvature enhancement optimization mechanism provides an effective solution for stable, fast and high-precision training of deep neural networks.
Related Results
NEURAL NETWORKS AND DEEP LEARNING: THEORITICAL INSIGHTS AND FRAMEWORKS
NEURAL NETWORKS AND DEEP LEARNING: THEORITICAL INSIGHTS AND FRAMEWORKS
“NEURAL NETWORKS AND DEEP LEARNING: THEORITICAL INSIGHTS AND FRAMEWORKS” is a comprehensive guide that dives deep into the world of neural networks and their applications in modern...
The connections between postural reactions, scoliosis postures and scoliosis in girls aged 12-15 years old examined using the Spearman’s Rank OrderCorrelation
The connections between postural reactions, scoliosis postures and scoliosis in girls aged 12-15 years old examined using the Spearman’s Rank OrderCorrelation
Abstract
The aim of the research was to analyse the Spearman's Rank Order Correlation between the postural reactions, scoliosis postures and scoliosis in girls aged ...
Study on the myopia control effect of OK lens on children with different corneal curvature
Study on the myopia control effect of OK lens on children with different corneal curvature
Abstract
Objective: To explore the effect of OK lens on myopia control in children with different corneal curvature.
Method: A total of 178 myopic children admitted to our...
Na+- and cGMP-induced Ca2+ fluxes in frog rod photoreceptors.
Na+- and cGMP-induced Ca2+ fluxes in frog rod photoreceptors.
We have examined the Ca2+ content and pathways of Ca2+ transport in frog rod outer segments using the Ca2+-indicating dye arsenazo III. The experiments employed suspensions of oute...
Fuzzy Chaotic Neural Networks
Fuzzy Chaotic Neural Networks
An understanding of the human brain’s local function has improved in recent years. But the cognition of human brain’s working process as a whole is still obscure. Both fuzzy logic ...
Over-correction of Curvature in Surgical Segments Cause the Non-surgical Curvature Loss in One- and Two-level Anterior Cervical Discectomy and Fusion
Over-correction of Curvature in Surgical Segments Cause the Non-surgical Curvature Loss in One- and Two-level Anterior Cervical Discectomy and Fusion
Abstract
Background: In some patients with anterior cervical disecetomy anf fusion (ACDF), the correction degree of cervical curvature after surgery is less than that of th...
On the role of network dynamics for information processing in artificial and biological neural networks
On the role of network dynamics for information processing in artificial and biological neural networks
Understanding how interactions in complex systems give rise to various collective behaviours has been of interest for researchers across a wide range of fields. However, despite ma...
ACM SIGCOMM computer communication review
ACM SIGCOMM computer communication review
At some point in the future, how far out we do not exactly know, wireless access to the Internet will outstrip all other forms of access bringing the freedom of mobility to the way...

