Search engine for discovering works of Art, research articles, and books related to Art and Culture
ShareThis
Javascript must be enabled to continue!

Norm-Hierarchy Transitions In Representation Learning: When And Why Neural Networks Abandon Shortcuts

View through CrossRef
Neural networks often rely on spurious shortcuts for hundreds of epochs before discovering structured representations. Yet the mechanism governing when this transition occurs-and whether its timing can be predicted-remains poorly understood. While prior work has established that gradient descent converges to low-norm solutions [?] and that networks exhibit simplicity bias [?], neither line of work characterises the timescale of the transition from simple to structured features. We propose a unifying framework-the Norm-Hierarchy Transition-which explains delayed representation learning as the slow traversal of a norm hierarchy under regularised optimisation. When multiple interpolating solutions exist with different norms, weight decay induces a slow contraction from high-norm shortcut solutions toward lower-norm structured representations. We prove a tight bound on the transition delay: T = Θ(γ-1 eff log(V sc /V st)), where V sc and V st are the characteristic norms of the shortcut and structured representations. The framework predicts three regimes as a function of regularisation strength: weak regularisation (shortcuts persist), intermediate regularisation (delayed transition), and strong regularisation (learning suppressed). We validate these predictions across four domains: modular arithmetic (where all six predictions hold with R 2 > 0.97), CIFAR-10 with spurious features (five of six, including 78% → 10% clean accuracy as shortcut strength increases), CelebA (which occupies an intermediate position on the norm separation spectrum), and Waterbirds (where norm dynamics transfer but representational transition does not, confirming the framework's boundary). The norm-hierarchy mechanism is robust across architectures: ResNet18 with standard batch normalisation exhibits the same peak-then-decay norm dynamics as models without normalisation, achieving 78% clean accuracy. The single prediction that fails to transfer-the precise delay scaling T ∝ 1/λ-is explained by a new condition we term clean norm separation, the first formal criterion distinguishing settings where implicit bias timescales are predictable from those where they are not. The framework further predicts that emergent abilities in large language models arise when model scale reduces the norm gap below a training-budget threshold, connecting scaling laws to the same norm-hierarchy mechanism. Our results suggest that grokking, shortcut learning, delayed feature discovery, and emergent abilities are manifestations of a single mechanism: the slow traversal of a norm hierarchy under regularised optimisation.
Title: Norm-Hierarchy Transitions In Representation Learning: When And Why Neural Networks Abandon Shortcuts
Description:
Neural networks often rely on spurious shortcuts for hundreds of epochs before discovering structured representations.
Yet the mechanism governing when this transition occurs-and whether its timing can be predicted-remains poorly understood.
While prior work has established that gradient descent converges to low-norm solutions [?] and that networks exhibit simplicity bias [?], neither line of work characterises the timescale of the transition from simple to structured features.
We propose a unifying framework-the Norm-Hierarchy Transition-which explains delayed representation learning as the slow traversal of a norm hierarchy under regularised optimisation.
When multiple interpolating solutions exist with different norms, weight decay induces a slow contraction from high-norm shortcut solutions toward lower-norm structured representations.
We prove a tight bound on the transition delay: T = Θ(γ-1 eff log(V sc /V st)), where V sc and V st are the characteristic norms of the shortcut and structured representations.
The framework predicts three regimes as a function of regularisation strength: weak regularisation (shortcuts persist), intermediate regularisation (delayed transition), and strong regularisation (learning suppressed).
We validate these predictions across four domains: modular arithmetic (where all six predictions hold with R 2 > 0.
97), CIFAR-10 with spurious features (five of six, including 78% → 10% clean accuracy as shortcut strength increases), CelebA (which occupies an intermediate position on the norm separation spectrum), and Waterbirds (where norm dynamics transfer but representational transition does not, confirming the framework's boundary).
The norm-hierarchy mechanism is robust across architectures: ResNet18 with standard batch normalisation exhibits the same peak-then-decay norm dynamics as models without normalisation, achieving 78% clean accuracy.
The single prediction that fails to transfer-the precise delay scaling T ∝ 1/λ-is explained by a new condition we term clean norm separation, the first formal criterion distinguishing settings where implicit bias timescales are predictable from those where they are not.
The framework further predicts that emergent abilities in large language models arise when model scale reduces the norm gap below a training-budget threshold, connecting scaling laws to the same norm-hierarchy mechanism.
Our results suggest that grokking, shortcut learning, delayed feature discovery, and emergent abilities are manifestations of a single mechanism: the slow traversal of a norm hierarchy under regularised optimisation.

Related Results

NEURAL NETWORKS AND DEEP LEARNING: THEORITICAL INSIGHTS AND FRAMEWORKS
NEURAL NETWORKS AND DEEP LEARNING: THEORITICAL INSIGHTS AND FRAMEWORKS
“NEURAL NETWORKS AND DEEP LEARNING: THEORITICAL INSIGHTS AND FRAMEWORKS” is a comprehensive guide that dives deep into the world of neural networks and their applications in modern...
CREATING LEARNING MEDIA IN TEACHING ENGLISH AT SMP MUHAMMADIYAH 2 PAGELARAN ACADEMIC YEAR 2020/2021
CREATING LEARNING MEDIA IN TEACHING ENGLISH AT SMP MUHAMMADIYAH 2 PAGELARAN ACADEMIC YEAR 2020/2021
The pandemic Covid-19 currently demands teachers to be able to use technology in teaching and learning process. But in reality there are still many teachers who have not been able ...
On the role of network dynamics for information processing in artificial and biological neural networks
On the role of network dynamics for information processing in artificial and biological neural networks
Understanding how interactions in complex systems give rise to various collective behaviours has been of interest for researchers across a wide range of fields. However, despite ma...
Fuzzy Chaotic Neural Networks
Fuzzy Chaotic Neural Networks
An understanding of the human brain’s local function has improved in recent years. But the cognition of human brain’s working process as a whole is still obscure. Both fuzzy logic ...
Net-Zero: A New Norm Analysis
Net-Zero: A New Norm Analysis
<p><strong>Despite its relative obscurity five years ago, four out of every five people on the planet now live under a Net-Zero target. Undoubtedly, Net-Zero has had a ...
Selection of Injectable Drug Product Composition using Machine Learning Models (Preprint)
Selection of Injectable Drug Product Composition using Machine Learning Models (Preprint)
BACKGROUND As of July 2020, a Web of Science search of “machine learning (ML)” nested within the search of “pharmacokinetics or pharmacodynamics” yielded over 100...
ACM SIGCOMM computer communication review
ACM SIGCOMM computer communication review
At some point in the future, how far out we do not exactly know, wireless access to the Internet will outstrip all other forms of access bringing the freedom of mobility to the way...
Goal Progress Velocity as a Determinant of Shortcut Behaviors
Goal Progress Velocity as a Determinant of Shortcut Behaviors
Employees often have a great deal of work to accomplish within stringent deadlines. Therefore, employees may engage in shortcut behaviors, which involve eschewing standard procedur...

Back to Top