Javascript must be enabled to continue!
Explainable Attention Pruning: A Meta-learning-based Approach
View through CrossRef
Pruning, as a technique to reduce the complexity and size of Transformer-based models, has gained significant attention in recent years. While various models have been successfully pruned, pruning BERT poses unique challenges due to their fine-grained structure and overparameterization. However, by carefully considering these factors, it is possible to prune BERT without significantly degrading its pre-trained loss. In this paper, we propose a Meta-learning-based pruning approach that can adaptively identify and eliminate insignificant attention weights. The performance of the proposed model is compared with several baseline models, as well as the default fine-tuned BERT model. The baseline pruning strategies employed low-level pruning techniques, targeting the removal of only 20% of the connections. The experimental results show that the proposed model outperforms the other baseline models, in terms of lower inference latency, higher MCC and lower loss. However, there is no significant improvement observed in terms of average FLOPs (floating-point operations per second). Furthermore, we conduct a comparative evaluation of the baseline models and our proposed model using two explainable (XAI) approaches. While other models allocate reasonable attention to less significant words for sentiment classification, our model assigns higher probabilities to the most significant sentimental words. Impact Statement-Efficient handling of inference time in pre-trained language models (PLMs) and the preservation of performance while reducing their size are important research considerations. Model compression techniques, such as pruning, are recognized as effective approaches for achieving memoryefficient, energy-efficient, computation-efficient, and storageefficient PLMs. Pruning addresses the need to create compact models without compromising their overall effectiveness. Existing pruning methods often rely on task and domain-specific approaches and therefore, it is important to explore a domainindependent pruning approach. We propose a new pruning strategy called Meta-Controller-based Attention Pruning (MCAP) for the BERT model targeting single-sentence prediction tasks. MCAP optimization strategy eliminates insignificant attention in the BERT by calculating their importance scores. The selfsupervised pruner in MCAP uses a meta-learning approach to identify and eliminate these insignificant attentions before finetuning. Our study compares MCAP with baseline models (both structured and unstructured pruning) and compared it with inference latency, MCC, and loss parameters. The results show that MCAP outperforms the baseline models in terms of inference latency, MCC, and loss. Explainable AI (XAI) techniques are used to interpret the model's decisions and predictions. MCAP focuses on significant words in sentiment classification, ensuring important model parameters are retained without a significant impact on output.
Title: Explainable Attention Pruning: A Meta-learning-based Approach
Description:
Pruning, as a technique to reduce the complexity and size of Transformer-based models, has gained significant attention in recent years.
While various models have been successfully pruned, pruning BERT poses unique challenges due to their fine-grained structure and overparameterization.
However, by carefully considering these factors, it is possible to prune BERT without significantly degrading its pre-trained loss.
In this paper, we propose a Meta-learning-based pruning approach that can adaptively identify and eliminate insignificant attention weights.
The performance of the proposed model is compared with several baseline models, as well as the default fine-tuned BERT model.
The baseline pruning strategies employed low-level pruning techniques, targeting the removal of only 20% of the connections.
The experimental results show that the proposed model outperforms the other baseline models, in terms of lower inference latency, higher MCC and lower loss.
However, there is no significant improvement observed in terms of average FLOPs (floating-point operations per second).
Furthermore, we conduct a comparative evaluation of the baseline models and our proposed model using two explainable (XAI) approaches.
While other models allocate reasonable attention to less significant words for sentiment classification, our model assigns higher probabilities to the most significant sentimental words.
Impact Statement-Efficient handling of inference time in pre-trained language models (PLMs) and the preservation of performance while reducing their size are important research considerations.
Model compression techniques, such as pruning, are recognized as effective approaches for achieving memoryefficient, energy-efficient, computation-efficient, and storageefficient PLMs.
Pruning addresses the need to create compact models without compromising their overall effectiveness.
Existing pruning methods often rely on task and domain-specific approaches and therefore, it is important to explore a domainindependent pruning approach.
We propose a new pruning strategy called Meta-Controller-based Attention Pruning (MCAP) for the BERT model targeting single-sentence prediction tasks.
MCAP optimization strategy eliminates insignificant attention in the BERT by calculating their importance scores.
The selfsupervised pruner in MCAP uses a meta-learning approach to identify and eliminate these insignificant attentions before finetuning.
Our study compares MCAP with baseline models (both structured and unstructured pruning) and compared it with inference latency, MCC, and loss parameters.
The results show that MCAP outperforms the baseline models in terms of inference latency, MCC, and loss.
Explainable AI (XAI) techniques are used to interpret the model's decisions and predictions.
MCAP focuses on significant words in sentiment classification, ensuring important model parameters are retained without a significant impact on output.
Related Results
Ground-Level Pruning at Right Time Improves Flower Yield of Old Plantation of Rosa damascena Without Compromising the Quality of Essential Oil
Ground-Level Pruning at Right Time Improves Flower Yield of Old Plantation of Rosa damascena Without Compromising the Quality of Essential Oil
The essential oil of Rosa damascena is extensively used as a key natural ingredient in the perfume and cosmetic industries. However, the productivity and quality of rose oil are a ...
DARB: A Density-Adaptive Regular-Block Pruning for Deep Neural Networks
DARB: A Density-Adaptive Regular-Block Pruning for Deep Neural Networks
The rapidly growing parameter volume of deep neural networks (DNNs) hinders the artificial intelligence applications on resource constrained devices, such as mobile and wearable de...
Effect of Pruning Intensities on the Performance of Fruit Plants under Mid-Hill Condition of Eastern Himalayas: Case Study on Guava
Effect of Pruning Intensities on the Performance of Fruit Plants under Mid-Hill Condition of Eastern Himalayas: Case Study on Guava
Current study was undertaken to highlight the effect of pruning on improving vigor of old orchards and increasing performance in terms of fruit yield and quality under water and nu...
A research on rejuvenation pruning of lavandin (Lavandula x intermedia Emeric ex Loisel.)
A research on rejuvenation pruning of lavandin (Lavandula x intermedia Emeric ex Loisel.)
Objective: The main purpose of the research was investigate whether to be renewed or not without the need for re-planting by rejuvenation pruning to the aged plantations of lavandi...
The Influence of Pruning on the Growth and Wood Properties of Populus deltoides “Nanlin 3804”
The Influence of Pruning on the Growth and Wood Properties of Populus deltoides “Nanlin 3804”
During the natural growth of trees, a large number of branches are formed, with a negative impact on timber quality. Therefore, pruning is an essential measure in forest cultivatio...
Sensing and Automation in Pruning of Tree Fruit Crops: A Review
Sensing and Automation in Pruning of Tree Fruit Crops: A Review
Pruning is one of the most important tree fruit production activities, which is highly dependent on human labor. Skilled labor is in short supply, and the increasing cost of labor ...
Advancing Transformer Efficiency with Token Pruning
Advancing Transformer Efficiency with Token Pruning
Transformer-based models have revolutionized natural language processing (NLP), achieving state-of-the-art performance across a wide range of tasks. However, their high computation...
Cost of pruning Douglas-fir in coastal British Columbia
Cost of pruning Douglas-fir in coastal British Columbia
Artificial pruning can increase the quantity of high-value clear lumber harvested from Douglas-fir, but the pruning cost per tree is relatively high. To prune a young Douglas-fir t...

