Javascript must be enabled to continue!
Dense Contrastive Visual-Linguistic Pretraining
View through CrossRef
Inspired by the success of BERT, several multimodal representation learning approaches have been proposed that jointly represent image and text. These approaches achieve superior performance by capturing high-level semantic information from large-scale multimodal pretraining. In particular, LXMERT and UNITER adopt visual region feature regression and label classification as pretext tasks. However, they tend to suffer from the problems of noisy labels and sparse semantic annotations, based on the visual features having been pretrained on a crowdsourced dataset with limited and inconsistent semantic labeling.To overcome these issues, we propose unbiased Dense Contrastive Visual-Linguistic Pretraining (DCVLP), which replaces the region regression and classification with cross-modality region contrastive learning that requires no annotations. Two data augmentation strategies (Mask Perturbation and Intra-/Inter-Adversarial Perturbation) are developed to improve the quality of negative samples used in contrastive learning. Overall, DCVLP allows cross-modality dense region contrastive learning in a self-supervised setting independent of any object annotations. We compare our method against prior visual-linguistic pretraining frameworks to validate the superiority of dense contrastive learning on multimodal representation learning.
Center for Open Science
Title: Dense Contrastive Visual-Linguistic Pretraining
Description:
Inspired by the success of BERT, several multimodal representation learning approaches have been proposed that jointly represent image and text.
These approaches achieve superior performance by capturing high-level semantic information from large-scale multimodal pretraining.
In particular, LXMERT and UNITER adopt visual region feature regression and label classification as pretext tasks.
However, they tend to suffer from the problems of noisy labels and sparse semantic annotations, based on the visual features having been pretrained on a crowdsourced dataset with limited and inconsistent semantic labeling.
To overcome these issues, we propose unbiased Dense Contrastive Visual-Linguistic Pretraining (DCVLP), which replaces the region regression and classification with cross-modality region contrastive learning that requires no annotations.
Two data augmentation strategies (Mask Perturbation and Intra-/Inter-Adversarial Perturbation) are developed to improve the quality of negative samples used in contrastive learning.
Overall, DCVLP allows cross-modality dense region contrastive learning in a self-supervised setting independent of any object annotations.
We compare our method against prior visual-linguistic pretraining frameworks to validate the superiority of dense contrastive learning on multimodal representation learning.
Related Results
Teoretyczne badania konfrontatywne
Teoretyczne badania konfrontatywne
Theoretical contrastive studiesThe contrastive studies criticism in the 60s – 70s of the 20th century was the only basis for quite a few researchers to formulate their opinion on t...
Temporal-Aware and Intent Contrastive Learning for Sequential Recommendation
Temporal-Aware and Intent Contrastive Learning for Sequential Recommendation
In recent years, research in sequential recommendation has primarily refined user intent by constructing sequence-level contrastive learning tasks through data augmentation or by e...
Grouped Contrastive Learning of Self-supervised Sentence Representation
Grouped Contrastive Learning of Self-supervised Sentence Representation
This paper proposes a Grouped Contrastive Learning of self-supervised Sentence Representation (GCLSR), which can learn an effective and meaningful representation of sentences. Prev...
Contrastive knowledge how
Contrastive knowledge how
Classical empiricists are notorious for claiming that knowledge-how is a complex disposition or ability that is directly exhibited in action. This dissertation evaluates that claim...
Comparative Evaluation of Self-Supervised Pretraining Strategies for Few-Shot Medical Image Analysis
Comparative Evaluation of Self-Supervised Pretraining Strategies for Few-Shot Medical Image Analysis
Self-supervised learning has emerged as a promising solution to address the chronic scarcity of labeled medical imaging data. This study presents a comprehensive evaluation of main...
Supervised Source-Task Pretraining for Low-Data Electrochemical Molecular Property Prediction
Supervised Source-Task Pretraining for Low-Data Electrochemical Molecular Property Prediction
Electrochemical molecular property data obtained from experiments or high-level quantum chemical calculations are often limited in size, which constrains the direct use of graph ne...
Prototype-Driven Dual-Perspective Collaborative Contrastive Fusion Network for Rotating Machinery Fault Diagnosis
Prototype-Driven Dual-Perspective Collaborative Contrastive Fusion Network for Rotating Machinery Fault Diagnosis
Recently, self-supervised learning frameworks based on contrastive learning have demonstrated superior performance in rotating machinery fault diagnosis with limited labeled data. ...
Different distributions of contrastive vowel nasalization in Basque
Different distributions of contrastive vowel nasalization in Basque
Contrastive vowel nasalization is usually a consequence of the reinterpretation of the phonetic nasalization of a vowel due to coarticulation with an adjacent nasal consonant as or...

