Javascript must be enabled to continue!
Generalized Hierarchical Co-Saliency Learning for Label-Efficient Tracking
View through CrossRef
Visual object tracking is one of the core techniques in human-centered artificial intelligence, which is very useful for human–machine interaction. State-of-the-art tracking methods have shown their robustness and accuracy on many challenges. However, a large amount of videos with precisely dense annotations are required for fully supervised training of their models. Considering that annotating videos frame-by-frame is a labor- and time-consuming workload, reducing the reliance on manual annotations during the tracking models’ training is an important problem to be resolved. To make a trade-off between the annotating costs and the tracking performance, we propose a weakly supervised tracking method based on co-saliency learning, which can be flexibly integrated into various tracking frameworks to reduce annotation costs and further enhance the target representation in current search images. Since our method enables the model to explore valuable visual information from unlabeled frames, and calculate co-salient attention maps based on multiple frames, our weakly supervised methods can obtain competitive performance compared to fully supervised baseline trackers, using only 3.33% of manual annotations. We integrate our method into two CNN-based trackers and a Transformer-based tracker; extensive experiments on four general tracking benchmarks demonstrate the effectiveness of our method. Furthermore, we also demonstrate the advantages of our method on egocentric tracking task; our weakly supervised method obtains 0.538 success on TREK-150, which is superior to prior state-of-the-art fully supervised tracker by 7.7%.
Title: Generalized Hierarchical Co-Saliency Learning for Label-Efficient Tracking
Description:
Visual object tracking is one of the core techniques in human-centered artificial intelligence, which is very useful for human–machine interaction.
State-of-the-art tracking methods have shown their robustness and accuracy on many challenges.
However, a large amount of videos with precisely dense annotations are required for fully supervised training of their models.
Considering that annotating videos frame-by-frame is a labor- and time-consuming workload, reducing the reliance on manual annotations during the tracking models’ training is an important problem to be resolved.
To make a trade-off between the annotating costs and the tracking performance, we propose a weakly supervised tracking method based on co-saliency learning, which can be flexibly integrated into various tracking frameworks to reduce annotation costs and further enhance the target representation in current search images.
Since our method enables the model to explore valuable visual information from unlabeled frames, and calculate co-salient attention maps based on multiple frames, our weakly supervised methods can obtain competitive performance compared to fully supervised baseline trackers, using only 3.
33% of manual annotations.
We integrate our method into two CNN-based trackers and a Transformer-based tracker; extensive experiments on four general tracking benchmarks demonstrate the effectiveness of our method.
Furthermore, we also demonstrate the advantages of our method on egocentric tracking task; our weakly supervised method obtains 0.
538 success on TREK-150, which is superior to prior state-of-the-art fully supervised tracker by 7.
7%.
Related Results
Depth-aware salient object segmentation
Depth-aware salient object segmentation
Object segmentation is an important task which is widely employed in many computer vision applications such as object detection, tracking, recognition, and ret...
Neural Correlates of High-Level Visual Saliency Models
Neural Correlates of High-Level Visual Saliency Models
Abstract
Visual saliency highlights regions in a scene that are most relevant to an observer. The process by which a saliency map is formed has been a crucial subje...
Recent Advances in Saliency Estimation for Omnidirectional Images, Image Groups, and Video Sequences
Recent Advances in Saliency Estimation for Omnidirectional Images, Image Groups, and Video Sequences
We present a review of methods for automatic estimation of visual saliency: the perceptual property that makes specific elements in a scene stand out and grab the attention of the ...
CREATING LEARNING MEDIA IN TEACHING ENGLISH AT SMP MUHAMMADIYAH 2 PAGELARAN ACADEMIC YEAR 2020/2021
CREATING LEARNING MEDIA IN TEACHING ENGLISH AT SMP MUHAMMADIYAH 2 PAGELARAN ACADEMIC YEAR 2020/2021
The pandemic Covid-19 currently demands teachers to be able to use technology in teaching and learning process. But in reality there are still many teachers who have not been able ...
A Dynamic Bottom-Up Saliency Detection Method for Still Images
A Dynamic Bottom-Up Saliency Detection Method for Still Images
AbstractIntroductionExisting saliency detection algorithms in the literature have ignored the importance of time. They create a static saliency map for the whole recording time. Ho...
Fitting‐based optimisation for image visual salient object detection
Fitting‐based optimisation for image visual salient object detection
To overcome some major problems with traditional saliency evaluation metrics, full‐reference image quality assessment (IQA) metrics, which have similar but stricter objectives, are...
CF‐based optimisation for saliency detection
CF‐based optimisation for saliency detection
In view of the observation that saliency maps generated by saliency detection algorithms usually show similarity imperfection against the ground truth, the authors propose an optim...
Is a Fitbit a Diary? Self-Tracking and Autobiography
Is a Fitbit a Diary? Self-Tracking and Autobiography
Data becomes something of a mirror in which people see themselves reflected. (Sorapure 270)In a 2014 essay for The New Yorker, the humourist David Sedaris recounts an obsession spu...

