Javascript must be enabled to continue!
COWO: towards real-time spatiotemporal action localization in videos
View through CrossRef
Purpose
The purpose of this paper is to provide a fast and accurate network for spatiotemporal action localization in videos. It detects human actions both in time and space simultaneously in real-time, which is applicable in real-world scenarios such as safety monitoring and collaborative assembly.
Design/methodology/approach
This paper design an end-to-end deep learning network called collaborator only watch once (COWO). COWO recognizes the ongoing human activities in real-time with enhanced accuracy. COWO inherits from the architecture of you only watch once (YOWO), known to be the best performing network for online action localization to date, but with three major structural modifications: COWO enhances the intraclass compactness and enlarges the interclass separability in the feature level. A new correlation channel fusion and attention mechanism are designed based on the Pearson correlation coefficient. Accordingly, a correction loss function is designed. This function minimizes the same class distance and enhances the intraclass compactness. Use a probabilistic K-means clustering technique for selecting the initial seed points. The idea behind this is that the initial distance between cluster centers should be as considerable as possible. CIOU regression loss function is applied instead of the Smooth L1 loss function to help the model converge stably.
Findings
COWO outperforms the original YOWO with improvements of frame mAP 3% and 2.1% at a speed of 35.12 fps. Compared with the two-stream, T-CNN, C3D, the improvement is about 5% and 14.5% when applied to J-HMDB-21, UCF101-24 and AGOT data sets.
Originality/value
COWO extends more flexibility for assembly scenarios as it perceives spatiotemporal human actions in real-time. It contributes to many real-world scenarios such as safety monitoring and collaborative assembly.
Title: COWO: towards real-time spatiotemporal action localization in videos
Description:
Purpose
The purpose of this paper is to provide a fast and accurate network for spatiotemporal action localization in videos.
It detects human actions both in time and space simultaneously in real-time, which is applicable in real-world scenarios such as safety monitoring and collaborative assembly.
Design/methodology/approach
This paper design an end-to-end deep learning network called collaborator only watch once (COWO).
COWO recognizes the ongoing human activities in real-time with enhanced accuracy.
COWO inherits from the architecture of you only watch once (YOWO), known to be the best performing network for online action localization to date, but with three major structural modifications: COWO enhances the intraclass compactness and enlarges the interclass separability in the feature level.
A new correlation channel fusion and attention mechanism are designed based on the Pearson correlation coefficient.
Accordingly, a correction loss function is designed.
This function minimizes the same class distance and enhances the intraclass compactness.
Use a probabilistic K-means clustering technique for selecting the initial seed points.
The idea behind this is that the initial distance between cluster centers should be as considerable as possible.
CIOU regression loss function is applied instead of the Smooth L1 loss function to help the model converge stably.
Findings
COWO outperforms the original YOWO with improvements of frame mAP 3% and 2.
1% at a speed of 35.
12 fps.
Compared with the two-stream, T-CNN, C3D, the improvement is about 5% and 14.
5% when applied to J-HMDB-21, UCF101-24 and AGOT data sets.
Originality/value
COWO extends more flexibility for assembly scenarios as it perceives spatiotemporal human actions in real-time.
It contributes to many real-world scenarios such as safety monitoring and collaborative assembly.
Related Results
Analyzing Quality of YouTube Videos about Premature Ovarian Failure in the Past Decade
Analyzing Quality of YouTube Videos about Premature Ovarian Failure in the Past Decade
Abstract
Background
To determine the quality of YouTube videos about premature ovarian failure (POF), and variations in quality of professional YouTube videos about POF.
M...
Rhytidectomy: Analysis of Videos Available Online
Rhytidectomy: Analysis of Videos Available Online
AbstractThe objective of this study was to examine YouTube videos related to rhytidectomy created by both physicians and nonphysicians to determine the content of the videos, the s...
Indoor Localization System Based on RSSI-APIT Algorithm
Indoor Localization System Based on RSSI-APIT Algorithm
An indoor localization system based on the RSSI-APIT algorithm is designed in this study. Integrated RSSI (received signal strength indication) and non-ranging APIT (approximate pe...
Short video platforms as sources of health information about cervical cancer: A content and quality analysis
Short video platforms as sources of health information about cervical cancer: A content and quality analysis
BackgroundThe development of short popular science video platforms helps people obtain health information, but no research has evaluated the information characteristics and quality...
Quais os comentários negativos e estratégias de enfrentamento mais comuns e eficazes na plataforma digital Youtube?
Quais os comentários negativos e estratégias de enfrentamento mais comuns e eficazes na plataforma digital Youtube?
O presente estudo tem como objetivo compreender qual a tipologia de comentário negativo mais comum nos comentários referentes a vídeos postados na plataforma YouTube, bem como as t...
Born To Die: Lana Del Rey, Beauty Queen or Gothic Princess?
Born To Die: Lana Del Rey, Beauty Queen or Gothic Princess?
Closer examination of contemporary art forms including music videos in addition to the Gothic’s literature legacy is essential, “as it is virtually impossible to ignore the relatio...
The associations between Screen Time, Screen Content, and ADHD risk based on the evidence of 41494 children from Longhua district, Shenzhen, China
The associations between Screen Time, Screen Content, and ADHD risk based on the evidence of 41494 children from Longhua district, Shenzhen, China
ABSTRACT
Objective
This study investigates the relationship between screen time, screen content, and the risk of Attention Defi...
Reliability of YouTube© Videos in terms of Cardiopulmonary Resuscitation Education (Preprint)
Reliability of YouTube© Videos in terms of Cardiopulmonary Resuscitation Education (Preprint)
BACKGROUND
People particularly choose video-sharing platforms for medical information.
...

