Javascript must be enabled to continue!
A Compact and Flexible FPGA Accelerator for Regular and Octave Convolutional Neural Networks
View through CrossRef
Convolutional Neural Networks has (CNN) have been widely used in various
Internet-of-Things (IoT) applications to offer smart solutions in our
daily lives. Many hardware accelerators were proposed to speed up CNN,
so that it can be used in the IoT sensor nodes. Recently, octave
convolution was proposed to remove spatial redundancy from input feature
maps, thus reducing the CNN computation and memory cost. Although octave
CNN is an attractive candidate for IoT applications, it cannot replace
regular CNN, which is already widely used in many applications. Hence,
it is desirable to have a hardware accelerator that can flexibly support
both regular and octave convolution. In this work, we present the first
compact and flexible CNN accelerator on an Field-Programmable-Gate-Array
(FPGA), which supports normal, pointwise and depthwise convolutions for
state-of-the-art octave convolution, as well as the existing regular
convolution. To ensure efficient accelerator utilization, a novel
adaptive scheduling scheme is presented to schedule the number of
computation tasks based on varying feature map dimensions across CNN
layers. A novel integration technique is presented to execute the
varying computational patterns in octave convolution due to multiple
branching in a single computation unit. An efficient memory scheme and
execution scheduling for octave convolution are also proposed to overlap
the long data fetch time with the on-going octave computation for better
overall computation performance. Comparing to other state-of-the-art
work, our proposed accelerator achieved 1.98x higher computation density
for MobileNetV2 implementation on XC7ZU9EG, as well as about 83.9%
lesser power for ResNet-50 implementation on VU9P.
Institute of Electrical and Electronics Engineers (IEEE)
Title: A Compact and Flexible FPGA Accelerator for Regular and Octave Convolutional Neural Networks
Description:
Convolutional Neural Networks has (CNN) have been widely used in various
Internet-of-Things (IoT) applications to offer smart solutions in our
daily lives.
Many hardware accelerators were proposed to speed up CNN,
so that it can be used in the IoT sensor nodes.
Recently, octave
convolution was proposed to remove spatial redundancy from input feature
maps, thus reducing the CNN computation and memory cost.
Although octave
CNN is an attractive candidate for IoT applications, it cannot replace
regular CNN, which is already widely used in many applications.
Hence,
it is desirable to have a hardware accelerator that can flexibly support
both regular and octave convolution.
In this work, we present the first
compact and flexible CNN accelerator on an Field-Programmable-Gate-Array
(FPGA), which supports normal, pointwise and depthwise convolutions for
state-of-the-art octave convolution, as well as the existing regular
convolution.
To ensure efficient accelerator utilization, a novel
adaptive scheduling scheme is presented to schedule the number of
computation tasks based on varying feature map dimensions across CNN
layers.
A novel integration technique is presented to execute the
varying computational patterns in octave convolution due to multiple
branching in a single computation unit.
An efficient memory scheme and
execution scheduling for octave convolution are also proposed to overlap
the long data fetch time with the on-going octave computation for better
overall computation performance.
Comparing to other state-of-the-art
work, our proposed accelerator achieved 1.
98x higher computation density
for MobileNetV2 implementation on XC7ZU9EG, as well as about 83.
9%
lesser power for ResNet-50 implementation on VU9P.
Related Results
A Compact and Flexible FPGA Accelerator for Regular and Octave Convolutional Neural Networks
A Compact and Flexible FPGA Accelerator for Regular and Octave Convolutional Neural Networks
<p>Convolutional Neural Networks has (CNN) have been widely used in various Internet-of-Things (IoT) applications to offer smart solutions in our daily lives. Many hardware a...
A Compact and Flexible FPGA Accelerator for Regular and Octave Convolutional Neural Networks
A Compact and Flexible FPGA Accelerator for Regular and Octave Convolutional Neural Networks
Convolutional Neural Networks has (CNN) have been widely used in various
Internet-of-Things (IoT) applications to offer smart solutions in our
daily lives. Many hardware accelerato...
Method of QoS evaluation of FPGA as a service
Method of QoS evaluation of FPGA as a service
The subject of study in this article is the evaluation of the performance issues of cloud services implemented using FPGA technology. The goal is to improve the performance of clou...
Аналіз застосування технологій ПЛІС в складі IoT
Аналіз застосування технологій ПЛІС в складі IoT
The subject of study in this article and work is the modern technologies of programmable logic devices (PLD) classified as FPGA, and the peculiarities of its application in Interne...
Graph convolutional neural networks for 3D data analysis
Graph convolutional neural networks for 3D data analysis
(English) Deep Learning allows the extraction of complex features directly from raw input data, eliminating the need for hand-crafted features from the classical Machine Learning p...
L'origine de l'octave courte
L'origine de l'octave courte
Contrairement à une idée répandue, l'octave courte n'est probablement pas née d'abord à l'orgue. Il n'en existe pratiquement aucune mention antérieure au XVIe siècle. Même ensuite,...
Methods of Deployment and Evaluation of FPGA as a Service Under Conditions of Changing Requirements and Environments
Methods of Deployment and Evaluation of FPGA as a Service Under Conditions of Changing Requirements and Environments
Applying Field Programmable Gate Array (FPGA) technology in cloud infrastructure and heterogeneous computations is of great interest today. FPGA as a Service assumes that the progr...
Development of a model for determining the necessary FPGA computing resource for placing a multilayer neural network on it
Development of a model for determining the necessary FPGA computing resource for placing a multilayer neural network on it
In this paper, the object of the research is the implementation of artificial neural networks (ANN) on FPGA. The problem to be solved is the construction of a mathematical model us...

