Javascript must be enabled to continue!
Latency-Critical Inference Serving for Deep Learning
View through CrossRef
Deep learning (DL) technology has made remarkable strides in terms of accuracy through the advancement of sophisticated and large deep neural networks (DNNs). Yet, its adoption in real-world applications is still challenging. DL-based applications, in particular video analytics, impose stringent low-latency and high-accuracy requirements on the DNN inference serving system that manages the deployment of DNNs and their inference. These application requirements are challenging for inference serving systems to satisfy as they are typically conflicting objectives. The challenge is exacerbated by the ever-increasing size of DNNs and the diverse set of computing platforms (on-device, edge, and cloud), of which none is best suited for serving DNNs independently. Therefore, exploring the design space of inference serving systems in an effort to identify a suitable solution that satisfies the application requirements by utilizing large DNNs and useful features of various computing platforms has emerged as an important problem that we study in this thesis.
This dissertation presents our study on how to design a latency-critical inference serving system for deep learning to meet the low-latency and high-accuracy requirements of DL-based video analytics that manifests itself in a large class of modern DL-based applications. The study is organized into four parts, each offering a step-by-step investigation toward designing a comprehensive solution.
First, we present a design for a hybrid networked system named Clownfish that deploys a small yet real-time DNN on the end device and a large and accurate DNN in the cloud. By leveraging the temporal correlation property in the video data, Clownfish enhances on-device analytics with the help of delayed and intermittent feedback from the cloud. Importantly, Clownfish removes the cloud from the critical path of the DNN inference pipeline. As a result, it always operates in real-time, depending on the latency performance of the small DNN deployed on the end device. Second, we remove the assumption of the temporal correlation property in Clownfish and present an edge inference serving system that is broadly applicable. Specifically, we offload DNN inference to edge servers and serve inference by streaming video data over the communication network, which can be highly dynamic in nature. Such a dynamic communication network leads to a variable data transfer time that may ultimately affect the latency of inference requests. To overcome this variability, we propose to conduct data and DNN adaptation jointly and present a feedback control mechanism that makes adaptation decisions efficiently. This system only applies to single-client and single-server setups; hence, we remove this limitation in the next part of the exploration. Third, we design a scalable edge inference serving system named Jellyfish to serve DNN inference for multiple clients on multiple edge resources. To this end, we propose a collective adaptation technique that conducts the data and DNN adaptation jointly for multiple clients. Finally, we optimize this scalable serving system and propose ideas to leverage dynamic DNNs to avoid the DNN adaptation overhead and improve the batched inference efficiency.
Overall, in an effort to enable the adoption of DL technology to broader video analytics applications, we propose solutions for designing a DNN inference serving system to deliver analytics results to multiple users while satisfying their latency requirements with high analytics accuracy. Based on these explorations, we also present a list of key takeaways and design implications that we hope will be helpful to designers and researchers of the DNN inference serving system.
Title: Latency-Critical Inference Serving for Deep Learning
Description:
Deep learning (DL) technology has made remarkable strides in terms of accuracy through the advancement of sophisticated and large deep neural networks (DNNs).
Yet, its adoption in real-world applications is still challenging.
DL-based applications, in particular video analytics, impose stringent low-latency and high-accuracy requirements on the DNN inference serving system that manages the deployment of DNNs and their inference.
These application requirements are challenging for inference serving systems to satisfy as they are typically conflicting objectives.
The challenge is exacerbated by the ever-increasing size of DNNs and the diverse set of computing platforms (on-device, edge, and cloud), of which none is best suited for serving DNNs independently.
Therefore, exploring the design space of inference serving systems in an effort to identify a suitable solution that satisfies the application requirements by utilizing large DNNs and useful features of various computing platforms has emerged as an important problem that we study in this thesis.
This dissertation presents our study on how to design a latency-critical inference serving system for deep learning to meet the low-latency and high-accuracy requirements of DL-based video analytics that manifests itself in a large class of modern DL-based applications.
The study is organized into four parts, each offering a step-by-step investigation toward designing a comprehensive solution.
First, we present a design for a hybrid networked system named Clownfish that deploys a small yet real-time DNN on the end device and a large and accurate DNN in the cloud.
By leveraging the temporal correlation property in the video data, Clownfish enhances on-device analytics with the help of delayed and intermittent feedback from the cloud.
Importantly, Clownfish removes the cloud from the critical path of the DNN inference pipeline.
As a result, it always operates in real-time, depending on the latency performance of the small DNN deployed on the end device.
Second, we remove the assumption of the temporal correlation property in Clownfish and present an edge inference serving system that is broadly applicable.
Specifically, we offload DNN inference to edge servers and serve inference by streaming video data over the communication network, which can be highly dynamic in nature.
Such a dynamic communication network leads to a variable data transfer time that may ultimately affect the latency of inference requests.
To overcome this variability, we propose to conduct data and DNN adaptation jointly and present a feedback control mechanism that makes adaptation decisions efficiently.
This system only applies to single-client and single-server setups; hence, we remove this limitation in the next part of the exploration.
Third, we design a scalable edge inference serving system named Jellyfish to serve DNN inference for multiple clients on multiple edge resources.
To this end, we propose a collective adaptation technique that conducts the data and DNN adaptation jointly for multiple clients.
Finally, we optimize this scalable serving system and propose ideas to leverage dynamic DNNs to avoid the DNN adaptation overhead and improve the batched inference efficiency.
Overall, in an effort to enable the adoption of DL technology to broader video analytics applications, we propose solutions for designing a DNN inference serving system to deliver analytics results to multiple users while satisfying their latency requirements with high analytics accuracy.
Based on these explorations, we also present a list of key takeaways and design implications that we hope will be helpful to designers and researchers of the DNN inference serving system.
Related Results
What's the Delay? Understanding Latency Across the Network
What's the Delay? Understanding Latency Across the Network
Network latency directly affects the performance of many applications that run over the Internet. While significant effort is spent on reducing network latency, the fundamental cap...
Flexible architecture for the future internet scalability of SDN control plane
Flexible architecture for the future internet scalability of SDN control plane
Software-Defined Networking (SDN) separates the control plane from the data plane. The initial SDN approach involves a single centralized controller, which may not scale properly a...
CREATING LEARNING MEDIA IN TEACHING ENGLISH AT SMP MUHAMMADIYAH 2 PAGELARAN ACADEMIC YEAR 2020/2021
CREATING LEARNING MEDIA IN TEACHING ENGLISH AT SMP MUHAMMADIYAH 2 PAGELARAN ACADEMIC YEAR 2020/2021
The pandemic Covid-19 currently demands teachers to be able to use technology in teaching and learning process. But in reality there are still many teachers who have not been able ...
Selection of Injectable Drug Product Composition using Machine Learning Models (Preprint)
Selection of Injectable Drug Product Composition using Machine Learning Models (Preprint)
BACKGROUND
As of July 2020, a Web of Science search of “machine learning (ML)” nested within the search of “pharmacokinetics or pharmacodynamics” yielded over 100...
Towards Ubiquitous and Continuous Network Latency Monitoring
Towards Ubiquitous and Continuous Network Latency Monitoring
The Internet plays an important role in modern society, and its network performance impacts billions of users every day. For many network applications, network latency has a large ...
Other-serving vs Self-serving Instructions in US College Commencement Speeches: A Quantitative Study
Other-serving vs Self-serving Instructions in US College Commencement Speeches: A Quantitative Study
INTRODUCTION: Research supports that serving others and practicing altruism is beneficial for one’s health, wellbeing, and success compared to solely serving oneself. However, it i...
Heat shock protein 90 is a master regulator of HIV-1 latency
Heat shock protein 90 is a master regulator of HIV-1 latency
Abstract
An estimated 32 million people live with HIV-1 globally. Combined antiretroviral therapy suppresses viral replication but therapy interruption results in v...
Deep Learning: Implications for Human Learning and Memory
Deep Learning: Implications for Human Learning and Memory
Recent years have seen an explosion of interest in deep learning and deep neural networks. Deep learning lies at the heart of unprecedented feats of machine intelligence as well as...

