Javascript must be enabled to continue!
Screening Deep Learning Inference Accelerators at the Production Lines
View through CrossRef
Artificial Intelligence (AI) accelerators can be divided into two main buckets, one for training and another for inference over the trained models. Computation results of AI inference chipsets are expected to be deterministic for a given input. There are different compute engines on the Inference chip which help in acceleration of the Arithmetic operations. The Inference output results are compared with a golden reference output for the accuracy measurements. There can be many errors which can occur during the Inference execution. These errors could be due to the faulty hardware units and these units should be thoroughly screened in the assembly line before they are deployed by the customers in the data centre. This paper talks about a generic Inference application that has been developed to execute inferences over multiple inputs for various real inference models and stress all the compute engines of the Inference chip. Inference outputs from a specific inference unit are stored and are assumed to be golden and further confirmed as golden statistically. Once the golden reference outputs are established, Inference application is deployed in the pre- and post-production environments to screen out defective units whose actual output do not match the reference. Strategy to compare against itself at mass scale resulted in achieving the Defects Per Million target for the customers.
Academy and Industry Research Collaboration Center (AIRCC)
Title: Screening Deep Learning Inference Accelerators at the Production Lines
Description:
Artificial Intelligence (AI) accelerators can be divided into two main buckets, one for training and another for inference over the trained models.
Computation results of AI inference chipsets are expected to be deterministic for a given input.
There are different compute engines on the Inference chip which help in acceleration of the Arithmetic operations.
The Inference output results are compared with a golden reference output for the accuracy measurements.
There can be many errors which can occur during the Inference execution.
These errors could be due to the faulty hardware units and these units should be thoroughly screened in the assembly line before they are deployed by the customers in the data centre.
This paper talks about a generic Inference application that has been developed to execute inferences over multiple inputs for various real inference models and stress all the compute engines of the Inference chip.
Inference outputs from a specific inference unit are stored and are assumed to be golden and further confirmed as golden statistically.
Once the golden reference outputs are established, Inference application is deployed in the pre- and post-production environments to screen out defective units whose actual output do not match the reference.
Strategy to compare against itself at mass scale resulted in achieving the Defects Per Million target for the customers.
Related Results
CREATING LEARNING MEDIA IN TEACHING ENGLISH AT SMP MUHAMMADIYAH 2 PAGELARAN ACADEMIC YEAR 2020/2021
CREATING LEARNING MEDIA IN TEACHING ENGLISH AT SMP MUHAMMADIYAH 2 PAGELARAN ACADEMIC YEAR 2020/2021
The pandemic Covid-19 currently demands teachers to be able to use technology in teaching and learning process. But in reality there are still many teachers who have not been able ...
The feasibility of risk-stratified screening as routine practice in the NHS Breast Screening Programme in England: the PROCAS2 research programme
The feasibility of risk-stratified screening as routine practice in the NHS Breast Screening Programme in England: the PROCAS2 research programme
Background
Screening for breast cancer produces benefits through cancers being detected earlier, thereby reducing premature deaths and the need for more intensi...
Latency-Critical Inference Serving for Deep Learning
Latency-Critical Inference Serving for Deep Learning
Deep learning (DL) technology has made remarkable strides in terms of accuracy through the advancement of sophisticated and large deep neural networks (DNNs). Yet, its adoption in ...
Nanosilicas as Accelerators in Oilwell Cementing at Low Temperatures
Nanosilicas as Accelerators in Oilwell Cementing at Low Temperatures
Abstract
Accelerators are important cementing additives in deepwater wells where low temperatures can lengthen the wait-on-cement (WOC) time, potentially increasing ...
Evaluating the Effectiveness of Randomized and Directed Testbenches in Stress Testing AI Accelerators
Evaluating the Effectiveness of Randomized and Directed Testbenches in Stress Testing AI Accelerators
As the demand for high-performance AI accelerators grows, ensuring their reliability under extreme computational loads becomes paramount. This study evaluates the effectiveness of ...
Selection of Injectable Drug Product Composition using Machine Learning Models (Preprint)
Selection of Injectable Drug Product Composition using Machine Learning Models (Preprint)
BACKGROUND
As of July 2020, a Web of Science search of “machine learning (ML)” nested within the search of “pharmacokinetics or pharmacodynamics” yielded over 100...
Determinants of breast and cervical cancer screening uptake among women in Zambia: analysis of the 2024 Zambia demographic and health survey
Determinants of breast and cervical cancer screening uptake among women in Zambia: analysis of the 2024 Zambia demographic and health survey
Abstract
Breast and cervical cancers remain leading causes of cancer-related morbidity and mortality among women globally, particularly in low- and middle-income ...
Genetic Diversity in Maintainer and Restorer Lines of Pearl Millet
Genetic Diversity in Maintainer and Restorer Lines of Pearl Millet
ABSTRACTMolecular markers facilitate rapid and environment‐neutral characterization of the pattern of genetic diversity. The International Crops Research Institute for the Semi‐Ari...

