Search engine for discovering works of Art, research articles, and books related to Art and Culture
ShareThis
Javascript must be enabled to continue!

Deploying and scaling distributed parallel deep neural networks on the Tianhe-3 prototype system

View through CrossRef
AbstractDue to the increase in computing power, it is possible to improve the feature extraction and data fitting capabilities of DNN networks by increasing their depth and model complexity. However, the big data and complex models greatly increase the training overhead of DNN, so accelerating their training process becomes a key task. The Tianhe-3 peak speed is designed to target E-class, and the huge computing power provides a potential opportunity for DNN training. We implement and extend LeNet, AlexNet, VGG, and ResNet model training for a single MT-2000+ and FT-2000+ compute nodes, as well as extended multi-node clusters, and propose an improved gradient synchronization process for Dynamic Allreduce communication optimization strategy for the gradient synchronization process base on the ARM architecture features of the Tianhe-3 prototype, providing experimental data and theoretical basis for further enhancing and improving the performance of the Tianhe-3 prototype in large-scale distributed training of neural networks.
Title: Deploying and scaling distributed parallel deep neural networks on the Tianhe-3 prototype system
Description:
AbstractDue to the increase in computing power, it is possible to improve the feature extraction and data fitting capabilities of DNN networks by increasing their depth and model complexity.
However, the big data and complex models greatly increase the training overhead of DNN, so accelerating their training process becomes a key task.
The Tianhe-3 peak speed is designed to target E-class, and the huge computing power provides a potential opportunity for DNN training.
We implement and extend LeNet, AlexNet, VGG, and ResNet model training for a single MT-2000+ and FT-2000+ compute nodes, as well as extended multi-node clusters, and propose an improved gradient synchronization process for Dynamic Allreduce communication optimization strategy for the gradient synchronization process base on the ARM architecture features of the Tianhe-3 prototype, providing experimental data and theoretical basis for further enhancing and improving the performance of the Tianhe-3 prototype in large-scale distributed training of neural networks.

Related Results

NEURAL NETWORKS AND DEEP LEARNING: THEORITICAL INSIGHTS AND FRAMEWORKS
NEURAL NETWORKS AND DEEP LEARNING: THEORITICAL INSIGHTS AND FRAMEWORKS
“NEURAL NETWORKS AND DEEP LEARNING: THEORITICAL INSIGHTS AND FRAMEWORKS” is a comprehensive guide that dives deep into the world of neural networks and their applications in modern...
Fuzzy Chaotic Neural Networks
Fuzzy Chaotic Neural Networks
An understanding of the human brain’s local function has improved in recent years. But the cognition of human brain’s working process as a whole is still obscure. Both fuzzy logic ...
On the role of network dynamics for information processing in artificial and biological neural networks
On the role of network dynamics for information processing in artificial and biological neural networks
Understanding how interactions in complex systems give rise to various collective behaviours has been of interest for researchers across a wide range of fields. However, despite ma...
Deep convolutional neural network and IoT technology for healthcare
Deep convolutional neural network and IoT technology for healthcare
Background Deep Learning is an AI technology that trains computers to analyze data in an approach similar to the human brain. Deep learning algorithms can find ...
REVIEW AND ANALYSIS OF APPROACHES AND PRACTICAL APPLICATIONS OF HUMAN EMOTION RECOGNITION
REVIEW AND ANALYSIS OF APPROACHES AND PRACTICAL APPLICATIONS OF HUMAN EMOTION RECOGNITION
Human emotions are complex and multifaceted, making them difficult to quantify and analyze. However, as technology advances, researchers are exploring the artificial intelligence u...
Neural Networks In Mining Sciences – General Overview And Some Representative Examples
Neural Networks In Mining Sciences – General Overview And Some Representative Examples
Abstract The many difficult problems that must now be addressed in mining sciences make us search for ever newer and more efficient computer tools that can be used to solve those p...
A Six‐Month Comparison of Three Periodontal Local Antimicrobial Therapies in Persistent Periodontal Pockets
A Six‐Month Comparison of Three Periodontal Local Antimicrobial Therapies in Persistent Periodontal Pockets
Background: Currently, several local antimicrobial delivery systems are available to periodontists. The aim of this 6‐month follow‐up parallel study was to evaluate the efficacy of...
On Robust and Efficient Parallel Reservoir Simulation on Tianhe-2
On Robust and Efficient Parallel Reservoir Simulation on Tianhe-2
Abstract Parallel reservoir simulators are now widely used with availability of super computers. Modern massively parallel supercomputers demonstrate great power for...

Back to Top