Towards Understanding Model Quantization for Reliable Deep Neural Network Deployment

Qiang Hu Yuejun Guo Maxime Cordy Xiaofei Xie Wei Ma Mike Papadakis Yves Le Traon

Citation

Reference

Related Paper

Citation Trend

Abstract:

Deep Neural Networks (DNNs) have gained considerable attention in the past decades due to their astounding performance in different applications, such as natural language modeling, self-driving assistance, and source code understanding. With rapid exploration, more and more complex DNN architectures have been proposed along with huge pre-trained model parameters. A common way to use such DNN models in user-friendly devices (e.g., mobile phones) is to perform model compression before deployment. However, recent research has demonstrated that model compression, e.g., model quantization, yields accuracy degradation as well as output disagreements when tested on unseen data. Since the unseen data always include distribution shifts and often appear in the wild, the quality and reliability of models after quantization are not ensured. In this paper, we conduct a comprehensive study to characterize and help users understand the behaviors of quantization models. Our study considers four datasets spanning from image to text, eight DNN architectures including both feed-forward neural networks and recurrent neural networks, and 42 shifted sets with both synthetic and natural distribution shifts. The results reveal that 1) data with distribution shifts lead to more disagreements than without. 2) Quantization-aware training can produce more stable models than standard, adversarial, and Mixup training. 3) Disagreements often have closer top-1 and top-2 output probabilities, and Margin is a better indicator than other uncertainty metrics to distinguish disagreements. 4) Retraining the model with disagreements has limited efficiency in removing disagreements. We release our code and models as a new benchmark for further study of model quantization.

Keywords:

Deep Neural Networks

Benchmark (surveying)

Topics:

Adversarial Robustness in Machine Learning

Anomaly Detection Techniques and Applications

Advanced Neural Network Applications

10.1109/cain58948.2023.00015

Cite

PDF

Serving Away From Home: How Deployments Influence Reenlistment

RAND Corporation eBooks (2002)

James Hosek Mark E. Totten

How does deployment affect reenlistment? The authors look at this particular issue in wake of the high rate of military deployment throughout the 1990s and with the prospect that deployment will rise even more in the coming years. The research finds that reenlistment was higher among members who deployed compared with those who did not. The analysis suggests that past deployment influences current reenlistment behavior because it enables members to learn about their preferences for deployment.

Military deployment

Affect

10.7249/mr1594

Cite

Citations (20)

Research on Accelerating Application Technology of Centralized ERP System Based on HANA

Journal of Physics Conference Series (2019)

Haohai Zhang Wei Wang Xinqiao Gu Hao Wang Wei Zhang

Abstract In order to effectively improve the processing speed of centralized deployment in enterprises, this paper proposes a research based on HANA to accelerate the application technology of centralized deployment of ERP system. By optimizing the configuration of data processor unit in ERP physical deployment module, the running efficiency of the system is accelerated. The optimized data processor is used to calculate the parameter index of centralized deployment acceleration authority. According to the parameter index, the deployment role framework of each department of the enterprise is reasonably allocated, so as to reduce the centralized deployment operation time, achieve the goal of accelerating the system operation, and finally realize the effective application of the centralized deployment ERP system acceleration technology. Finally, through comparative experimental tests, it is confirmed that the actual application effect of the HANA - based accelerated application technology for centralized deployment of ERP system can reach more than 90 %, which is significantly improved compared with the traditional centralized deployment accelerated technology.

System deployment

10.1088/1742-6596/1314/1/012143

Cite

Citations (1)

AWS Deployment Strategies

Apress eBooks (2023)

O Guler Mustafa

In the previous chapter, I covered deployment and how to create a pipeline. Still, one of the central concepts when building the pipeline is understanding the deployment strategy and which type you will use, because there are multiple deployment types, each serving a different purpose depending on the use case and the company approach.

10.1007/978-1-4842-9303-4_5

Cite

Citations (1)

MEASUREMENTS OF WAVE IMPACTS AT FULL SCALE: RESULTS OF FIELDWORK ON CONCRETE ARMOUR UNITS

N. W. H. Allsop A. M. Vann M. Howarth RJ. Jones J. P. Davis

: Introduction The Problem Fieldwork Measurements: The Approach The First Two Deployments Third Deployment, September 1989 to June 1990 Fourth Deployment, November 1993 to April 1994 Results from Third Deployment Results from Fourth Deployment Further Work Acknowledgements

Armour

10.1680/aicsab.25097.0027

Cite

Citations (9)

Heuristic algorithms for effective broker deployment

Information Technology and Management (2011)

Yifeng Qian Beihong Jin Wenjing Fang

10.1007/s10799-011-0095-4

Cite

Citations (13)

Large-Scale Deployment of Tablet Computers in Brazilian Public Schools: Decisive Factors and an Implementation Model

Perspectives on rethinking and reforming education (2017)

Giovanni Ferreira de Farias Mohamed Ally Fernando José Spanhol

10.1007/978-981-10-6144-8_16

Cite

Citations (1)

Understanding deployment from the perspective of those who have served

Nursing Outlook (2016)

Bonnie Mowinski Jennings Kristal C. Melvin Donna L. Belew

Military deployment

Peacekeeping

10.1016/j.outlook.2016.12.005

Cite

Citations (5)

Evaluation of Quantization Techniques for Deep Neural Networks

Communications in computer and information science (2021)

Zhiyuan Li

Deep Neural Networks

10.1007/978-981-16-8885-0_12

Cite

Citations (0)

NATIONAL ITS PROGRAM PLAN : SYNOPSIS

G W Euler H D Robertson

This document provides a fifty page encapsulation of the major subject areas within the National ITS (Intelligent Transportation Systems) Program Plan, with special emphasis on the area of deployment. The document is organized into the following areas of discussion: ITS User Services, ITS National Compatibility, Current Deployment, Future Deployment, Scenarios of Deployment, Deployment Support, and Recommendations for Deployment.

Source

Cite

Citations (6)

Improving Performance of Direct-Detection Terahertz Communication System based on k-Means Adaptive Vector Quantization

2022 20th International Conference on Optical Communications and Networks (ICOCN) (2021)

Linghao Yue Yuancheng Cai Min Zhu Pengyuan Wang Liyao Zhang

The performance of the adaptive vector quantization based on k-means clustering for direct-detection terahertz communication in 0.3 THz band is studied by simulation. Compared with the traditional uniform quantization, the two-dimensional k-means quantization can effectively improve the bit error rate by more than an order of magnitude with 2 quantization bits per sample, which can facilitate the low-power and low-cost receivers.

Linde–Buzo–Gray algorithm

K-Means Clustering

10.1109/icocn53177.2021.9563683

Cite

Citations (1)