nn-Meter: towards accurate latency prediction of deep-learning model inference on diverse edge devices

Li Lyna Zhang,Shihao Han,Jianyu Wei,Ningxin Zheng,Ting Cao,Yuqing Yang,Yunxin Liu

nn-Meter: towards accurate latency prediction of deep-learning model inference on diverse edge devices

2021

Li Lyna Zhang
Shihao Han
Jianyu Wei
Ningxin Zheng
Ting Cao
Yuqing Yang
Yunxin Liu

With the recent trend of on-device deep learning, inference latency has become a crucial metric in running Deep Neural Network (DNN) models on various mobile and edge devices. To this end, latency prediction of DNN model inference is highly desirable for many tasks where measuring the latency on real devices is infeasible or too costly, such as searching for efficient DNN models with latency constraints from a huge model-design space. Yet it is very challenging and existing approaches fail to achieve a high accuracy of prediction, due to the varying model-inference latency caused by the runtime optimizations on diverse edge devices.

Correction
Source
Cite
Save
Machine Reading By IdeaReader

References

Citations