A GPU inference system scheduling algorithm with asynchronous data transfer

Qin Zhang; Li Zha; Xiaohua Wan; Boqun Cheng

doi:10.1109/IPDPSW.2019.00083

A GPU inference system scheduling algorithm with asynchronous data transfer

Qin Zhang, Li Zha^*, Xiaohua Wan, Boqun Cheng

^*此作品的通讯作者

科研成果: 书/报告/会议事项章节 › 会议稿件 › 同行评审

1 引用（Scopus）

摘要

With the rapid expansion of application range, Deep-Learning has increasingly become an indispensable practical method to solve problems in various industries. In different application scenarios, especially in high concurrency areas such as search and recommendation, deep learning inference system is required to have high throughput and low latency, which can not be easily obtained at the same time. In this paper, we build a model to quantify the relationship between concurrency, throughput and job latency. Then we implement a GPU scheduling algorithm for inference jobs in deep learning inference system based on the model. The algorithm predicts the completion time of batch jobs being executed, and reasonably chooses the batch size of the next batch jobs according to the concurrency and upload data to GPU memory ahead of time. So that the system can hide the data transfer delay of GPU and achieve the minimum job latency under the premise of meetingthethroughputrequirements.Experimentsshowthatthe proposed GPU asynchronous data transfer scheduling algorithm improves throughput by 9% compared with the traditional synchronous algorithm, reduces the latency by 3%-76% under different concurrency, and can better suppress the job latency fluctuation caused by concurrency changing.

源语言	英语
主期刊名	Proceedings - 2019 IEEE 33rd International Parallel and Distributed Processing Symposium Workshops, IPDPSW 2019
出版商	Institute of Electrical and Electronics Engineers Inc.
页	438-445
页数	8
ISBN（电子版）	9781728135106
DOI	https://doi.org/10.1109/IPDPSW.2019.00083
出版状态	已出版 - 5月 2019
已对外发布	是
活动	33rd IEEE International Parallel and Distributed Processing Symposium Workshops, IPDPSW 2019 - Rio de Janeiro, 巴西期限: 20 5月 2019 → 24 5月 2019

出版系列

姓名	Proceedings - 2019 IEEE 33rd International Parallel and Distributed Processing Symposium Workshops, IPDPSW 2019

会议

会议	33rd IEEE International Parallel and Distributed Processing Symposium Workshops, IPDPSW 2019
国家/地区	巴西
市	Rio de Janeiro
时期	20/05/19 → 24/05/19

访问文件

10.1109/IPDPSW.2019.00083

其它文件与链接

链接到 Scopus 的出版物

引用此

Zhang, Q., Zha, L., Wan, X., & Cheng, B. (2019). A GPU inference system scheduling algorithm with asynchronous data transfer. 在 Proceedings - 2019 IEEE 33rd International Parallel and Distributed Processing Symposium Workshops, IPDPSW 2019 (页码 438-445). 文章 8778365 (Proceedings - 2019 IEEE 33rd International Parallel and Distributed Processing Symposium Workshops, IPDPSW 2019). Institute of Electrical and Electronics Engineers Inc.. https://doi.org/10.1109/IPDPSW.2019.00083

Zhang, Qin ; Zha, Li ; Wan, Xiaohua 等. / A GPU inference system scheduling algorithm with asynchronous data transfer. Proceedings - 2019 IEEE 33rd International Parallel and Distributed Processing Symposium Workshops, IPDPSW 2019. Institute of Electrical and Electronics Engineers Inc., 2019. 页码 438-445 (Proceedings - 2019 IEEE 33rd International Parallel and Distributed Processing Symposium Workshops, IPDPSW 2019).

@inproceedings{b136be08b696415dbd7e88a40059c076,

title = "A GPU inference system scheduling algorithm with asynchronous data transfer",

abstract = "With the rapid expansion of application range, Deep-Learning has increasingly become an indispensable practical method to solve problems in various industries. In different application scenarios, especially in high concurrency areas such as search and recommendation, deep learning inference system is required to have high throughput and low latency, which can not be easily obtained at the same time. In this paper, we build a model to quantify the relationship between concurrency, throughput and job latency. Then we implement a GPU scheduling algorithm for inference jobs in deep learning inference system based on the model. The algorithm predicts the completion time of batch jobs being executed, and reasonably chooses the batch size of the next batch jobs according to the concurrency and upload data to GPU memory ahead of time. So that the system can hide the data transfer delay of GPU and achieve the minimum job latency under the premise of meetingthethroughputrequirements.Experimentsshowthatthe proposed GPU asynchronous data transfer scheduling algorithm improves throughput by 9% compared with the traditional synchronous algorithm, reduces the latency by 3%-76% under different concurrency, and can better suppress the job latency fluctuation caused by concurrency changing.",

keywords = "Deep Learning, GPU, Inference, Latency, Scheduling Algorithm",

author = "Qin Zhang and Li Zha and Xiaohua Wan and Boqun Cheng",

note = "Publisher Copyright: {\textcopyright} 2019 IEEE.; 33rd IEEE International Parallel and Distributed Processing Symposium Workshops, IPDPSW 2019 ; Conference date: 20-05-2019 Through 24-05-2019",

year = "2019",

month = may,

doi = "10.1109/IPDPSW.2019.00083",

language = "English",

series = "Proceedings - 2019 IEEE 33rd International Parallel and Distributed Processing Symposium Workshops, IPDPSW 2019",

publisher = "Institute of Electrical and Electronics Engineers Inc.",

pages = "438--445",

booktitle = "Proceedings - 2019 IEEE 33rd International Parallel and Distributed Processing Symposium Workshops, IPDPSW 2019",

address = "United States",

}

Zhang, Q, Zha, L, Wan, X & Cheng, B 2019, A GPU inference system scheduling algorithm with asynchronous data transfer. 在 Proceedings - 2019 IEEE 33rd International Parallel and Distributed Processing Symposium Workshops, IPDPSW 2019., 8778365, Proceedings - 2019 IEEE 33rd International Parallel and Distributed Processing Symposium Workshops, IPDPSW 2019, Institute of Electrical and Electronics Engineers Inc., 页码 438-445, 33rd IEEE International Parallel and Distributed Processing Symposium Workshops, IPDPSW 2019, Rio de Janeiro, 巴西, 20/05/19. https://doi.org/10.1109/IPDPSW.2019.00083

A GPU inference system scheduling algorithm with asynchronous data transfer. / Zhang, Qin; Zha, Li; Wan, Xiaohua 等.
Proceedings - 2019 IEEE 33rd International Parallel and Distributed Processing Symposium Workshops, IPDPSW 2019. Institute of Electrical and Electronics Engineers Inc., 2019. 页码 438-445 8778365 (Proceedings - 2019 IEEE 33rd International Parallel and Distributed Processing Symposium Workshops, IPDPSW 2019).

科研成果: 书/报告/会议事项章节 › 会议稿件 › 同行评审

TY - GEN

T1 - A GPU inference system scheduling algorithm with asynchronous data transfer

AU - Zhang, Qin

AU - Zha, Li

AU - Wan, Xiaohua

AU - Cheng, Boqun

PY - 2019/5

Y1 - 2019/5

N2 - With the rapid expansion of application range, Deep-Learning has increasingly become an indispensable practical method to solve problems in various industries. In different application scenarios, especially in high concurrency areas such as search and recommendation, deep learning inference system is required to have high throughput and low latency, which can not be easily obtained at the same time. In this paper, we build a model to quantify the relationship between concurrency, throughput and job latency. Then we implement a GPU scheduling algorithm for inference jobs in deep learning inference system based on the model. The algorithm predicts the completion time of batch jobs being executed, and reasonably chooses the batch size of the next batch jobs according to the concurrency and upload data to GPU memory ahead of time. So that the system can hide the data transfer delay of GPU and achieve the minimum job latency under the premise of meetingthethroughputrequirements.Experimentsshowthatthe proposed GPU asynchronous data transfer scheduling algorithm improves throughput by 9% compared with the traditional synchronous algorithm, reduces the latency by 3%-76% under different concurrency, and can better suppress the job latency fluctuation caused by concurrency changing.

AB - With the rapid expansion of application range, Deep-Learning has increasingly become an indispensable practical method to solve problems in various industries. In different application scenarios, especially in high concurrency areas such as search and recommendation, deep learning inference system is required to have high throughput and low latency, which can not be easily obtained at the same time. In this paper, we build a model to quantify the relationship between concurrency, throughput and job latency. Then we implement a GPU scheduling algorithm for inference jobs in deep learning inference system based on the model. The algorithm predicts the completion time of batch jobs being executed, and reasonably chooses the batch size of the next batch jobs according to the concurrency and upload data to GPU memory ahead of time. So that the system can hide the data transfer delay of GPU and achieve the minimum job latency under the premise of meetingthethroughputrequirements.Experimentsshowthatthe proposed GPU asynchronous data transfer scheduling algorithm improves throughput by 9% compared with the traditional synchronous algorithm, reduces the latency by 3%-76% under different concurrency, and can better suppress the job latency fluctuation caused by concurrency changing.

KW - Deep Learning

KW - GPU

KW - Inference

KW - Latency

KW - Scheduling Algorithm

UR - http://www.scopus.com/inward/record.url?scp=85070419330&partnerID=8YFLogxK

U2 - 10.1109/IPDPSW.2019.00083

DO - 10.1109/IPDPSW.2019.00083

M3 - Conference contribution

AN - SCOPUS:85070419330

T3 - Proceedings - 2019 IEEE 33rd International Parallel and Distributed Processing Symposium Workshops, IPDPSW 2019

SP - 438

EP - 445

BT - Proceedings - 2019 IEEE 33rd International Parallel and Distributed Processing Symposium Workshops, IPDPSW 2019

PB - Institute of Electrical and Electronics Engineers Inc.

T2 - 33rd IEEE International Parallel and Distributed Processing Symposium Workshops, IPDPSW 2019

Y2 - 20 May 2019 through 24 May 2019

ER -

Zhang Q, Zha L, Wan X, Cheng B. A GPU inference system scheduling algorithm with asynchronous data transfer. 在 Proceedings - 2019 IEEE 33rd International Parallel and Distributed Processing Symposium Workshops, IPDPSW 2019. Institute of Electrical and Electronics Engineers Inc. 2019. 页码 438-445. 8778365. (Proceedings - 2019 IEEE 33rd International Parallel and Distributed Processing Symposium Workshops, IPDPSW 2019). doi: 10.1109/IPDPSW.2019.00083

A GPU inference system scheduling algorithm with asynchronous data transfer

摘要

出版系列

会议

访问文件

其它文件与链接

指纹

引用此