跳到主要导航 跳到搜索 跳到主要内容

How Speculative Can Speculative Decoding Be?

  • Zhuorui Liu
  • , Chen Zhang
  • , Dawei Song*
  • *此作品的通讯作者
  • Beijing Institute of Technology
  • Open University Milton Keynes

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

Large language models (LLMs) have drawn great attention from the field of natural language processing and beyond, due to their impressive capability of autoregressive modeling, yet bringing an obvious problem, i.e., the largely increased latency. An emerging idea to alleviate this problem is speculative decoding, which first uses a draft model to draft tokens autoregressively and then makes the target model verify these tokens in parallel. The draft model is typically smaller than the target model, and it essentially trades generation quality for speed. Thereby, speculative decoding can be viewed as a speculative game for the target model in term of verification failures. That is, the lengthy draft tokens proposed by the small draft models could fail in the verification stage. Naturally, a critical question arises: how speculative can speculative decoding be, or in other words, how small can an adequate draft model be and how large can an appropriate number of draft tokens be? This work aims to investigate these questions and demonstrate how the scale of the draft model and the number of draft tokens would have an impact on the overall latency of the speculative decoding. We theoretically show that neither of above two factors will be infinitely speculative. Namely, there is a certain turning point for each of them. We then empirically show that the scale of the draft model could be 10-20× smaller than the target model and the optimal number of draft tokens should lie in 3-5.

源语言英语
主期刊名2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation, LREC-COLING 2024 - Main Conference Proceedings
编辑Nicoletta Calzolari, Min-Yen Kan, Veronique Hoste, Alessandro Lenci, Sakriani Sakti, Nianwen Xue
出版商European Language Resources Association (ELRA)
8265-8275
页数11
ISBN(电子版)9782493814104
出版状态已出版 - 2024
活动Joint 30th International Conference on Computational Linguistics and 14th International Conference on Language Resources and Evaluation, LREC-COLING 2024 - Hybrid, Torino, 意大利
期限: 20 5月 202425 5月 2024

丛书

姓名2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation, LREC-COLING 2024 - Main Conference Proceedings

会议

会议Joint 30th International Conference on Computational Linguistics and 14th International Conference on Language Resources and Evaluation, LREC-COLING 2024
国家/地区意大利
Hybrid, Torino
时期20/05/2425/05/24

学术指纹

探究 'How Speculative Can Speculative Decoding Be?' 的科研主题。它们共同构成独一无二的学术指纹。

引用此