跳到主要导航 跳到搜索 跳到主要内容

FLASHBACK: Efficient Retrieval-Augmented Language Modeling for Fast Inference

  • Beijing Institute of Technology

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

Retrieval-Augmented Language Modeling (RALM) by integrating large language models (LLM) with relevant documents from an external corpus is a proven methodology for enabling the LLM to generate information beyond the scope of its pre-training corpus. Previous work by retrieving a set of tokens iteratively with retrieved content prepending to the input poses a high run-time issue, which degrades the inference efficiency of the LLMs because they fail to use the Key-Value (KV) cache efficiently. We propose FLASHBACK, a modular RALM designed to improve the inference efficiency of RALM with the appending context pattern while maintaining decent performance after fine-tuning by Low-Rank Adaption. FLASHBACK appends retrieved documents at the end of the context to efficiently utilize the KV cache. We also introduce the Marking Token as two special prompt tokens for marking the appending context during fine-tuning. Our experiments show that FLASHBACK can improve language modeling performance in the perplexity metric. We proved that the Marking Token is a usable add-on when fine-tuning models on specific context patterns. By bypassing unnecessary recomputation, FLASHBACK achieves fast inference speed with long context input. The inference speed is up to 4× faster than the prepending counterpart on a 7B LLM (Llama 2) in the runtime test.

源语言英语
主期刊名Findings of the Association for Computational Linguistics
主期刊副标题ACL 2025
编辑Wanxiang Che, Joyce Nabende, Ekaterina Shutova, Mohammad Taher Pilehvar
出版商Association for Computational Linguistics (ACL)
595-608
页数14
ISBN(电子版)9798891762565
DOI
出版状态已出版 - 2025
已对外发布
活动63rd Annual Meeting of the Association for Computational Linguistics, ACL 2025 - Vienna, 奥地利
期限: 27 7月 20251 8月 2025

出版系列

姓名Proceedings of the Annual Meeting of the Association for Computational Linguistics
ISSN(印刷版)0736-587X

会议

会议63rd Annual Meeting of the Association for Computational Linguistics, ACL 2025
国家/地区奥地利
Vienna
时期27/07/251/08/25

指纹

探究 'FLASHBACK: Efficient Retrieval-Augmented Language Modeling for Fast Inference' 的科研主题。它们共同构成独一无二的指纹。

引用此