TY - JOUR
T1 - Retrieval is NOT always needed
T2 - Exploring the timing of retrieval in dynamic retrieval-augmented generation
AU - Liu, Yuhang
AU - Huang, Heyan
AU - Yang, Yizhe
AU - Zeng, Zhizhuo
AU - Zhou, Youchao
AU - Wu, Zhijing
AU - Gao, Yang
N1 - Publisher Copyright:
Elsevier Ltd
PY - 2026/1/1
Y1 - 2026/1/1
N2 - Retrieval-augmented Generation (RAG) has been proven to expand knowledge boundaries of large language models (LLMs). However, indiscriminate retrieval throughout the generation process can increase costs and lead to potentially incorrect answers. To identify the optimal retrieval timing within the generation process, we introduce a novel framework named SC-DRAG, which employs sentence-level confidence estimation. Based on the estimated confidence of each reasoning sentence, SC DRAG adopts a set of adaptive strategies that correspond to different confidence levels, including retrieval, reflection and progression. A key aspect of our framework is the accurate estimation of confidence, which is enhanced by considering not only the generated responses but also the semantics of the query, thereby improving confidence evaluation. This dual consideration ensures that the retrieval process is activated only when truly beneficial, preventing unnecessary retrievals that may introduce noise into the generation. Furthermore, SC-DRAG provides a flexible solution that adapts to the evolving demands of the generation process, enhancing its robustness in complex knowledge-intensive tasks. Compared to previous works, our framework exhibits exceptional adaptability across a variety of white-box and black-box models. Experimental results on three knowledge-intensive datasets show that the proposed framework achieves favorable accuracy-efficiency trade-offs and competitive performance across the evaluated settings by improving retrieval timing. Code will be avaliable at https://gitub.com/yifraternity/conf-rag.git.
AB - Retrieval-augmented Generation (RAG) has been proven to expand knowledge boundaries of large language models (LLMs). However, indiscriminate retrieval throughout the generation process can increase costs and lead to potentially incorrect answers. To identify the optimal retrieval timing within the generation process, we introduce a novel framework named SC-DRAG, which employs sentence-level confidence estimation. Based on the estimated confidence of each reasoning sentence, SC DRAG adopts a set of adaptive strategies that correspond to different confidence levels, including retrieval, reflection and progression. A key aspect of our framework is the accurate estimation of confidence, which is enhanced by considering not only the generated responses but also the semantics of the query, thereby improving confidence evaluation. This dual consideration ensures that the retrieval process is activated only when truly beneficial, preventing unnecessary retrievals that may introduce noise into the generation. Furthermore, SC-DRAG provides a flexible solution that adapts to the evolving demands of the generation process, enhancing its robustness in complex knowledge-intensive tasks. Compared to previous works, our framework exhibits exceptional adaptability across a variety of white-box and black-box models. Experimental results on three knowledge-intensive datasets show that the proposed framework achieves favorable accuracy-efficiency trade-offs and competitive performance across the evaluated settings by improving retrieval timing. Code will be avaliable at https://gitub.com/yifraternity/conf-rag.git.
KW - Confidence
KW - Large language model
KW - Query-response
KW - Retrieval timing
KW - Retrieval-augmented generation
UR - https://www.scopus.com/pages/publications/105047282475
U2 - 10.1016/j.eswa.2026.133921
DO - 10.1016/j.eswa.2026.133921
M3 - Article
AN - SCOPUS:105047282475
SN - 0957-4174
VL - 333
JO - Expert Systems with Applications
JF - Expert Systems with Applications
M1 - 133921
ER -