Abstract
Retrieval-augmented Generation (RAG) has been proven to expand knowledge boundaries of large language models (LLMs). However, indiscriminate retrieval throughout the generation process can increase costs and lead to potentially incorrect answers. To identify the optimal retrieval timing within the generation process, we introduce a novel framework named SC-DRAG, which employs sentence-level confidence estimation. Based on the estimated confidence of each reasoning sentence, SC DRAG adopts a set of adaptive strategies that correspond to different confidence levels, including retrieval, reflection and progression. A key aspect of our framework is the accurate estimation of confidence, which is enhanced by considering not only the generated responses but also the semantics of the query, thereby improving confidence evaluation. This dual consideration ensures that the retrieval process is activated only when truly beneficial, preventing unnecessary retrievals that may introduce noise into the generation. Furthermore, SC-DRAG provides a flexible solution that adapts to the evolving demands of the generation process, enhancing its robustness in complex knowledge-intensive tasks. Compared to previous works, our framework exhibits exceptional adaptability across a variety of white-box and black-box models. Experimental results on three knowledge-intensive datasets show that the proposed framework achieves favorable accuracy-efficiency trade-offs and competitive performance across the evaluated settings by improving retrieval timing. Code will be avaliable at https://gitub.com/yifraternity/conf-rag.git.
| Original language | English |
|---|---|
| Article number | 133921 |
| Journal | Expert Systems with Applications |
| Volume | 333 |
| DOIs | |
| Publication status | Published - 1 Jan 2026 |
Keywords
- Confidence
- Large language model
- Query-response
- Retrieval timing
- Retrieval-augmented generation
Fingerprint
Dive into the research topics of 'Retrieval is NOT always needed: Exploring the timing of retrieval in dynamic retrieval-augmented generation'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver