TY - JOUR
T1 - Towards Efficient Coarse-grained Dialogue Response Selection
AU - Lan, Tian
AU - Mao, Xian Ling
AU - Wei, Wei
AU - Gao, Xiaoyan
AU - Huang, Heyan
N1 - Publisher Copyright:
© 2023 Association for Computing Machinery. All rights reserved.
PY - 2023/9/27
Y1 - 2023/9/27
N2 - oarse-grained response selection is a fundamental and essential subsystem for the widely used retrievalbased chatbots, aiming to recall a coarse-grained candidate set from a large-scale dataset. The dense retrievaltechnique has recently been proven very effective in building such a subsystem. However, dialogue denseretrieval models face two problems in real scenarios: (1) the multi-turn dialogue history is re-computed ineach turn, leading to inefficient inference; (2) the index storage of the offline index is enormous, significantlyincreasing the deployment cost. To address these problems, we propose an efficient coarse-grained responseselection subsystem consisting of two novel methods. Specifically, to address the first problem, we proposethe Hierarchical Dense Retrieval. It caches rich multi-vector representations of the dialogue history and onlyencodes the latest user’s utterance, leading to better inference efficiency. Then, to address the second problem,we design the Deep Semantic Hashing to reduce the index storage while effectively saving its recall accuracynotably. Extensive experimental results prove the advantages of the two proposed methods over previousworks. Specifically, with the limited performance loss, our proposed coarse-grained response selection modelachieves over 5x FLOPs speedup and over 192x storage compression ratio. Moreover, our source codes havebeen publicly released.
AB - oarse-grained response selection is a fundamental and essential subsystem for the widely used retrievalbased chatbots, aiming to recall a coarse-grained candidate set from a large-scale dataset. The dense retrievaltechnique has recently been proven very effective in building such a subsystem. However, dialogue denseretrieval models face two problems in real scenarios: (1) the multi-turn dialogue history is re-computed ineach turn, leading to inefficient inference; (2) the index storage of the offline index is enormous, significantlyincreasing the deployment cost. To address these problems, we propose an efficient coarse-grained responseselection subsystem consisting of two novel methods. Specifically, to address the first problem, we proposethe Hierarchical Dense Retrieval. It caches rich multi-vector representations of the dialogue history and onlyencodes the latest user’s utterance, leading to better inference efficiency. Then, to address the second problem,we design the Deep Semantic Hashing to reduce the index storage while effectively saving its recall accuracynotably. Extensive experimental results prove the advantages of the two proposed methods over previousworks. Specifically, with the limited performance loss, our proposed coarse-grained response selection modelachieves over 5x FLOPs speedup and over 192x storage compression ratio. Moreover, our source codes havebeen publicly released.
KW - Retrieval-based dialogue system
KW - deep semantic hashing
KW - dense retrieval
UR - https://www.scopus.com/pages/publications/85181538476
U2 - 10.1145/3597609
DO - 10.1145/3597609
M3 - Article
AN - SCOPUS:85181538476
SN - 1046-8188
VL - 42
JO - ACM Transactions on Information Systems
JF - ACM Transactions on Information Systems
IS - 2
M1 - 35
ER -