TY - JOUR
T1 - Time-Perception Enhanced Video Grounding Framework with Density-Aware Position Encoding
AU - Wang, Bo
AU - Huang, Heyan
AU - Li, Xuefen
AU - Shi, Ge
AU - Teng, Jiahao
AU - Feng, Chong
N1 - Publisher Copyright:
© 1991-2012 IEEE.
PY - 2026
Y1 - 2026
N2 - Existing video-language large models (video-LLMs) have made notable progress in understanding global video content but still struggle to capture fine-grained temporal dynamics and accurately align visual events with textual descriptions. In this work, we introduce DAPE-VLLM (Density-Aware Position Encoding enhanced Video-Language Large Model), a time-perception enhanced video grounding framework that integrates density-aware position encoding with boundary perception, to address these limitations and bridge the temporal discrepancy between video and language modalities. We propose a suite of boundary-aware tasks that explicitly model event durations, locations, and transitions, enabling large language models to better understand temporal structure within videos. In addition, we develop a novel density-aware temporal encoding mechanism that dynamically adapts position representations to varying temporal information density, enhancing the model’s ability to localize events precisely across diverse video content. Extensive experiments across ActivityNet, Charades, and DiDeMo datasets demonstrate the effectiveness of our approach, yielding consistent performance improvements and achieving up to 11.2% relative gain in R@0.3. These results validate the importance of integrating density-aware modeling and boundary perception to advance fine-grained video grounding capabilities in large language models.
AB - Existing video-language large models (video-LLMs) have made notable progress in understanding global video content but still struggle to capture fine-grained temporal dynamics and accurately align visual events with textual descriptions. In this work, we introduce DAPE-VLLM (Density-Aware Position Encoding enhanced Video-Language Large Model), a time-perception enhanced video grounding framework that integrates density-aware position encoding with boundary perception, to address these limitations and bridge the temporal discrepancy between video and language modalities. We propose a suite of boundary-aware tasks that explicitly model event durations, locations, and transitions, enabling large language models to better understand temporal structure within videos. In addition, we develop a novel density-aware temporal encoding mechanism that dynamically adapts position representations to varying temporal information density, enhancing the model’s ability to localize events precisely across diverse video content. Extensive experiments across ActivityNet, Charades, and DiDeMo datasets demonstrate the effectiveness of our approach, yielding consistent performance improvements and achieving up to 11.2% relative gain in R@0.3. These results validate the importance of integrating density-aware modeling and boundary perception to advance fine-grained video grounding capabilities in large language models.
UR - https://www.scopus.com/pages/publications/105043680574
U2 - 10.1109/TCSVT.2026.3708456
DO - 10.1109/TCSVT.2026.3708456
M3 - Article
AN - SCOPUS:105043680574
SN - 1051-8215
JO - IEEE Transactions on Circuits and Systems for Video Technology
JF - IEEE Transactions on Circuits and Systems for Video Technology
ER -