Abstract
Existing video-language large models (video-LLMs) have made notable progress in understanding global video content but still struggle to capture fine-grained temporal dynamics and accurately align visual events with textual descriptions. In this work, we introduce DAPE-VLLM (Density-Aware Position Encoding enhanced Video-Language Large Model), a time-perception enhanced video grounding framework that integrates density-aware position encoding with boundary perception, to address these limitations and bridge the temporal discrepancy between video and language modalities. We propose a suite of boundary-aware tasks that explicitly model event durations, locations, and transitions, enabling large language models to better understand temporal structure within videos. In addition, we develop a novel density-aware temporal encoding mechanism that dynamically adapts position representations to varying temporal information density, enhancing the model’s ability to localize events precisely across diverse video content. Extensive experiments across ActivityNet, Charades, and DiDeMo datasets demonstrate the effectiveness of our approach, yielding consistent performance improvements and achieving up to 11.2% relative gain in R@0.3. These results validate the importance of integrating density-aware modeling and boundary perception to advance fine-grained video grounding capabilities in large language models.
| Original language | English |
|---|---|
| Journal | IEEE Transactions on Circuits and Systems for Video Technology |
| DOIs | |
| Publication status | Accepted/In press - 2026 |
| Externally published | Yes |
Fingerprint
Dive into the research topics of 'Time-Perception Enhanced Video Grounding Framework with Density-Aware Position Encoding'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver