Skip to main navigation Skip to search Skip to main content

Time-Perception Enhanced Video Grounding Framework with Density-Aware Position Encoding

  • Beijing Institute of Technology
  • Beijing University of Technology

Research output: Contribution to journalArticlepeer-review

Abstract

Existing video-language large models (video-LLMs) have made notable progress in understanding global video content but still struggle to capture fine-grained temporal dynamics and accurately align visual events with textual descriptions. In this work, we introduce DAPE-VLLM (Density-Aware Position Encoding enhanced Video-Language Large Model), a time-perception enhanced video grounding framework that integrates density-aware position encoding with boundary perception, to address these limitations and bridge the temporal discrepancy between video and language modalities. We propose a suite of boundary-aware tasks that explicitly model event durations, locations, and transitions, enabling large language models to better understand temporal structure within videos. In addition, we develop a novel density-aware temporal encoding mechanism that dynamically adapts position representations to varying temporal information density, enhancing the model’s ability to localize events precisely across diverse video content. Extensive experiments across ActivityNet, Charades, and DiDeMo datasets demonstrate the effectiveness of our approach, yielding consistent performance improvements and achieving up to 11.2% relative gain in R@0.3. These results validate the importance of integrating density-aware modeling and boundary perception to advance fine-grained video grounding capabilities in large language models.

Original languageEnglish
JournalIEEE Transactions on Circuits and Systems for Video Technology
DOIs
Publication statusAccepted/In press - 2026
Externally publishedYes

Fingerprint

Dive into the research topics of 'Time-Perception Enhanced Video Grounding Framework with Density-Aware Position Encoding'. Together they form a unique fingerprint.

Cite this