Skip to main navigation Skip to search Skip to main content

MoVi: Real-Time Large Multimodal Model-Driven Interactive Video Analytics on Mobile Devices

  • Zekai Li
  • , Xiaoyi Fan
  • , Xiping Hu
  • , Yifei Zhu*
  • *Corresponding author for this work
  • Shanghai Jiao Tong University
  • Shenzhen MSU-BIT University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Large multimodal models (LMMs) are transforming traditional mobile video analytics into interactive services, where wearable cameras stream live videos and users pose free-form queries and receive responses in natural language. Due to the high deployment costs, LMMs are primarily deployed on cloud servers. Real-time video streaming under dynamic network conditions and the afterward inference thus significantly affect quality of experience (QoE). However, existing video analytics systems are designed for task-specific, frame-independent, single-modal tasks, making them unsuitable for interaction with LMMs over multimodal, contextual inputs. Even worse, the autoregressive decoding of LMMs delays visual token processing during text generation, affecting real-time response. To bridge these gaps, we present MoVi, the first collaborative system for real-time LMM-driven interactive video analytics over dynamic networks. MoVi first establishes an interaction-oriented QoE model tailored to emerging interactive video analytics applications based on real-world user studies. It then jointly designs the video streaming and inference stages for QoE optimization. Specifically, MoVi employs a causal-aware streaming controller to adapt video configurations under dynamic networks. A query-assisted token manager further reduces response latency by dynamically pruning buffered tokens. Extensive experiments on real-world datasets and user studies demonstrate that MoVi achieves 40.1% higher QoE and 38.9% higher user opinion scores than existing state-of-the-art systems.

Original languageEnglish
Title of host publicationINFOCOM 2026 - IEEE Conference on Computer Communications
PublisherInstitute of Electrical and Electronics Engineers Inc.
ISBN (Electronic)9798331549619
DOIs
Publication statusPublished - 2026
Event2026 IEEE Conference on Computer Communications, INFOCOM 2026 - Tokyo, Japan
Duration: 18 May 202621 May 2026

Publication series

NameProceedings - IEEE INFOCOM
ISSN (Print)0743-166X

Conference

Conference2026 IEEE Conference on Computer Communications, INFOCOM 2026
Country/TerritoryJapan
CityTokyo
Period18/05/2621/05/26

Keywords

  • edge-cloud collaboration
  • large multimodal model
  • video analytics
  • video streaming

Fingerprint

Dive into the research topics of 'MoVi: Real-Time Large Multimodal Model-Driven Interactive Video Analytics on Mobile Devices'. Together they form a unique fingerprint.

Cite this