跳到主要导航 跳到搜索 跳到主要内容

MMTL-UniAD: A Unified Framework for Multimodal and Multi-Task Learning in Assistive Driving Perception

  • Wenzhuo Liu
  • , Wenshuo Wang*
  • , Yicheng Qiao
  • , Qiannan Guo
  • , Jiayin Zhu
  • , Pengfei Li
  • , Zilong Chen
  • , Huiming Yang
  • , Zhiwei Li
  • , Lening Wang
  • , Tiao Tan
  • , Huaping Liu
  • *此作品的通讯作者
  • Beijing Institute of Technology
  • Tsinghua University
  • The Hong Kong University of Science and Technology (Guangzhou)
  • Beijing University of Chemical Technology
  • Beihang University

科研成果: 期刊稿件会议文章同行评审

摘要

Advanced driver assistance systems require a comprehensive understanding of the driver's mental/physical state and traffic context but existing works often neglect the potential benefits of joint learning between these tasks. This paper proposes MMTL-UniAD, a unified multi-modal multitask learning framework that simultaneously recognizes driver behavior (e.g., looking around, talking), driver emotion (e.g., anxiety, happiness), vehicle behavior (e.g., parking, turning), and traffic context (e.g., traffic jam, traffic smooth). A key challenge is avoiding negative transfer between tasks, which can impair learning performance. To address this, we introduce two key components into the framework: one is the multi-axis region attention network to extract global context-sensitive features, and the other is the dual-branch multimodal embedding to learn multi-modal embeddings from both task-shared and task-specific features. The former uses a multi-attention mechanism to extract task-relevant features, mitigating negative transfer caused by task-unrelated features. The latter employs a dual-branch structure to adaptively adjust task-shared and task-specific parameters, enhancing cross-task knowledge transfer while reducing task conflicts. We assess MMTL-UniAD on the AIDE dataset, using a series of ablation studies, and show that it outperforms state-of-the-art methods across all four tasks. The code is available on https://github.com/Wenzhuo-Liu/MMTL-UniAD.

源语言英语
页(从-至)6864-6874
页数11
期刊Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition
DOI
出版状态已出版 - 2025
活动2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2025 - Nashville, 美国
期限: 11 6月 202515 6月 2025

学术指纹

探究 'MMTL-UniAD: A Unified Framework for Multimodal and Multi-Task Learning in Assistive Driving Perception' 的科研主题。它们共同构成独一无二的学术指纹。

引用此