跳到主要导航 跳到搜索 跳到主要内容

SSTFormer: Bridging Spiking Neural Network and Memory Support Transformer for Frame-Event Based-Recognition

  • Xiao Wang
  • , Yao Rong
  • , Zongzhen Wu
  • , Lin Zhu
  • , Bo Jiang*
  • , Jin Tang
  • , Yonghong Tian
  • *此作品的通讯作者
  • School of Computer Science and Technology, Anhui University
  • Peng Cheng Laboratory
  • Peking University

科研成果: 期刊稿件文章同行评审

摘要

Event camera-based pattern recognition is a newly arising research topic in recent years. Current researchers usually transform the event streams into images, graphs, or voxels, and adopt deep neural networks for event-based classification. Although good performance can be achieved on simple event recognition datasets, however, their results may still be limited due to the following two issues. First, they adopt spatial sparse event streams for recognition only, which may fail to capture the color and detailed texture information well. Second, they adopt either spiking neural networks (SNN) for energy-efficient recognition with suboptimal results, or artificial neural networks (ANN) for energy-intensive, high-performance recognition. However, few of them consider achieving a balance between these two aspects. In this article, we formally propose to recognize patterns by fusing RGB frames and event streams simultaneously and propose a new RGB frame-event recognition framework to address the aforementioned issues. The proposed method contains four main modules, i.e., memory support Transformer network for RGB frame encoding, spiking neural network for raw event stream encoding, multimodal bottleneck fusion module for RGB-Event feature aggregation, and prediction head. Due to the scarcity of RGB-Event based classification dataset, we also propose a large-scale PokerEvent dataset which contains 114 classes, and 27 102 frame-event pairs recorded using a DVS346 event camera. Extensive experiments on two RGB-event based classification datasets fully validated the effectiveness of our proposed framework. We hope this work will boost the development of pattern recognition by fusing RGB frames and event streams. Both our dataset and source code of this work will be released at https://github.com/Event-AHU/SSTFormer.

源语言英语
页(从-至)1488-1502
页数15
期刊IEEE Transactions on Cognitive and Developmental Systems
17
6
DOI
出版状态已出版 - 2025

学术指纹

探究 'SSTFormer: Bridging Spiking Neural Network and Memory Support Transformer for Frame-Event Based-Recognition' 的科研主题。它们共同构成独一无二的学术指纹。

引用此