HAFNet: Hierarchical Attentive Fusion Network for Multispectral Pedestrian Detection

Peiran Peng; Tingfa Xu; Bo Huang; Jianan Li

doi:10.3390/rs15082041

HAFNet: Hierarchical Attentive Fusion Network for Multispectral Pedestrian Detection

Peiran Peng, Tingfa Xu, Bo Huang, Jianan Li^*

^*此作品的通讯作者

光电学院

科研成果: 期刊稿件 › 文章 › 同行评审

7 引用（Scopus）

摘要

Multispectral pedestrian detection via visible and thermal image pairs has received widespread attention in recent years. It provides a promising multi-modality solution to address the challenges of pedestrian detection in low-light environments and occlusion situations. Most existing methods directly blend the results of the two modalities or combine the visible and thermal features via a linear interpolation. However, such fusion strategies tend to extract coarser features corresponding to the positions of different modalities, which may lead to degraded detection performance. To mitigate this, this paper proposes a novel and adaptive cross-modality fusion framework, named Hierarchical Attentive Fusion Network (HAFNet), which fully exploits the multispectral attention knowledge to inspire pedestrian detection in the decision-making process. Concretely, we introduce a Hierarchical Content-dependent Attentive Fusion (HCAF) module to extract top-level features as a guide to pixel-wise blending features of two modalities to enhance the quality of the feature representation and a plug-in multi-modality feature alignment (MFA) block to fine-tune the feature alignment of two modalities. Experiments on the challenging KAIST and CVC-14 datasets demonstrate the superior performance of our method with satisfactory speed.

源语言	英语
文章编号	2041
期刊	Remote Sensing
卷	15
期	8
DOI	https://doi.org/10.3390/rs15082041
出版状态	已出版 - 4月 2023

访问文件

10.3390/rs15082041

其它文件与链接

链接到 Scopus 的出版物

引用此

Peng, P., Xu, T., Huang, B., & Li, J. (2023). HAFNet: Hierarchical Attentive Fusion Network for Multispectral Pedestrian Detection. Remote Sensing, 15(8), 文章 2041. https://doi.org/10.3390/rs15082041

@article{8280eeddb0874a94a50526ea9f6d95c7,

title = "HAFNet: Hierarchical Attentive Fusion Network for Multispectral Pedestrian Detection",

abstract = "Multispectral pedestrian detection via visible and thermal image pairs has received widespread attention in recent years. It provides a promising multi-modality solution to address the challenges of pedestrian detection in low-light environments and occlusion situations. Most existing methods directly blend the results of the two modalities or combine the visible and thermal features via a linear interpolation. However, such fusion strategies tend to extract coarser features corresponding to the positions of different modalities, which may lead to degraded detection performance. To mitigate this, this paper proposes a novel and adaptive cross-modality fusion framework, named Hierarchical Attentive Fusion Network (HAFNet), which fully exploits the multispectral attention knowledge to inspire pedestrian detection in the decision-making process. Concretely, we introduce a Hierarchical Content-dependent Attentive Fusion (HCAF) module to extract top-level features as a guide to pixel-wise blending features of two modalities to enhance the quality of the feature representation and a plug-in multi-modality feature alignment (MFA) block to fine-tune the feature alignment of two modalities. Experiments on the challenging KAIST and CVC-14 datasets demonstrate the superior performance of our method with satisfactory speed.",

keywords = "content-dependent, feature alignment, multispectral pedestrian detection",

author = "Peiran Peng and Tingfa Xu and Bo Huang and Jianan Li",

note = "Publisher Copyright: {\textcopyright} 2023 by the authors.",

year = "2023",

month = apr,

doi = "10.3390/rs15082041",

language = "English",

volume = "15",

journal = "Remote Sensing",

issn = "2072-4292",

publisher = "Multidisciplinary Digital Publishing Institute (MDPI)",

number = "8",

}

TY - JOUR

T1 - HAFNet

T2 - Hierarchical Attentive Fusion Network for Multispectral Pedestrian Detection

AU - Peng, Peiran

AU - Xu, Tingfa

AU - Huang, Bo

AU - Li, Jianan

PY - 2023/4

Y1 - 2023/4

N2 - Multispectral pedestrian detection via visible and thermal image pairs has received widespread attention in recent years. It provides a promising multi-modality solution to address the challenges of pedestrian detection in low-light environments and occlusion situations. Most existing methods directly blend the results of the two modalities or combine the visible and thermal features via a linear interpolation. However, such fusion strategies tend to extract coarser features corresponding to the positions of different modalities, which may lead to degraded detection performance. To mitigate this, this paper proposes a novel and adaptive cross-modality fusion framework, named Hierarchical Attentive Fusion Network (HAFNet), which fully exploits the multispectral attention knowledge to inspire pedestrian detection in the decision-making process. Concretely, we introduce a Hierarchical Content-dependent Attentive Fusion (HCAF) module to extract top-level features as a guide to pixel-wise blending features of two modalities to enhance the quality of the feature representation and a plug-in multi-modality feature alignment (MFA) block to fine-tune the feature alignment of two modalities. Experiments on the challenging KAIST and CVC-14 datasets demonstrate the superior performance of our method with satisfactory speed.

AB - Multispectral pedestrian detection via visible and thermal image pairs has received widespread attention in recent years. It provides a promising multi-modality solution to address the challenges of pedestrian detection in low-light environments and occlusion situations. Most existing methods directly blend the results of the two modalities or combine the visible and thermal features via a linear interpolation. However, such fusion strategies tend to extract coarser features corresponding to the positions of different modalities, which may lead to degraded detection performance. To mitigate this, this paper proposes a novel and adaptive cross-modality fusion framework, named Hierarchical Attentive Fusion Network (HAFNet), which fully exploits the multispectral attention knowledge to inspire pedestrian detection in the decision-making process. Concretely, we introduce a Hierarchical Content-dependent Attentive Fusion (HCAF) module to extract top-level features as a guide to pixel-wise blending features of two modalities to enhance the quality of the feature representation and a plug-in multi-modality feature alignment (MFA) block to fine-tune the feature alignment of two modalities. Experiments on the challenging KAIST and CVC-14 datasets demonstrate the superior performance of our method with satisfactory speed.

KW - content-dependent

KW - feature alignment

KW - multispectral pedestrian detection

UR - http://www.scopus.com/inward/record.url?scp=85156098772&partnerID=8YFLogxK

U2 - 10.3390/rs15082041

DO - 10.3390/rs15082041

M3 - Article

AN - SCOPUS:85156098772

SN - 2072-4292

VL - 15

JO - Remote Sensing

JF - Remote Sensing

IS - 8

M1 - 2041

ER -

HAFNet: Hierarchical Attentive Fusion Network for Multispectral Pedestrian Detection

摘要

访问文件

其它文件与链接

指纹

引用此