Hierarchical collaboration for referring image segmentation

Wei Zhang; Zesen Cheng; Jie Chen; Wen Gao

doi:10.1016/j.neucom.2024.128632

Hierarchical collaboration for referring image segmentation

Wei Zhang, Zesen Cheng, Jie Chen, Wen Gao^*

^*此作品的通讯作者

科研成果: 期刊稿件 › 文章 › 同行评审

摘要

In the field of referring segmentation, top-down methods and bottom-up methods are the two prevailing approaches. Both of these methods inevitably exhibit certain drawbacks. Top-down methods are susceptible to Polar Negative (PN) errors due to their limited understanding of multi-modal fine-grained features. Bottom-up methods lack macro-level object positional information, making them susceptible to Inferior Positive (IP) errors. However, we find that the two approaches are highly complementary in addressing their respective weaknesses, but combining them directly through a simple average does not yield complementary advantages. Therefore, we proposed a hierarchical collaboration approach to explore the complementary characteristics of the existing two methods from the perspectives of fusion and interaction, aiming to achieve more precise segmentation results. We proposed the Complementary Feature Interaction (CFI) module, which enables top-down methods to access fine-grained information and allows bottom-up approaches to obtain object positional information interactively. Regarding integration, Gaussian Scoring Integration (GSI) models the Gaussian performance distributions of two branches and performs weighted integration by sampling confidence scores from these distributions. We integrate various top-down and bottom-up methods within the proposed architecture and conduct experiments on three standard datasets. The experimental results demonstrate that our method outperforms the state-of-the-art independent segmentation algorithms. On the RefCOCO validation, test A and test B datasets, our proposed method achieved IoU scores of 77.51, 79.12, and 72.79, respectively. Extensive experiments demonstrate that our method can significantly improve segmentation accuracy when fusing different sub-methods.

源语言	英语
文章编号	128632
期刊	Neurocomputing
卷	613
DOI	https://doi.org/10.1016/j.neucom.2024.128632
出版状态	已出版 - 14 1月 2025
已对外发布	是

访问文件

10.1016/j.neucom.2024.128632

其它文件与链接

链接到 Scopus 的出版物

引用此

@article{601aaa5c40634f7fb96d64f5550f1b20,

title = "Hierarchical collaboration for referring image segmentation",

abstract = "In the field of referring segmentation, top-down methods and bottom-up methods are the two prevailing approaches. Both of these methods inevitably exhibit certain drawbacks. Top-down methods are susceptible to Polar Negative (PN) errors due to their limited understanding of multi-modal fine-grained features. Bottom-up methods lack macro-level object positional information, making them susceptible to Inferior Positive (IP) errors. However, we find that the two approaches are highly complementary in addressing their respective weaknesses, but combining them directly through a simple average does not yield complementary advantages. Therefore, we proposed a hierarchical collaboration approach to explore the complementary characteristics of the existing two methods from the perspectives of fusion and interaction, aiming to achieve more precise segmentation results. We proposed the Complementary Feature Interaction (CFI) module, which enables top-down methods to access fine-grained information and allows bottom-up approaches to obtain object positional information interactively. Regarding integration, Gaussian Scoring Integration (GSI) models the Gaussian performance distributions of two branches and performs weighted integration by sampling confidence scores from these distributions. We integrate various top-down and bottom-up methods within the proposed architecture and conduct experiments on three standard datasets. The experimental results demonstrate that our method outperforms the state-of-the-art independent segmentation algorithms. On the RefCOCO validation, test A and test B datasets, our proposed method achieved IoU scores of 77.51, 79.12, and 72.79, respectively. Extensive experiments demonstrate that our method can significantly improve segmentation accuracy when fusing different sub-methods.",

keywords = "Cross-modal, Image understanding, Referring image segmentation",

author = "Wei Zhang and Zesen Cheng and Jie Chen and Wen Gao",

note = "Publisher Copyright: {\textcopyright} 2024 Elsevier B.V.",

year = "2025",

month = jan,

day = "14",

doi = "10.1016/j.neucom.2024.128632",

language = "English",

volume = "613",

journal = "Neurocomputing",

issn = "0925-2312",

publisher = "Elsevier B.V.",

}

TY - JOUR

T1 - Hierarchical collaboration for referring image segmentation

AU - Zhang, Wei

AU - Cheng, Zesen

AU - Chen, Jie

AU - Gao, Wen

PY - 2025/1/14

Y1 - 2025/1/14

N2 - In the field of referring segmentation, top-down methods and bottom-up methods are the two prevailing approaches. Both of these methods inevitably exhibit certain drawbacks. Top-down methods are susceptible to Polar Negative (PN) errors due to their limited understanding of multi-modal fine-grained features. Bottom-up methods lack macro-level object positional information, making them susceptible to Inferior Positive (IP) errors. However, we find that the two approaches are highly complementary in addressing their respective weaknesses, but combining them directly through a simple average does not yield complementary advantages. Therefore, we proposed a hierarchical collaboration approach to explore the complementary characteristics of the existing two methods from the perspectives of fusion and interaction, aiming to achieve more precise segmentation results. We proposed the Complementary Feature Interaction (CFI) module, which enables top-down methods to access fine-grained information and allows bottom-up approaches to obtain object positional information interactively. Regarding integration, Gaussian Scoring Integration (GSI) models the Gaussian performance distributions of two branches and performs weighted integration by sampling confidence scores from these distributions. We integrate various top-down and bottom-up methods within the proposed architecture and conduct experiments on three standard datasets. The experimental results demonstrate that our method outperforms the state-of-the-art independent segmentation algorithms. On the RefCOCO validation, test A and test B datasets, our proposed method achieved IoU scores of 77.51, 79.12, and 72.79, respectively. Extensive experiments demonstrate that our method can significantly improve segmentation accuracy when fusing different sub-methods.

AB - In the field of referring segmentation, top-down methods and bottom-up methods are the two prevailing approaches. Both of these methods inevitably exhibit certain drawbacks. Top-down methods are susceptible to Polar Negative (PN) errors due to their limited understanding of multi-modal fine-grained features. Bottom-up methods lack macro-level object positional information, making them susceptible to Inferior Positive (IP) errors. However, we find that the two approaches are highly complementary in addressing their respective weaknesses, but combining them directly through a simple average does not yield complementary advantages. Therefore, we proposed a hierarchical collaboration approach to explore the complementary characteristics of the existing two methods from the perspectives of fusion and interaction, aiming to achieve more precise segmentation results. We proposed the Complementary Feature Interaction (CFI) module, which enables top-down methods to access fine-grained information and allows bottom-up approaches to obtain object positional information interactively. Regarding integration, Gaussian Scoring Integration (GSI) models the Gaussian performance distributions of two branches and performs weighted integration by sampling confidence scores from these distributions. We integrate various top-down and bottom-up methods within the proposed architecture and conduct experiments on three standard datasets. The experimental results demonstrate that our method outperforms the state-of-the-art independent segmentation algorithms. On the RefCOCO validation, test A and test B datasets, our proposed method achieved IoU scores of 77.51, 79.12, and 72.79, respectively. Extensive experiments demonstrate that our method can significantly improve segmentation accuracy when fusing different sub-methods.

KW - Cross-modal

KW - Image understanding

KW - Referring image segmentation

UR - http://www.scopus.com/inward/record.url?scp=85207080766&partnerID=8YFLogxK

U2 - 10.1016/j.neucom.2024.128632

DO - 10.1016/j.neucom.2024.128632

M3 - Article

AN - SCOPUS:85207080766

SN - 0925-2312

VL - 613

JO - Neurocomputing

JF - Neurocomputing

M1 - 128632

ER -

Hierarchical collaboration for referring image segmentation

摘要

访问文件

其它文件与链接

指纹

引用此