Preprocessing Enhanced Image Compression for Machine Vision

Guo Lu; Xingtong Ge; Tianxiong Zhong; Qiang Hu; Jing Geng

doi:10.1109/TCSVT.2024.3441049

Preprocessing Enhanced Image Compression for Machine Vision

Guo Lu, Xingtong Ge, Tianxiong Zhong, Qiang Hu^*, Jing Geng^*

^*此作品的通讯作者

计算机学院

科研成果: 期刊稿件 › 文章 › 同行评审

1 引用（Scopus）

摘要

Recently, more and more images are compressed and sent to the back-end devices for machine analysis tasks (e.g., object detection) instead of being purely watched by humans. However, most traditional or learned image codecs are designed to minimize the distortion of the human visual system without considering the increased demand from machine vision systems. In this work, we propose a preprocessing enhanced image compression method for machine vision tasks to address this challenge. Instead of relying on the learned image codecs for end-to-end optimization, our framework is built upon the traditional non-differential codecs, which means it is standard compatible and can be easily deployed in practical applications. Specifically, we propose a neural preprocessing module before the encoder to maintain the useful semantic information for the downstream tasks and suppress the irrelevant information for bitrate saving. Furthermore, our neural preprocessing module is quantization adaptive and can be used in different compression ratios. More importantly, to jointly optimize the preprocessing module with the downstream machine vision tasks, we introduce the proxy network for the traditional non-differential codecs in the back-propagation stage. We provide extensive experiments by evaluating our compression method for several representative downstream tasks with different backbone networks. Experimental results show our method achieves a better trade-off between the coding bitrate and the performance of the downstream machine vision tasks by saving about 20% bitrate.

源语言	英语
页（从-至）	13556-13568
页数	13
期刊	IEEE Transactions on Circuits and Systems for Video Technology
卷	34
期	12
DOI	https://doi.org/10.1109/TCSVT.2024.3441049
出版状态	已出版 - 2024

访问文件

10.1109/TCSVT.2024.3441049

其它文件与链接

链接到 Scopus 的出版物

引用此

Lu, G., Ge, X., Zhong, T., Hu, Q., & Geng, J. (2024). Preprocessing Enhanced Image Compression for Machine Vision. IEEE Transactions on Circuits and Systems for Video Technology, 34(12), 13556-13568. https://doi.org/10.1109/TCSVT.2024.3441049

@article{06b2c0bf3a91491aa41a528bba116a14,

title = "Preprocessing Enhanced Image Compression for Machine Vision",

abstract = "Recently, more and more images are compressed and sent to the back-end devices for machine analysis tasks (e.g., object detection) instead of being purely watched by humans. However, most traditional or learned image codecs are designed to minimize the distortion of the human visual system without considering the increased demand from machine vision systems. In this work, we propose a preprocessing enhanced image compression method for machine vision tasks to address this challenge. Instead of relying on the learned image codecs for end-to-end optimization, our framework is built upon the traditional non-differential codecs, which means it is standard compatible and can be easily deployed in practical applications. Specifically, we propose a neural preprocessing module before the encoder to maintain the useful semantic information for the downstream tasks and suppress the irrelevant information for bitrate saving. Furthermore, our neural preprocessing module is quantization adaptive and can be used in different compression ratios. More importantly, to jointly optimize the preprocessing module with the downstream machine vision tasks, we introduce the proxy network for the traditional non-differential codecs in the back-propagation stage. We provide extensive experiments by evaluating our compression method for several representative downstream tasks with different backbone networks. Experimental results show our method achieves a better trade-off between the coding bitrate and the performance of the downstream machine vision tasks by saving about 20% bitrate.",

keywords = "Image compression, deep learning, machine vision, preprocessing",

author = "Guo Lu and Xingtong Ge and Tianxiong Zhong and Qiang Hu and Jing Geng",

note = "Publisher Copyright: {\textcopyright} 1991-2012 IEEE.",

year = "2024",

doi = "10.1109/TCSVT.2024.3441049",

language = "English",

volume = "34",

pages = "13556--13568",

journal = "IEEE Transactions on Circuits and Systems for Video Technology",

issn = "1051-8215",

publisher = "Institute of Electrical and Electronics Engineers Inc.",

number = "12",

}

TY - JOUR

T1 - Preprocessing Enhanced Image Compression for Machine Vision

AU - Lu, Guo

AU - Ge, Xingtong

AU - Zhong, Tianxiong

AU - Hu, Qiang

AU - Geng, Jing

PY - 2024

Y1 - 2024

N2 - Recently, more and more images are compressed and sent to the back-end devices for machine analysis tasks (e.g., object detection) instead of being purely watched by humans. However, most traditional or learned image codecs are designed to minimize the distortion of the human visual system without considering the increased demand from machine vision systems. In this work, we propose a preprocessing enhanced image compression method for machine vision tasks to address this challenge. Instead of relying on the learned image codecs for end-to-end optimization, our framework is built upon the traditional non-differential codecs, which means it is standard compatible and can be easily deployed in practical applications. Specifically, we propose a neural preprocessing module before the encoder to maintain the useful semantic information for the downstream tasks and suppress the irrelevant information for bitrate saving. Furthermore, our neural preprocessing module is quantization adaptive and can be used in different compression ratios. More importantly, to jointly optimize the preprocessing module with the downstream machine vision tasks, we introduce the proxy network for the traditional non-differential codecs in the back-propagation stage. We provide extensive experiments by evaluating our compression method for several representative downstream tasks with different backbone networks. Experimental results show our method achieves a better trade-off between the coding bitrate and the performance of the downstream machine vision tasks by saving about 20% bitrate.

AB - Recently, more and more images are compressed and sent to the back-end devices for machine analysis tasks (e.g., object detection) instead of being purely watched by humans. However, most traditional or learned image codecs are designed to minimize the distortion of the human visual system without considering the increased demand from machine vision systems. In this work, we propose a preprocessing enhanced image compression method for machine vision tasks to address this challenge. Instead of relying on the learned image codecs for end-to-end optimization, our framework is built upon the traditional non-differential codecs, which means it is standard compatible and can be easily deployed in practical applications. Specifically, we propose a neural preprocessing module before the encoder to maintain the useful semantic information for the downstream tasks and suppress the irrelevant information for bitrate saving. Furthermore, our neural preprocessing module is quantization adaptive and can be used in different compression ratios. More importantly, to jointly optimize the preprocessing module with the downstream machine vision tasks, we introduce the proxy network for the traditional non-differential codecs in the back-propagation stage. We provide extensive experiments by evaluating our compression method for several representative downstream tasks with different backbone networks. Experimental results show our method achieves a better trade-off between the coding bitrate and the performance of the downstream machine vision tasks by saving about 20% bitrate.

KW - Image compression

KW - deep learning

KW - machine vision

KW - preprocessing

UR - http://www.scopus.com/inward/record.url?scp=85200825172&partnerID=8YFLogxK

U2 - 10.1109/TCSVT.2024.3441049

DO - 10.1109/TCSVT.2024.3441049

M3 - Article

AN - SCOPUS:85200825172

SN - 1051-8215

VL - 34

SP - 13556

EP - 13568

JO - IEEE Transactions on Circuits and Systems for Video Technology

JF - IEEE Transactions on Circuits and Systems for Video Technology

IS - 12

ER -

Preprocessing Enhanced Image Compression for Machine Vision

摘要

访问文件

其它文件与链接

指纹

引用此