TY - GEN
T1 - SDPS-M2CAN
T2 - 2024 China Automation Congress, CAC 2024
AU - Wang, Rensheng
AU - Zhen, Chen
AU - Dong, Ning
N1 - Publisher Copyright:
© 2024 IEEE.
PY - 2024
Y1 - 2024
N2 - Photometric stereo is a method used to estimate surface normals of objects, commonly employed in detecting surface defects on industrial products. To enhance the detection of surface features on smoother and more continuous industrial components, this study improves the learning-based photometric stereo approach, SDPS-Net, resulting in the model SDPS-M2CAN. Firstly, we introduce the Multi-scale Channel Attention Module (MCAB) to optimize the extraction of photometric stereo features, allowing the model to effectively adaptively extract useful features across multiple scales. Secondly, in the upsampling network, we employ a multi-scale feature fusion strategy based on maxpooling to fully utilize photometric stereo features at different scales, thereby enhancing the prediction capability of surface normals. Finally, by incorporating the ResNet architecture to simplify the model training process and expanding the network channels, we further improve prediction accuracy. Quantitative and qualitative experiments demonstrate that our enhanced model outperforms the original SDPS-Net in predicting concave-convex areas and normals near edges of smooth and continuous broad-spectrum reflective materials, achieving more accurate overall normal predictions.
AB - Photometric stereo is a method used to estimate surface normals of objects, commonly employed in detecting surface defects on industrial products. To enhance the detection of surface features on smoother and more continuous industrial components, this study improves the learning-based photometric stereo approach, SDPS-Net, resulting in the model SDPS-M2CAN. Firstly, we introduce the Multi-scale Channel Attention Module (MCAB) to optimize the extraction of photometric stereo features, allowing the model to effectively adaptively extract useful features across multiple scales. Secondly, in the upsampling network, we employ a multi-scale feature fusion strategy based on maxpooling to fully utilize photometric stereo features at different scales, thereby enhancing the prediction capability of surface normals. Finally, by incorporating the ResNet architecture to simplify the model training process and expanding the network channels, we further improve prediction accuracy. Quantitative and qualitative experiments demonstrate that our enhanced model outperforms the original SDPS-Net in predicting concave-convex areas and normals near edges of smooth and continuous broad-spectrum reflective materials, achieving more accurate overall normal predictions.
KW - Channel attention
KW - Multi-scale convolution
KW - Multi-scale features fusion
KW - Photometric Stereo
KW - Surface Normal Estimation
UR - https://www.scopus.com/pages/publications/86000792838
U2 - 10.1109/CAC63892.2024.10864779
DO - 10.1109/CAC63892.2024.10864779
M3 - Conference contribution
AN - SCOPUS:86000792838
T3 - Proceedings - 2024 China Automation Congress, CAC 2024
SP - 3347
EP - 3352
BT - Proceedings - 2024 China Automation Congress, CAC 2024
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 1 November 2024 through 3 November 2024
ER -