Skip to main navigation Skip to search Skip to main content

视觉语言模型在焊接缺陷检测中的应用现状与展望

Translated title of the contribution: Application status and prospects of vision-language models in welding defect detection
  • Haoyu Wen
  • , Xiaopeng Wang
  • , Xinghua Yu*
  • *Corresponding author for this work
  • Beijing Institute of Technology
  • Hebei University of Science and Technology

Research output: Contribution to journalArticlepeer-review

Abstract

With a focus on the issue of whether vision-language models (VLMs) can provide substantial gains beyond traditional deep vision for welding defect detection, relevant literature was systematically analyzed within the scope of three types of detection objects: surface defects, internal defects, and weld formation anomalies. The results indicate that this research direction is still in the early exploratory stage, is complementary to existing defect detection algorithms, and has great development potential. The advantages of existing VLMs are manifested in high-level tasks such as semantic interpretation, few-shot recognition guidance, and structured report generation; the disadvantages are manifested as follows: insufficient feature extraction ability for non-natural image data such as X-rays and ultrasound; detection accuracy in fine-grained pixel localization lagging behind object detection algorithms based on CNNs; large inference latency in real-time edge deployment scenarios. Therefore, it is considered that a more reasonable engineering path at present is a hybrid complementary architecture: Traditional vision methods are responsible for precise localization and real-time front-end detection, while VLM is responsible for upper-layer semantic interpretation and structured report output, forming a hierarchical collaborative relationship rather than a replacement relationship between the two. Highlights: (1) A three-tier evidence grading framework oriented to welding defect detection was established. (2) Based on the graded evidence, the relative advantages of VLM at the semantic interpretation level and its capability bottlenecks in fine-grained localization and real-time deployment were systematically evaluated. (3) A hybrid complementary architecture direction synergizing traditional vision front-end localization and VLM back-end semantic interpretation was proposed, which provided a phased reference for engineering implementation in this field.

Translated title of the contributionApplication status and prospects of vision-language models in welding defect detection
Original languageChinese (Traditional)
Pages (from-to)75-85
Number of pages11
JournalHanjie Xuebao/Transactions of the China Welding Institution
Volume47
Issue number6
DOIs
Publication statusPublished - 2026
Externally publishedYes

Fingerprint

Dive into the research topics of 'Application status and prospects of vision-language models in welding defect detection'. Together they form a unique fingerprint.

Cite this