Skip to main navigation Skip to search Skip to main content

Weld-LLaVA: a visual-prompt-guided vision-language assistant for welding X-ray defect decision support

  • Haoyu Wen
  • , Baoxin Zhang
  • , Xuefeng Zhao
  • , Juntao Wu
  • , Na Dong
  • , Xinghua Yu*
  • *Corresponding author for this work
  • Beijing Institute of Technology
  • China Nuclear Power Engineering Co. Ltd.
  • State Grid Corporation of China
  • Ltd.
  • Beijing Institute of Technology (Zhuhai)

Research output: Contribution to journalArticlepeer-review

Abstract

Welding X-ray inspection is essential for ensuring joint integrity and process reliability in manufacturing, yet conventional vision-only detectors may struggle with ambiguous indications, overlapping defect candidates, and limited interpretability. This paper presents Weld-LLaVA, a visual-prompt-guided vision-language framework for welding X-ray defect recognition and decision support in visual question answering tasks. The proposed workflow integrates radiographic image enhancement, YOLOv8-assisted automatic candidate localization, colored bounding-box visual prompting, chain-of-thought (CoT)-style dialogue construction, and domain-specific fine-tuning of LLaVA-1.5-7B. On the human-annotated validation set containing 1798 images and 3465 defect instances, Weld-LLaVA achieves 87.13% VQA classification accuracy, outperforming GPT-4o, Claude-3.7-sonnet, Qwen2.5-VL-72B-Instruct, Mistral-Small, and the baseline LLaVA-7B model. Ablation experiments show that the proposed visual prompting strategy improves the VQA accuracy from 81.96% under the coordinate-prompt setting to 87.13%, and that the multi-turn CoT dialogue design improves defect recognition compared with single-turn inference. For 41 ambiguous YOLOv8 cases with highly overlapping boxes and conflicting labels, the proposed reasoning-assisted strategy improves the accuracy from 36.59 to 53.66% and macro-precision from 34.31 to 47.28%. Heatmap visualizations further indicate that Weld-LLaVA can use spatial semantics and defect morphology to distinguish visually similar indications in low-contrast radiographs. These findings demonstrate that explicit visual prompts and vision-language reasoning can provide interpretable decision support for ambiguous welding defect cases, thereby improving inspection consistency and supporting near-line industrial quality assurance.

Original languageEnglish
JournalWelding in the World
DOIs
Publication statusAccepted/In press - 2026
Externally publishedYes

Keywords

  • Chain-of-thought reasoning
  • Non-destructive testing
  • Vision-language model
  • Visual prompting
  • Welding radiographic testing
  • X-ray defect recognition

Fingerprint

Dive into the research topics of 'Weld-LLaVA: a visual-prompt-guided vision-language assistant for welding X-ray defect decision support'. Together they form a unique fingerprint.

Cite this