TY - JOUR
T1 - Weld-LLaVA
T2 - a visual-prompt-guided vision-language assistant for welding X-ray defect decision support
AU - Wen, Haoyu
AU - Zhang, Baoxin
AU - Zhao, Xuefeng
AU - Wu, Juntao
AU - Dong, Na
AU - Yu, Xinghua
N1 - Publisher Copyright:
© International Institute of Welding 2026.
PY - 2026
Y1 - 2026
N2 - Welding X-ray inspection is essential for ensuring joint integrity and process reliability in manufacturing, yet conventional vision-only detectors may struggle with ambiguous indications, overlapping defect candidates, and limited interpretability. This paper presents Weld-LLaVA, a visual-prompt-guided vision-language framework for welding X-ray defect recognition and decision support in visual question answering tasks. The proposed workflow integrates radiographic image enhancement, YOLOv8-assisted automatic candidate localization, colored bounding-box visual prompting, chain-of-thought (CoT)-style dialogue construction, and domain-specific fine-tuning of LLaVA-1.5-7B. On the human-annotated validation set containing 1798 images and 3465 defect instances, Weld-LLaVA achieves 87.13% VQA classification accuracy, outperforming GPT-4o, Claude-3.7-sonnet, Qwen2.5-VL-72B-Instruct, Mistral-Small, and the baseline LLaVA-7B model. Ablation experiments show that the proposed visual prompting strategy improves the VQA accuracy from 81.96% under the coordinate-prompt setting to 87.13%, and that the multi-turn CoT dialogue design improves defect recognition compared with single-turn inference. For 41 ambiguous YOLOv8 cases with highly overlapping boxes and conflicting labels, the proposed reasoning-assisted strategy improves the accuracy from 36.59 to 53.66% and macro-precision from 34.31 to 47.28%. Heatmap visualizations further indicate that Weld-LLaVA can use spatial semantics and defect morphology to distinguish visually similar indications in low-contrast radiographs. These findings demonstrate that explicit visual prompts and vision-language reasoning can provide interpretable decision support for ambiguous welding defect cases, thereby improving inspection consistency and supporting near-line industrial quality assurance.
AB - Welding X-ray inspection is essential for ensuring joint integrity and process reliability in manufacturing, yet conventional vision-only detectors may struggle with ambiguous indications, overlapping defect candidates, and limited interpretability. This paper presents Weld-LLaVA, a visual-prompt-guided vision-language framework for welding X-ray defect recognition and decision support in visual question answering tasks. The proposed workflow integrates radiographic image enhancement, YOLOv8-assisted automatic candidate localization, colored bounding-box visual prompting, chain-of-thought (CoT)-style dialogue construction, and domain-specific fine-tuning of LLaVA-1.5-7B. On the human-annotated validation set containing 1798 images and 3465 defect instances, Weld-LLaVA achieves 87.13% VQA classification accuracy, outperforming GPT-4o, Claude-3.7-sonnet, Qwen2.5-VL-72B-Instruct, Mistral-Small, and the baseline LLaVA-7B model. Ablation experiments show that the proposed visual prompting strategy improves the VQA accuracy from 81.96% under the coordinate-prompt setting to 87.13%, and that the multi-turn CoT dialogue design improves defect recognition compared with single-turn inference. For 41 ambiguous YOLOv8 cases with highly overlapping boxes and conflicting labels, the proposed reasoning-assisted strategy improves the accuracy from 36.59 to 53.66% and macro-precision from 34.31 to 47.28%. Heatmap visualizations further indicate that Weld-LLaVA can use spatial semantics and defect morphology to distinguish visually similar indications in low-contrast radiographs. These findings demonstrate that explicit visual prompts and vision-language reasoning can provide interpretable decision support for ambiguous welding defect cases, thereby improving inspection consistency and supporting near-line industrial quality assurance.
KW - Chain-of-thought reasoning
KW - Non-destructive testing
KW - Vision-language model
KW - Visual prompting
KW - Welding radiographic testing
KW - X-ray defect recognition
UR - https://www.scopus.com/pages/publications/105048147222
U2 - 10.1007/s40194-026-02557-1
DO - 10.1007/s40194-026-02557-1
M3 - Article
AN - SCOPUS:105048147222
SN - 0043-2288
JO - Welding in the World
JF - Welding in the World
ER -