TY - JOUR
T1 - LLM Collective Intelligence for Open-Ended Medical Diagnosis
AU - Yan, Zhijun
AU - Zhang, Zhu
AU - Zhu, Bo
AU - Tang, Mingrui
AU - Wang, Tianmei
N1 - Publisher Copyright:
© 2026, Association for Information Systems. All rights reserved.
PY - 2026
Y1 - 2026
N2 - Large Language Models (LLMs) show promise for open-ended medical diagnosis, yet single-model predictions often suffer from instability, inconsistency, and limited robustness in complex clinical cases. To address this challenge, this study proposes Probabilistic Multi-LLM Fusion (PMLF), an unsupervised collective intelligence framework that aggregates diagnostic outputs from multiple LLMs at the probability-distribution level. The framework first standardizes model-generated diagnoses into a unified medical terminology space and calibrates confidence scores into comparable probability distributions. It then infers a latent consensus diagnostic distribution while jointly modeling model competence and case difficulty. Using 2,497 real-world clinical cases, experimental results show that PMLF consistently outperforms mean probability averaging, majority voting, and frequency-based aggregation across Top-K accuracy, mean reciprocal rank, and coverage. The findings demonstrate the potential of probabilistic collective intelligence to improve the reliability and robustness of LLM-assisted open-ended medical diagnosis.
AB - Large Language Models (LLMs) show promise for open-ended medical diagnosis, yet single-model predictions often suffer from instability, inconsistency, and limited robustness in complex clinical cases. To address this challenge, this study proposes Probabilistic Multi-LLM Fusion (PMLF), an unsupervised collective intelligence framework that aggregates diagnostic outputs from multiple LLMs at the probability-distribution level. The framework first standardizes model-generated diagnoses into a unified medical terminology space and calibrates confidence scores into comparable probability distributions. It then infers a latent consensus diagnostic distribution while jointly modeling model competence and case difficulty. Using 2,497 real-world clinical cases, experimental results show that PMLF consistently outperforms mean probability averaging, majority voting, and frequency-based aggregation across Top-K accuracy, mean reciprocal rank, and coverage. The findings demonstrate the potential of probabilistic collective intelligence to improve the reliability and robustness of LLM-assisted open-ended medical diagnosis.
KW - Collective Intelligence
KW - Large Language Models
KW - Medical Diagnosis
KW - Open-ended Diagnosis
KW - Probabilistic Fusion
KW - Unsupervised Learning
UR - https://www.scopus.com/pages/publications/105046068694
M3 - Conference article
AN - SCOPUS:105046068694
SN - 2689-6354
VL - PartF1
JO - Pacific Asia Conference on Information Systems
JF - Pacific Asia Conference on Information Systems
T2 - 30th Pacific Asia Conference on Information Systems, PACIS 2026
Y2 - 4 July 2026 through 8 July 2026
ER -