基于专家评分的三种大语言模型在前交叉韧带重建术后康复问答中的比较
简介
该研究比较了三种公开可用的大语言模型(GPT-5.4、Doubao、MiniMax-M2.7)在ACL重建术后标准化康复问答中的表现。研究于2026年3月某预设日,将30个涵盖五个术后康复阶段的英文问题分别提交给三个模型,由五位骨科临床医生从准确性、安全性、阶段适配性、完整性和可理解性五个维度进行盲法评分。结果显示:GPT-5.4总分最高(4.61±0.13…
英文摘要
BACKGROUND: Large language models (LLMs) are increasingly used by patients to obtain health information. Postoperative rehabilitation after anterior cruciate ligament reconstruction (ACLR) has distinct phase boundaries and safety considerations. Therefore, responses should be not only clear and understandable, but also medically accurate, safe, and stage-fit. This study compared the performance of three publicly accessible LLMs in standardized post-ACLR rehabilitation question answering. METHODS: This was a standardized, blinded, expert-rated comparative evaluation study. On a single prespecified data collection day in March 2026, 30 English-language rehabilitation questions were submitted separately to GPT-5.4, Doubao, and MiniMax-M2.7. The questions covered five postoperative rehabilitation phases. Responses were anonymized and randomly reordered before blinded rating by five orthopaedic clinicians across five domains: Accuracy, Safety, Stage-fit, Completeness, and Understandability. Paired non-parametric tests, effect size analyses, intraclass correlation coefficients, and linear mixed-effects modelling were used for statistical analysis. RESULTS: A total of 90 model-generated responses and 450 expert rating records were included. Overall scores differed significantly among the three models (Friedman χ² = 46.067, P < 0.001; Kendall's W = 0.768). GPT-5.4 achieved the highest overall score (4.61 ± 0.13), followed by MiniMax-M2.7 (4.53 ± 0.19), whereas Doubao had the lowest score (3.86 ± 0.29). GPT-5.4 performed best in Accuracy, Safety, and Stage-fit; MiniMax-M2.7 achieved the highest score for Completeness; and Doubao achieved the highest mean score for Understandability. Inter-rater agreement was good [ICC(3,k) = 0.893], and sensitivity analysis supported the primary findings. CONCLUSIONS: The three models showed distinct rating profiles in standardized single-turn post-ACLR rehabilitation question answering. Evaluation of patient-facing rehabilitation information should not rely solely on linguistic fluency, but should prioritize medical accuracy, safety, and Stage-fit. These findings provide preliminary benchmark evidence in a phase-sensitive rehabilitation setting, but they should not be interpreted as evidence supporting clinical implementation, clinician substitution, or patient benefit.