KneeDaily · 膝关节运动医学文献获取 App

ChatGPT provides accurate and safe responses to patient questions on hip arthroscopy, while completeness remains variable: A systematic review and single-arm meta-analysis.

Knee surgery, sports traumatology, arthroscopy : official journal of the ESSKA · 2026-Apr-14
阅读数 0
Ramadanov Nikolai, Penchev Plamen, Voss Maximilian, Heinz Maximilian, Prill Robert, Salzmann Mikhail et al.

简介

PURPOSE: Large language models, such as ChatGPT, are increasingly used by patients seeking information on hip arthroscopy (HAS) and femoroacetabular impingement (FAI). Despite thei…

英文摘要

PURPOSE: Large language models, such as ChatGPT, are increasingly used by patients seeking information on hip arthroscopy (HAS) and femoroacetabular impingement (FAI). Despite their linguistic fluency, the accuracy, completeness and safety of procedure-specific patient information remain unclear. Although orthopaedic studies report variable performance across subspecialties, no systematic evaluation has specifically addressed HAS. METHODS: PubMed, Embase, Scopus, CINAHL and Epistemonikos were searched to 10 February 2026 for studies evaluating ChatGPT responses to patient-oriented HAS or FAI questions. Randomized and non-randomized studies, observational cohorts and case series were eligible. Data on question sources, model versions, rating systems and performance domains (accuracy, relevance, completeness, safety, readability and clarity) were extracted. Heterogeneous rating scales were dichotomized into high- versus low-quality responses. Risk of bias was assessed using QUADAS-2 and ROBINS-I. Random-effects single-arm meta-analyses (REML) were conducted for each domain. RESULTS: Eight studies met eligibility criteria. Accuracy was high (pooled 88.6%). Relevance, safety, readability and clarity reached pooled values of 100% with low heterogeneity. Completeness was lower (83.8%) with moderate heterogeneity, mainly driven by early GPT-3.5 studies. Funnel plots showed no clear small-study effects, although interpretation was limited by the small number of studies. Risk of bias was predominantly high or moderate, largely due to non-systematic question selection and heterogeneous rating tools. Later models (GPT-4/4o and beyond) demonstrated higher performance compared with GPT-3.5. CONCLUSION: ChatGPT provides accurate, relevant, safe and clear responses to patient questions about HAS, while completeness shows moderate variability. Although LLMs appear promising as adjuncts to patient education, methodological limitations in the current evidence base underscore the need for expert clinical counselling and more rigorous, standardized evaluation frameworks. LEVEL OF EVIDENCE: Level III, systematic review and meta-analysis of non-randomized studies.

关键词

ChatGPT accuracy completeness femoroacetabular impingement hip arthroscopy meta‐analysis

链接

获取 KneeDaily

选择手机系统

iPhone前往 App Store ↗ 安卓查看下载与安装说明 →
微信内无法打开?

在微信右上角菜单选择“在浏览器打开”,或复制下面的网址到浏览器。

https://kneedaily.com/kneedaily