基于深度强化学习与可穿戴传感的闭环数字疗法实现精准骨科康复:一项基于模拟的概念验证研究
简介
本研究提出一种基于深度强化学习(DRL)与可穿戴传感的闭环康复管理系统,并在前交叉韧带重建(ACLR)术后模拟环境中进行概念验证。该系统通过惯性测量单元(IMU)和表面肌电图(sEMG)采集患者多维度数据,利用近端策略优化(PPO)算法动态调整康复方案。模拟结果显示,与传统静态方案相比,DRL驱动的策略平均缩短15%的“安全恢复轻活动”时间,并相对减少40%…
英文摘要
OBJECTIVE: Static, one-size-fits-all protocols in postoperative orthopedic rehabilitation fail to adapt to individual recovery dynamics, potentially leading to suboptimal rehabilitation efficiency or an elevated risk of secondary injury. To address this critical gap, we propose and provide a simulation-based proof-of-concept validation for a novel closed-loop management system that deeply integrates patient-generated health data (PGHD) with deep reinforcement learning (DRL), offering a potential technical pathway for real-time personalized optimization of rehabilitation regimens and continuous prediction of long-term functional outcomes. METHODS: A three-tier system architecture was constructed, comprising an intelligent sensing layer, an AI decision-making and prediction layer, and an interactive feedback layer. Through wearable inertial measurement units (IMUs) and surface electromyography (sEMG) devices, the system continuously collected multi-dimensional PGHD, including movement quality, training intensity, adherence, and pain feedback. These heterogeneous data were encoded into a comprehensive "patient state space" through a standardized feature engineering pipeline. A proximal policy optimization (PPO) algorithm was employed to train a DRL agent to learn the optimal policy for dynamically adjusting the next-cycle rehabilitation prescription (including exercise type, intensity, frequency, and progression pace). The agent aimed to maximize a hierarchical cumulative reward function that integrates short-term safety (ΔVAS pain score monitoring), mid-term adherence (training completion rate), and long-term functional improvement. Critically, a temporal convolutional network (TCN) prognostic module was deeply coupled with the DRL agent, providing prospective predictions of future functional recovery curves that inform the agent's long-term reward calculation, equipping the system with the capability of "making decisions based on predictions." RESULTS: A proof-of-concept validation was conducted in a simulated environment for a post-operative anterior cruciate ligament reconstruction (ACLR) scenario. The trained DRL agent demonstrated the ability to generate differentiated rehabilitation strategies: in stratified analysis, it prescribed distinct progression paces for virtual patients with fast vs. slow recovery trajectories. Under this simulation setting, compared with a static conservative protocol, the DRL-driven strategy reduced the simulated time to "safe return to light activity" by an average of 15% and, compared with a static aggressive protocol, relatively decreased simulated "re-injury" events by 40%. The mean absolute errors (MAEs) of the TCN prognostic module for predicting functional scores at 2, 4, and 8 weeks into the future were 3.2, 4.8, and 6.5 points (on a 100-point scale), respectively, outperforming the ARIMA, LSTM, and GRU baseline models. An ablation study confirmed the TCN module's independent contribution, as its removal led to a relative increase in the simulated re-injury rate. CONCLUSION: This proof-of-concept study provides foundational evidence for the technical feasibility of a DRL-based closed-loop rehabilitation system. The proposed framework uniquely couples a wearable sensing layer with a symbiotic DRL-TCN architecture, demonstrating the potential to safely and dynamically personalize rehabilitation strategies in a simulated environment. These findings lay the groundwork for future prospective clinical trials, which are the necessary next step to validate safety, efficacy, and clinical utility in real-world settings.