收录解读
血液恶性肿瘤的诊疗决策高度依赖多学科肿瘤委员会,涉及纵向治疗史、分子分型和快速演变的临床证据,但专业决策资源的可及性日益不均。
本文开发了HemaGuide,一个可本地部署的模块化大语言模型代理,能够将非结构化临床文档转化为结构化病例表示,自主将病例路由至特定决策模式(指南、高级、分子),并基于疾病特异性指南流程和超过2000例真实肿瘤委员会病例的临床决策记忆生成推荐。
在45例高复杂度病例的专家盲评中,HemaGuide显著提升了与肿瘤委员会决策的一致性。系统性消融实验证实性能提升依赖路由类型,无单一组件足以应对所有病例。外部验证555例独立病例达81.8%一致性,前瞻性静默试验64例连续病例达82.8%一致性,幻觉率仅0.3%,中位延迟39秒,可在消费级硬件上实时运行。
该工作展示了可本地部署、可审计的临床决策支持代理在血液恶性肿瘤中的可行性与稳健性,具有重要的临床转化价值。但其方法原创性基于现有LLM和模块化架构,未达到定义新研究方向的范式级贡献。
原始摘要
Abstract Multidisciplinary tumor boards integrate longitudinal treatment histories, molecular profiling and rapidly evolving evidence to guide decisions in hematological malignancies, yet access to this level of subspecialty deliberation is increasingly uneven. Here we develop HemaGuide, a locally deployable, modular large language model agent that converts unstructured clinical documents into structured case representations, autonomously routes cases to specialized decision modes (‘guideline’, ‘advanced’ and ‘molecular’) and grounds recommendations in disease-specific guideline flowcharts and a clinical decision memory of >2,000 real-world tumor board cases. In expert-blinded benchmarking on 45 high-complexity cases across six foundation models, HemaGuide substantially improved concordance with tumor board decisions. A systematic ablation study across 11 layers confirmed that performance gains were routing-type-dependent, with no single component sufficient across case types. Automated classification of 70 clinically relevant missense variants showed high concordance with expert standards; no oncogenic variant was downgraded to benign and the whole workflow was completed under real-time conditions on commodity hardware with a median latency of 39 s rather than the hours typically required for manual molecular board workflows. In a simulated practice study, agent-assisted resident physicians achieved near-senior concordance and partially outperformed senior physicians in their subspecialty. External validation on 555 independent cases from a second academic center yielded 81.8% concordance across 47 entities, and a prospective 1-month silent trial on 64 consecutive, unselected cases achieved 82.8% concordance. Hallucinations occurred in 2 of 664 evaluated cases (0.3%). Together these data provide evidence that locally deployable, case-grounded large language model agents can deliver auditable clinical decision support across hematological malignancies, with concordance maintained across institutions and under real-time conditions on commodity hardware.