从病例到论文:医生的AI临床科研工作流
人工智能医学应用的历史演进与技术版图
这些文献从人工智能的历史演进、技术类型、深度学习发展及其在诊断、治疗、风险预测、影像、医疗管理和医学教育中的应用进行综述,为理解AI进入临床科研工作流提供宏观技术背景和临床转化基础。
- The history of artificial intelligence in medicine.(V. Kaul, Sarah Enslin, S. Gross, 2020, Gastrointestinal Endoscopy)
- Artificial intelligence in medicine: What is it doing for us today?(A. Becker, 2019, Health Policy and Technology)
- Artificial Intelligence in Medicine: Where Are We Now?(S. Kulkarni, N. Seneviratne, Mirza Shaheer Baig, Ameer Hamid Ahmed Khan, 2020, Academic Radiology)
- The Coming of Age of Artificial Intelligence in Medicine(V. Patel, E. Shortliffe, M. Stefanelli, Peter Szolovits, M. Berthold, R. Bellazzi, A. Abu-Hanna, 2008, Artificial Intelligence in Medicine)
- Artificial Intelligence in Medicine(John H. Holmes, Riccardo Bellazzi, Lucia Sacchi, Niels Peek, 2015, Lecture Notes in Computer Science)
- Artificial Intelligence in Medicine: Today and Tomorrow(G. Briganti, O. le Moine, 2020, Frontiers in Medicine)
- Artificial intelligence in medicine. Where do we stand?(W. B. Schwartz, R. Patil, Peter Szolovits, 1987, New England Journal of Medicine)
- Overview of Artificial Intelligence in Medicine(Chi Liu, Zachary Tan, Mingguang He, 2022, Artificial Intelligence in Medicine)
- Research on Artificial-Intelligence-Assisted Medicine: A Survey on Medical Artificial Intelligence(Fangfang Gou, Jun Liu, Chunwen Xiao, Jia Wu, 2024, Diagnostics)
- Introduction to artificial intelligence in medicine(Y. Mintz, Ronit Brodie, 2019, Minimally Invasive Therapy & Allied Technologies)
- Application of Artificial Intelligence in Medicine: An Overview(Peng Liu, Lin Lu, Jiayao Zhang, Tongtong Huo, Songxiang Liu, Z. Ye, 2021, Current Medical Science)
病例报告规范、学术价值与临床叙事表达
这些文献共同讨论病例报告的历史价值、学术功能、叙事结构、写作方法及CARE规范,强调病例信息的完整采集、透明呈现、知情同意和规范报告,是将临床观察转化为可发表论文的基础。
- The CARE Guidelines: Consensus-based Clinical Case Reporting Guideline Development(J. Gagnier, G. Kienle, D. Altman, D. Moher, H. Sox, David Riley, 2013, Global Advances in Health and Medicine)
- The CARE guidelines: consensus-based clinical case report guideline development.(Joel J. Gagnier, G. Kienle, Douglas G. Altman, David Moher, Harold C. Sox, David Riley, 2014, Journal of Clinical Epidemiology)
- The history of the case report: a selective review(T. Nissen, R. Wynn, 2014, JRSM Open)
- The recent history of the clinical case report: a narrative review(T. Nissen, R. Wynn, 2012, JRSM Short Reports)
- The clinical case report: a review of its merits and limitations(T. Nissen, R. Wynn, 2014, BMC Research Notes)
- How to Write a Case Report(M. Rossor, 2012, How to Write a Paper)
- Form and Representation in Clinical Case Reports(B. Hurwitz, 2007, Literature and Medicine)
临床科研工作流整合、信息基础设施与流程设计
这些研究关注如何把临床照护、临床研究、临床试验、电子病历、研究信息系统和人员协作连接起来,涉及工作流建模、用户参与、跨中心协作、数据质量和研究基础设施,是病例发现、研究组织与科研实施的系统基础。
- Connecting healthcare and clinical research: Workflow optimizations through seamless integration of EHR, pseudonymization services and EDC systems(P. Bruland, Justin Doods, T. Brix, M. Dugas, M. Storck, 2018, International Journal of Medical Informatics)
- Conceptual Framework to Support Clinical Trial Optimization and End-to-End Enrollment Workflow(Neha M. Jain, Alison Culley, T. Knoop, C. Micheel, T. Osterman, M. Levy, 2019, JCO Clinical Cancer Informatics)
- Physician participation in clinical research and trials: issues and approaches(Sayeeda Rahman, Azim Majumder, Sami Shaban, Nuzhat Rahman, SM Moslehuddin Ahmed, Khalid A. Bin Abdulrahman, D’Souza, 2011, Advances in Medical Education and Practice)
- Clinical Research Informatics: Challenges, Opportunities and Definition for an Emerging Domain(Peter J. Embí, Philip Payne, 2009, Journal of the American Medical Informatics Association)
- Leveraging a clinical research information system to assist biospecimen data and workflow management: a hybrid approach(P. Nadkarni, R. Kemp, C. Parikh, 2011, Journal of Clinical Bioinformatics)
- Workflow in Clinical Trial Sites & Its Association with Near Miss Events for Data Quality: Ethnographic, Workflow & Systems Simulation(Elias Cesar Araujo de Carvalho, A. Batilana, W. Claudino, L. F. Lima Reis, R. Schmerling, Jatin Shah, R. Pietrobon, 2012, PLoS ONE)
- Integrating clinical research into electronic health record workflows to support a learning health system(N. Goldhaber, Marni B. Jacobs, Louise C. Laurent, R. Knight, Wenhong Zhu, D. Pham, Allen Tran, S. P. Patel, Michael Hogarth, Christopher A. Longhurst, 2024, JAMIA Open)
- Design for improved workflow(M. Ozkaynak, B. Reeder, S. Park, Jina Huh-Yoo, 2020, Design for Health)
- Adapting day to day clinical workflow for clinical research(Rohan J. Kurian, Steven Falowski, 2026, Clinical Research in Private Practice)
- Standardizing Clinical Trials Workflow Representation in UML for International Site Comparison(E. C. A. De Carvalho, M. K. Jayanti, A. Batilana, Andreia M. O. Kozan, M. J. Rodrigues, Jatin Shah, M. R. Loures, Sunita Patil, Philip R. O. Payne, R. Pietrobon, 2010, PLoS ONE)
- Measuring Clinical Workflow to Improve Quality and Safety(M. Tanzini, J. Westbrook, S. Guidi, Neroli S Sunderland, M. Prgomet, 2020, Textbook of Patient Safety and Clinical Risk Management)
- An Integrated Model for Patient Care and Clinical Trials (IMPACT) to Support Clinical Research Visit Scheduling Workflow for Future Learning Health Systems(C. Weng, Yu Li, S. Berhe, M. Boland, Junfeng Gao, G. Hruby, R. Steinman, C. López-Jiménez, Linda Busacca, G. Hripcsak, Suzanne Bakken, J. Bigger, 2013, Journal of Biomedical Informatics)
- Principles for Designing and Developing a Workflow Monitoring Tool to Enable and Enhance Clinical Workflow Automation(Danny T. Y. Wu, Lindsey Barrick, M. Ozkaynak, K. Blondon, Kai Zheng, 2022, Applied Clinical Informatics)
临床试验运营监测与研究流程自动化
这些文献聚焦临床试验运营和临床照护过程中的自动化数据采集、不良事件报告及流程衔接,体现AI和信息技术在研究执行、质量监测及运营效率提升中的作用。
- The automation of clinical trial serious adverse event reporting workflow(J. London, K. J. Smalley, K. Conner, J. B. Smith, 2009, Clinical Trials)
- Automated data extraction: merging clinical care with real-time cohort-specific research and quality improvement data.(F. Hebal, Elizabeth Nanney, C. Stake, Michael L. Miller, George Lales, K. Barsness, 2017, Journal of Pediatric Surgery)
临床非结构化文本的信息抽取与变量标准化
这些研究主要从电子病历、病例记录、病理或其他临床文本中识别、抽取和规范化研究变量,涵盖规则方法、自然语言处理和生成式AI,直接支撑病例筛选、队列构建和科研数据集准备。
- Automated data extraction of electronic medical records: Validity of data mining to construct research databases for eligibility in gastroenterological clinical trials(Nora Joseph, Ida Lindblad, S. Zaker, Sharareh Elfversson, Maria Albinzon, Øyvind Ødegård, Li Hantler, P. M. Hellström, 2022, Upsala Journal of Medical Sciences)
- Clinical Data Extraction and Normalization of Cyrillic Electronic Health Records Via Deep-Learning Natural Language Processing(Boyang Zhao, 2019, JCO Clinical Cancer Informatics)
- Information Extraction for Clinical Data Mining: A Mammography Case Study(Houssam Nassif, R. Woods, E. Burnside, Mehmet Ayvaci, J. Shavlik, David Page, 2009, 2009 IEEE International Conference on Data Mining Workshops)
- Ad Hoc Information Extraction for Clinical Data Warehouses(Georg Dietrich, Jonathan Krebs, G. Fette, Maximilian Ertl, Mathias Kaspar, S. Störk, F. Puppe, 2018, Methods of Information in Medicine)
- Data extraction for epidemiological research (DExtER): a novel tool for automated clinical epidemiology studies(K. Gokhale, J. Chandan, K. Toulis, G. Gkoutos, P. Tiňo, K. Nirantharakumar, 2020, European Journal of Epidemiology)
- Using Generative AI to Extract Structured Information from Free Text Pathology Reports(Fahad Shahid, Min-Huei Hsu, Yung-Chun Chang, W. Jian, 2025, Journal of Medical Systems)
- Extraction of clinical data on major pulmonary diseases from unstructured radiologic reports using a large language model(H. Park, Jin-Young Huh, G. Chae, M. Choi, 2024, PLOS ONE)
- Rule-based information extraction from patients' clinical data(A. Mykowiecka, M. Marciniak, Anna Kupsc, 2009, Journal of Biomedical Informatics)
- Text data extraction for a prospective, research-focused data mart: implementation and validation(M. Hinchcliff, E. M. Just, Sofia Podlusky, J. Varga, R. Chang, W. Kibbe, 2012, BMC Medical Informatics and Decision Making)
- Development and validation of a novel AI framework using NLP with LLM integration for relevant clinical data extraction through automated chart review(M. M. Dagli, Yohannes G Ghenbot, H. Ahmad, Daksh Chauhan, Ryan W Turlip, Patrick T Wang, William C. Welch, Ali K. Ozturk, J. Yoon, 2024, Scientific Reports)
临床科研数据工程、ETL与可分析数据集构建
这些文献关注临床数据从业务系统进入研究数据库的ETL、预处理、数据仓库建设和信息技术辅助抽取,重点解决多源数据整合、格式转换、质量控制及可分析性问题。
- Dynamic-ETL: a hybrid approach for health data extraction, transformation and loading(Toan C Ong, M. Kahn, Bethany M. Kwan, Traci Yamashita, E. Brandt, Patrick Hosokawa, Christopher A. Uhrich, L. Schilling, 2017, BMC Medical Informatics and Decision Making)
- MIMIC-Extract: a data extraction, preprocessing, and representation pipeline for MIMIC-III(Shirly Wang, Matthew B. A. McDermott, Geeticka Chauhan, Michael C. Hughes, Tristan Naumann, M. Ghassemi, 2019, Proceedings of the ACM Conference on Health, Inference, and Learning)
- Application of Information Technology: Data Extraction and Ad Hoc Query of an Entity - Attribute - Value Database(P. Nadkarni, 1998, Journal of the American Medical Informatics Association)
医学文献检索、系统综述与循证证据生产自动化
这些文献共同研究AI在医学文献检索、PICO问题处理、文献筛选、数据提取和偏倚风险评价中的应用,并讨论召回率、精确率、敏感性、可靠性和人工复核要求,支撑医生从病例问题到循证论证的过程。
- The Impact of Systematic Review Automation Tools on Methodological Quality and Time Taken to Complete Systematic Review Tasks: Case Study(J. Clark, Catherine McFarlane, Gina Cleo, Christiane Ishikawa Ramos, Skye Marshall, 2020, JMIR Medical Education)
- Generative AI for Evidence-Based Medicine: A PICO GenAI for Synthesizing Clinical Case Reports(S. Mohammed, J. Fiaidhi, 2024, ICC 2024 - IEEE International Conference on Communications)
- Automated medical literature screening using artificial intelligence: a systematic review and meta-analysis(Yunying Feng, Siyu Liang, Yuelun Zhang, Shi Chen, Qing Wang, Tianze Huang, F. Sun, Xiaoqing Liu, Huijuan Zhu, Hui Pan, 2022, Journal of the American Medical Informatics Association)
- Comparison of a traditional systematic review approach with review-of-reviews and semi-automation as strategies to update the evidence(Shivani M. Reddy, Sheila V. Patel, Meghan S. Weyrich, Joshua Fenton, M. Viswanathan, 2020, Systematic Reviews)
- Advanced analytics for the automation of medical systematic reviews(Prem Timsina, Jun Liu, O. El-Gayar, 2015, Information Systems Frontiers)
- Automation in Healthcare Systematic Review(R. Ruiz, V. Duffy, 2021, Lecture Notes in Computer Science)
- Automation of systematic reviews of biomedical literature: a scoping review of studies indexed in PubMed(Barbara Tóth, László Berek, L. Gulácsi, M. Péntek, Z. Zrubka, 2024, Systematic Reviews)
- Systematic review automation technologies(G. Tsafnat, P. Glasziou, M. K. Choong, A. Dunn, Filippo Galgani, E. Coiera, 2014, Systematic Reviews)
- Automation of literature screening using machine learning in medical evidence synthesis: a diagnostic test accuracy systematic review protocol(Yuelun Zhang, Siyu Liang, Yunying Feng, Qing Wang, F. Sun, Shi Chen, Yiying Yang, Xin He, Huijuan Zhu, Hui Pan, 2022, Systematic Reviews)
- Assessing the Reliability of Large Language Models for Evaluation of Risk of Bias in Randomized Clinical Trials(Takeshi Nagao, Tetsuya Kawakita, 2026, American Journal of Perinatology)
科研分析自动化、工作流管理与可复现性
这些研究覆盖组学实验、标志物发现、生物信息学工作流管理和自动机器学习,强调模块化流程、自动化分析、质量控制、模型性能评估和结果可复现性,代表AI科研工作流中的分析建模层。
- Proteomics of Human Milk: Definition of a Discovery Workflow for Clinical Research Studies.(L. Dayon, Charlotte Macron, S. Lahrichi, A. Núñez Galindo, M. Affolter, 2021, Journal of Proteome Research)
- Small molecule biomarker discovery: Proposed workflow for LC-MS-based clinical research projects(S. Rischke, L. Hahnefeld, B. Burla, F. Behrens, R. Gurke, T. Garrett, 2023, Journal of Mass Spectrometry and Advances in the Clinical Lab)
- Design considerations for workflow management systems use in production genomics research and the clinic(Azza E. Ahmed, J. M. Allen, Tajesvi Bhat, P. Burra, Christina E. Fliege, S. Hart, Jacob R Heldenbrand, M. Hudson, Dave D. Istanto, Michael Kalmbach, Gregory D. Kapraun, Katherine I Kendig, Matthew C. Kendzior, E. Klee, Nathan R Mattson, Christian A. Ross, S. Sharif, R. Venkatakrishnan, Faisal M. Fadlelmola, L. Mainzer, 2021, Scientific Reports)
- Clinical performance of automated machine learning: a systematic review(A. Thirunavukarasu, K. Elangovan, Laura Gutierrez, Refaat Hassan, Yong Li, Ting Fang Tan, Haoran Cheng, Zhen Ling Teo, Gilbert Lim, D. Ting, 2023, medRxiv)
- Large language models streamline automated machine learning for clinical studies(Soroosh Tayebi Arasteh, T. Han, Mahshad Lotfinia, C. Kuhl, Jakob Nikolas Kather, D. Truhn, S. Nebelung, 2023, Nature Communications)
大语言模型在医疗科研中的证据评价与风险边界
这些文献从综述、证据评价和前景分析角度考察大语言模型在医疗和生物医学中的能力、应用成熟度与评价方法,同时总结幻觉、偏倚、隐私、责任和实施障碍,为LLM进入临床科研流程划定能力边界。
- The future landscape of large language models in medicine(J. Clusmann, F. Kolbinger, H. Muti, Zunamys I. Carrero, J. Eckardt, Narmin Ghaffari Laleh, C. Löffler, Sophie-Caroline Schwarzkopf, Michaela Unger, G. Veldhuizen, Sophia J. Wagner, Jakob Nikolas Kather, 2023, Communications Medicine)
- Large Language Models in Healthcare and Medical Domain: A Review(Zabir Al Nazi, Wei Peng, 2024, Informatics)
- Large language models in biomedicine and health: current research landscape and future directions(Zhiyong Lu, Yifan Peng, T. Cohen, Marzyeh Ghassemi, C. Weng, Shubo Tian, 2024, Journal of the American Medical Informatics Association)
- Assessing the research landscape and clinical utility of large language models: a scoping review(Ye-Jean Park, Abhinav Pillai, J. Deng, Eddie Guo, Mehul Gupta, M. Paget, Christopher Naugler, 2024, BMC Medical Informatics and Decision Making)
- Evaluating the Application of Large Language Models in Clinical Research Contexts.(R. Perlis, Stephan D. Fihn, 2023, JAMA Network Open)
大语言模型驱动的临床研究设计、试验匹配与决策支持
这些文献聚焦大语言模型在临床决策支持、知识问答、预后与治疗建议、临床试验设计、患者—试验匹配、虚拟试验和临床机器学习中的具体应用,展示LLM参与科研问题分析和研究执行的方式。
- Application of large language models in medicine(Fenglin Liu, Hongjian Zhou, Boyang Gu, Xinyu Zou, Jinfa Huang, Jinge Wu, Yiru Li, Sam S. Chen, Y. Hua, Peilin Zhou, Junling Liu, Chengfeng Mao, Chenyu You, Xian Wu, Yefeng Zheng, Lei A. Clifton, Zheng Li, Jiebo Luo, David A. Clifton, 2025, Nature Reviews Bioengineering)
- Large language models in medicine(A. Thirunavukarasu, D. S. J. Ting, K. Elangovan, Laura Gutierrez, Ting Fang Tan, D. Ting, 2023, Nature Medicine)
- Generative AI for Transformative Healthcare: A Comprehensive Study of Emerging Models, Applications, Case Studies, and Limitations(Siva Sai, Aanchal Gaur, Revant Sai, V. Chamola, Mohsen Guizani, J. Rodrigues, 2024, IEEE Access)
- Large language models in clinical trials: applications, technical advances, and future directions(Anqi Lin, Zhihan Wang, Aimin Jiang, Li Chen, Chang Qi, Lingxuan Zhu, W. Mou, Wenyi Gan, Dongqiang Zeng, Mingjia Xiao, Guangdi Chu, Shengkun Peng, Hank Z. H. Wong, Lin Zhang, Hengguo Zhang, Xinpei Deng, Yaxuan Wang, Jian Zhang, Quan Cheng, Bufu Tang, Peng Luo, 2025, BMC Medicine)
- Large language models encode clinical knowledge(K. Singhal, Shekoofeh Azizi, T. Tu, S. Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, A. Tanwani, H. Cole-Lewis, S. Pfohl, P. Payne, Martin G. Seneviratne, P. Gamble, Chris Kelly, Nathaneal Scharli, A. Chowdhery, P. A. Mansfield, B. A. Y. Arcas, D. Webster, Greg S. Corrado, Yossi Matias, K. Chou, Juraj Gottweis, Nenad Tomašev, Yun Liu, A. Rajkomar, J. Barral, Christopher Semturs, A. Karthikesalingam, Vivek Natarajan, 2022, Nature)
- Transforming clinical trials: the emerging roles of large language models(Jong-Lyul Ghim, S. Ahn, 2023, Translational and Clinical Pharmacology)
- Large language models in medicine: A review of current clinical trials across healthcare applications(Mahmud Omar, Girish N. Nadkarni, Eyal Klang, B. Glicksberg, 2024, PLOS Digital Health)
- Applications of Large Language Models in Medical Research: From Systematic Reviews to Clinical Studies(Eun Jeong Gong, Chang Seok Bang, Yen Shin, 2026, Bioengineering)
- Distilling large language models for matching patients to clinical trials(Mauro Nievas, Aditya Basu, Yanshan Wang, Hrituraj Singh, 2024, Journal of the American Medical Informatics Association)
生成式AI辅助医学写作、病例整理与临床沟通
这些研究突出生成式AI在检索增强生成、医学知识问答、研究方案和论文写作、影像报告生成、病例资料整理及患者友好沟通中的作用,体现从科研内容生产到临床成果表达的应用环节。
- From RAGs to riches: Utilizing large language models to write documents for clinical trials(N. Markey, Ilyass El-Mansouri, Gaëtan Rensonnet, Casper van Langen, Christoph Meier, 2024, Clinical Trials)
- A study of generative large language model for medical research and healthcare(Cheng Peng, Xi Yang, Aokun Chen, Kaleb E. Smith, Nima M. Pournejatian, Anthony B Costa, Cheryl Martin, Mona G. Flores, Ying Zhang, Tanja Magoc, Gloria P. Lipori, Duane A. Mitchell, N. Ospina, Mustafa M. Ahmed, W. Hogan, E. Shenkman, Yi Guo, Jiang Bian, Yonghui Wu, 2023, npj Digital Medicine)
- Strategies for integrating ChatGPT and generative AI into clinical studies(Jeong-Moo Lee, 2024, Blood Research)
- Patient-centered radiology reports with generative artificial intelligence: adding value to radiology reporting(Jiwoo Park, K. Oh, Kyunghwa Han, Young Han Lee, 2024, Scientific Reports)
- A Case Report on Ground-Level Alternobaric Vertigo Due to Eustachian Tube Dysfunction With the Assistance of Conversational Generative Pre-trained Transformer (ChatGPT)(Hee-Young Kim, 2023, Cureus)
AI临床科研的可解释性、公平性、伦理与责任治理
这些文献讨论AI临床科研应用中的可解释性、因果可理解性、公平性、偏倚、隐私、数据共享、监管、患者—医生关系、研究诚信和组织实施,构成工作流中人工监督、伦理审查与责任治理的核心依据。
- Artificial intelligence in medicine: the challenges ahead.(E. Coiera, 1996, Journal of the American Medical Informatics Association)
- Causability and explainability of artificial intelligence in medicine(Andreas Holzinger, G. Langs, H. Denk, K. Zatloukal, Heimo Müller, 2019, WIREs Data Mining and Knowledge Discovery)
- Automation and artificial intelligence in the clinical laboratory(C. Naugler, D. Church, 2019, Critical Reviews in Clinical Laboratory Sciences)
- Physician clinical decision modification and bias assessment in a randomized controlled trial of AI assistance(Ethan Goh, B. Bunning, Elaine C. Khoong, Robert J. Gallo, Arnold Milstein, Damon Centola, Jonathan H. Chen, 2025, Communications Medicine)
- A manifesto on explainability for artificial intelligence in medicine(Carlo Combi, Beatrice Amico, R. Bellazzi, Andrea Holzinger, Jason W. Moore, M. Zitnik, J. H. Holmes, 2022, Artificial Intelligence in Medicine)
- The practical implementation of artificial intelligence technologies in medicine(J. He, Sally L. Baxter, Jie Xu, Jiming Xu, Xingtao Zhou, Kang Zhang, 2019, Nature Medicine)
- Clinical research and the physician–patient relationship: the dual roles of physician and researcher(N. King, L. Churchill, 2008, The Cambridge Textbook of Bioethics)
- Professional integrity in clinical research.(F. Miller, D. Rosenstein, Evan G. DeRenzo, 1998, JAMA)
- Ethics of large language models in medicine and medical research.(H. Li, J. Moon, S. Purkayastha, L. Celi, H. Trivedi, J. Gichoya, 2023, The Lancet Digital Health)
医生科研角色、能力培养与AI协作基础
这些文献关注医生科学家的科研角色、临床医生参与研究的必要性以及科研自我效能和能力培养,为分析AI如何降低技术门槛、提升医生科研参与度和构建人机协作模式提供人员基础。
- Encouraging clinical research by physician scientists.(K. Shine, 1998, JAMA)
- Assessing Research Self-Efficacy in Physician-Scientists: The Clinical Research APPraisal Inventory(Elizabeth A. Mullikin, L. Bakken, N. Betz, 2007, Journal of Career Assessment)
合并后形成一条由“技术背景—病例规范—工作流基础设施—数据抽取与工程—循证检索—自动化分析—LLM科研应用—医学写作—治理与人员能力”组成的完整链路。分组将病例报告规范、临床数据准备、研究流程整合、证据综合、计算分析及大语言模型应用分别处理,避免把不同环节过度合并;同时以可解释性、伦理责任、人工复核和医生科研能力作为贯穿全流程的质量保障。
总计 92 篇相关文献
The clinical case report has a long-standing tradition in the medical literature. While its scientific significance has become smaller as more advanced research methods have gained ground, case reports are still presented in many medical journals. Some scholars point to its limited value for medical progress, while others assert that the genre is undervalued. We aimed to present the various points of view regarding the merits and limitations of the case report genre. We searched Google Scholar, PubMed and select textbooks on epidemiology and medical research for articles and book-chapters discussing the merits and limitations of clinical case reports and case series. The major merits of case reporting were these: Detecting novelties, generating hypotheses, pharmacovigilance, high applicability when other research designs are not possible to carry out, allowing emphasis on the narrative aspect (in-depth understanding), and educational value. The major limitations were: Lack of ability to generalize, no possibility to establish cause-effect relationship, danger of over-interpretation, publication bias, retrospective design, and distraction of reader when focusing on the unusual. Despite having lost its central role in medical literature in the 20th century, the genre still appears popular. It is a valuable part of the various research methods, especially since it complements other approaches. Furthermore, it also contributes in areas of medicine that are not specifically research-related, e.g. as an educational tool. Revision of the case report genre has been attempted in order to integrate the biomedical model with the narrative approach, but without significant success. The future prospects of the case report could possibly be in new applications of the genre, i.e. exclusive case report databases available online, and open access for clinicians and researchers.
Background: A case report is a narrative that describes, for medical, scientific, or educational purposes, a medical problem experienced by one or more patients. Case reports written without guidance from reporting standards are insufficiently rigorous to guide clinical practice or to inform clinical study design. Primary Objective: Develop, disseminate, and implement systematic reporting guidelines for case reports. Methods: We used a three-phase consensus process consisting of (1) premeeting literature review and interviews to generate items for the reporting guidelines, (2) a face-to-face consensus meeting to draft the reporting guidelines, and (3) postmeeting feedback, review, and pilot testing, followed by finalization of the case report guidelines. Results: This consensus process involved 27 participants and resulted in a 13-item checklist—a reporting guideline for case reports. The primary items of the checklist are title, key words, abstract, introduction, patient information, clinical findings, timeline, diagnostic assessment, therapeutic interventions, follow-up and outcomes, discussion, patient perspective, and informed consent. Conclusions: We believe the implementation of the CARE (CAse REport) guidelines by medical journals will improve the completeness and transparency of published case reports and that the systematic aggregation of information from case reports will inform clinical study design, provide early signals of effectiveness and harms, and improve healthcare delivery.
… Case reports may generate hypotheses for future clinical studies, prove useful in the … of treatments in clinical practice [4], [5]. Furthermore, case reports offer a structure for case-based …
The clinical case report is a literary tool that enables clinicians to depict and think medically about a sick person's situation. Composed for the most part as chronologies set out as observations, case reports are marked by various forms of literalism contrasted with more expressive forms of writing. In this way, they adopt literary framings and dramatic devices to help to convey the patient's overall situation and the physician's reaction to it. Encoding the clinically essential for the purposes of record, demonstration, and communication, the shape and emphases of case reports today build on, and contribute to, a long train of developments in medical theory, practice, and case description. My aim in this essay is not to suggest a single unbroken continuity between past and present case reports, but rather to explore the variety of textual representations deployed in their composition and to suggest that the plurality of past forms licenses renewed experimentation in the writing of clinical case reports today.
… citations about writing case reports.We also combined the term case report with keywords for … to the clinical question and the purpose of the case report, and it defies the critical need for …
The clinical case report is a popular genre in medical writing. While authors and editors have debated the justification for the clinical case report, few have attempted to examine the long history of this genre in medical literature. By reviewing selected literature and presenting and discussing excerpts of clinical case reports from Egyptian antiquity to the 20th century, we illustrate the presence of the genre in medical science and how its form developed. Central features of the clinical case report in different time periods are discussed, including its main components, structure, style and author presence.
Clinical case reporting in the form of case reports and case series reports has always been an integral part of medical literature. From the late 1970s the genre appeared to fall from grace and was marginalized in many medical journals. There was controversy as to its value as a research method. From the late 1990s and onwards, there has been an increased demand for and publication of case reports and case series. The various causes for its decline and subsequent return are discussed with an emphasis on the recent historical context.
Artificial intelligence assistance in clinical decision making shows promise, but concerns exist about potential exacerbation of demographic biases in healthcare. This study aims to evaluate how physician clinical decisions and biases are influenced by AI assistance in a chest pain triage scenario. A randomized, pre post-intervention study was conducted with 50 US-licensed physicians who reviewed standardized chest pain video vignettes featuring either a white male or Black female patient. Participants answered clinical questions about triage, risk assessment, and treatment before and after receiving GPT-4 generated recommendations. Clinical decision accuracy was evaluated against evidence-based guidelines. Here we show that physicians are willing to modify their clinical decisions based on GPT-4 assistance, leading to improved accuracy scores from 47% to 65% in the white male patient group and 63% to 80% in the Black female patient group. The accuracy improvement occurs without introducing or exacerbating demographic biases, with both groups showing similar magnitudes of improvement (18%). A post-study survey indicates that 90% of physicians expect AI tools to play a significant role in future clinical decision making. Physician clinical decision making can be augmented by AI assistance while maintaining equitable care across patient demographics. These findings suggest a path forward for AI clinical decision support that improves medical care without amplifying healthcare disparities. Doctors sometimes make different medical decisions for patients based on their race or gender, even when the symptoms are the same and the advice should be similar. New artificial intelligence (AI) tools such as GPT-4 are becoming available to assist doctors when making clinical decisions. Our study looked at whether using AI would impact performance and bias during doctor decision making. We investigated how doctors respond to AI suggestions when evaluating chest pain, a common but serious medical concern. We showed 50 doctors a video of either a white male or Black female patient describing chest pain symptoms, and asked them to make medical decisions. The doctors then received suggestions from an AI system and could change their decisions. We found that doctors were willing to consider the AI’s suggestions and made more accurate medical decisions after receiving this help. This improvement in decision-making happened equally for all patients, regardless of their race or gender, suggesting AI tools could help improve medical care without increasing bias. Goh et al. evaluate physician responses to GPT-4 assistance in standardized chest pain cases. The study demonstrates that physicians show willingness to modify their clinical decisions based on GPT-4 assistance, leading to improved clinical decision accuracy without introducing or exacerbating demographic biases in patient care.
The rapid development of new drugs, therapies, and devices has created a dramatic increase in the number of clinical research studies that highlights the need for greater participation in research by physicians as well as patients. Furthermore, the potential of clinical research is unlikely to be reached without greater participation of physicians in research. Physicians face a variety of barriers with regard to participation in clinical research. These barriers are system-or organization-related as well as research-and physician-related. To encourage physician participation, appropriate organizational and operational infrastructures are needed in health care institutes to support research planning and management. All physicians should receive education and training in the fundamentals of research design and methodology, which need to be incorporated into undergraduate medical education and postgraduate training curricula and then reinforced through continuing medical education. Medical schools need to analyze current practices of teaching-learning and research, and reflect upon possible changes needed to develop a 'student-focused teaching-learning and research culture'. This article examines the barriers to and benefits of physician participation in clinical research as well as interventions needed to increase their participation, including the specific role of undergraduate medical education. The main challenge is the unwillingness of many physicians and patients to participate in clinical trials. Barriers to participation include lack of time, lack of resources, trial-specific issues, communication difficulties, conflicts between the role of clinician and scientist, inadequate research experience and training for physicians, lack of rewards and recognition for physicians, and sometimes a scientifically uninteresting research question, among others. Strategies to encourage physician participation in clinical research include financial and nonfinancial incentives, adequate training, research questions that are in line with physician interests and have clear potential to improve patient care, and regular feedback. Finally, encouraging research culture and fostering the development of inquiry and research-based learning among medical students is now a high priority in order to develop more and better clinician-researchers.
… Clinical Research by Physician Scientists Why is it important that physicians actively conduct clinical research… independently or in collaboration with physicians. But only the well-trained …
Clinical research and the physician–patient relationship: the dual roles of physician and researcher
… that physicians think of clinical research as distinguishable from, and complementary to, medical treatment. All clinical research is essentially future oriented: after the research is …
… In this article we focus on challenging issues of professional integrity posed by a significant portion of clinical research studies conducted by physician investigators with patient …
… , on the career development of physician researchers, especially women and people of color… clinical research self-efficacy in a population of physicians training for clinical research …
Artificial intelligence (AI) was first described in 1950; however, several limitations in early models prevented widespread acceptance and application to medicine. In the early 2000s, many of these limitations were overcome by the advent of deep learning. Now AI systems are capable of analyzing complex algorithms and self-learning, we enter a new age in medicine where AI can be applied to clinical practice through risk assessment models, improving diagnostic accuracy and improving workflow efficiency. This article presents a brief historical perspective on the evolution of AI over the last several decades and the introduction and development of AI in medicine in recent years. A brief summary of the major applications of AI in gastroenterology and endoscopy are also presented, which will be reviewed in further detail by several other articles in this issue of GIE.
… The principal aims of AIME are to foster fundamental and applied research in the application of artificial intelligence (AI) techniques to medical care and medical research, and to …
The development of artificial intelligence (AI)-based technologies in medicine is advancing rapidly, but real-world clinical implementation has not yet become a reality. Here we review some of the key practical issues surrounding the implementation of AI into existing clinical workflows, including data sharing and privacy, transparency of algorithms, data standardization, and interoperability across multiple platforms, and concern for patient safety. We summarize the current regulatory environment in the United States and highlight comparisons with other regions in the world, notably Europe and China. To bring the full potential of AI to the clinic, practical and regulatory changes need to be made in health systems globally.
Abstract The term Artificial Intelligence (AI) was coined by John McCarthy in 1956 during a conference held on this subject. However, the possibility of machines being able to simulate human behavior and actually think was raised earlier by Alan Turing who developed the Turing test in order to differentiate humans from machines. Since then, computational power has grown to the point of instant calculations and the ability evaluate new data, according to previously assessed data, in real time. Today, AI is integrated into our daily lives in many forms, such as personal assistants (Siri, Alexa, Google assistant etc.), automated mass transportation, aviation and computer gaming. More recently, AI has also begun to be incorporated into medicine to improve patient care by speeding up processes and achieving greater accuracy, opening the path to providing better healthcare overall. Radiological images, pathology slides, and patients’ electronic medical records (EMR) are being evaluated by machine learning, aiding in the process of diagnosis and treatment of patients and augmenting physicians’ capabilities. Herein we describe the current status of AI in medicine, the way it is used in the different disciplines and future trends.
… Artificial intelligent techniques such as fuzzy expert systems, Bayesian networks, artificial neural networks, and hybrid intelligent … AI in medicine can be dichotomized into two subtypes: …
Artificial intelligence-powered medical technologies are rapidly evolving into applicable solutions for clinical practice. Deep learning algorithms can deal with increasing amounts of data provided by wearables, smartphones, and other mobile monitoring sensors in different areas of medicine. Currently, only very specific settings in clinical practice benefit from the application of artificial intelligence, such as the detection of atrial fibrillation, epilepsy seizures, and hypoglycemia, or the diagnosis of disease based on histopathological examination or medical imaging. The implementation of augmented medicine is long-awaited by patients because it allows for a greater autonomy and a more personalized treatment, however, it is met with resistance from physicians which were not prepared for such an evolution of clinical practice. This phenomenon also creates the need to validate these modern tools with traditional clinical trials, debate the educational upgrade of the medical curriculum in light of digital medicine as well as ethical consideration of the ongoing connected monitoring. The aim of this paper is to discuss recent scientific literature and provide a perspective on the benefits, future opportunities and risks of established artificial intelligence applications in clinical practice on physicians, healthcare institutions, medical education, and bioethics.
Artificial intelligence in medicine has made dramatic progress in recent years. However, much of this progress is seemingly scattered, lacking a cohesive structure for the discerning observer. In this article, we will provide an up-to-date review of artificial intelligence in medicine, with a specific focus on its application to radiology, pathology, ophthalmology, and dermatology. We will discuss a range of selected papers that illustrate the potential uses of artificial intelligence in a technologically advanced future.
With the improvement of economic conditions and the increase in living standards, people’s attention in regard to health is also continuously increasing. They are beginning to place their hopes on machines, expecting artificial intelligence (AI) to provide a more humanized medical environment and personalized services, thus greatly expanding the supply and bridging the gap between resource supply and demand. With the development of IoT technology, the arrival of the 5G and 6G communication era, and the enhancement of computing capabilities in particular, the development and application of AI-assisted healthcare have been further promoted. Currently, research on and the application of artificial intelligence in the field of medical assistance are continuously deepening and expanding. AI holds immense economic value and has many potential applications in regard to medical institutions, patients, and healthcare professionals. It has the ability to enhance medical efficiency, reduce healthcare costs, improve the quality of healthcare services, and provide a more intelligent and humanized service experience for healthcare professionals and patients. This study elaborates on AI development history and development timelines in the medical field, types of AI technologies in healthcare informatics, the application of AI in the medical field, and opportunities and challenges of AI in the field of medicine. The combination of healthcare and artificial intelligence has a profound impact on human life, improving human health levels and quality of life and changing human lifestyles.
Abstract With its origins in the mid- to late-1900s, today, artificial intelligence (AI) is used in a wide range of medical fields for varying purposes. This review first covers the early work regarding AI in medicine, then aims to elucidate some of the most current applications of machine learning in medicine according to the following four specific categories: (1) its use in assessing the risk of disease onset and in estimating treatment success; (2) its use in managing or alleviating complications; (3) its role in ongoing patient care; and (4) its use in ongoing pathology and treatment efficacy research. Lastly, this paper clarifies some of the potential drawbacks, concerns, and uncertainties surrounding the use of AI in medicine and briefly discusses some of the efforts being made to prepare the health care industry for the implementation of AI.
Artificial intelligence (AI) is a new technical discipline that uses computer technology to research and develop the theory, method, technique, and application system for the simulation, extension, and expansion of human intelligence. With the assistance of new AI technology, the traditional medical environment has changed a lot. For example, a patient’s diagnosis based on radiological, pathological, endoscopic, ultrasonographic, and biochemical examinations has been effectively promoted with a higher accuracy and a lower human workload. The medical treatments during the perioperative period, including the preoperative preparation, surgical period, and postoperative recovery period, have been significantly enhanced with better surgical effects. In addition, AI technology has also played a crucial role in medical drug production, medical management, and medical education, taking them into a new direction. The purpose of this review is to introduce the application of AI in medicine and to provide an outlook of future trends.
Summary This paper is based on a panel discussion held at the Artificial Intelligence in Medicine Europe (AIME) conference in Amsterdam, The Netherlands, in July 2007. It had been more than 15 years since Edward Shortliffe gave a talk at AIME in which he characterized artificial intelligence (AI) in medicine as being in its “adolescence” (Shortliffe EH. The adolescence of AI in medicine: Will the field come of age in the ‘90s? Artificial Intelligence in Medicine 1993; 5:93–106). In this article, the discussants reflect on medical AI research during the subsequent years and attempt to characterize the maturity and influence that has been achieved to date. Participants focus on their personal areas of expertise, ranging from clinical decision making, reasoning under uncertainty, and knowledge representation to systems integration, translational bioinformatics, and cognitive issues in both the modeling of expertise and the creation of acceptable systems.
… parent disciplines of medicine and artificial intelligence. Now, … at the heart of medical practice. The growing emphasis within … Medical artificial intelligence is primarily concerned with the …
The rapid increase of interest in, and use of, artificial intelligence (AI) in computer applications has raised a parallel concern about its ability (or lack thereof) to provide understandable, or explainable, output to users. This concern is especially legitimate in biomedical contexts, where patient safety is of paramount importance. This position paper brings together seven researchers working in the field with different roles and perspectives, to explore in depth the concept of explainable AI, or XAI, offering a functional definition and conceptual framework or model that can be used when considering XAI. This is followed by a series of desiderata for attaining explainability in AI, each of which touches upon a key domain in biomedicine.
Explainable artificial intelligence (AI) is attracting much interest in medicine. Technically, the problem of explainability is as old as AI itself and classic AI represented comprehensible retraceable approaches. However, their weakness was in dealing with uncertainties of the real world. Through the introduction of probabilistic learning, applications became increasingly successful, but increasingly opaque. Explainable AI deals with the implementation of transparency and traceability of statistical black‐box machine learning methods, particularly deep learning (DL). We argue that there is a need to go beyond explainable AI. To reach a level of explainable medicine we need causability. In the same way that usability encompasses measurements for the quality of use, causability encompasses measurements for the quality of explanations. In this article, we provide some necessary definitions to discriminate between explainability and causability as well as a use‐case of DL interpretation and of human explanation in histopathology. The main contribution of this article is the notion of causability, which is differentiated from explainability in that causability is a property of a person, while explainability is a property of a system
After hearing for several decades that computers will soon be able to assist with difficult diagnoses, the practicing physician may well wonder why the revolution has not occurred. …
OBJECTIVE In the last years, several projects promote the secondary use of routine healthcare data based on electronic health record (EHR) data. In multicenter studies, dedicated pseudonymization services are applied for unified pseudonym handling. Healthcare, clinical research and pseudonymization systems are generally disconnected. Hence, the aim of this research work is to integrate these applications and to evaluate the workflow of clinical research. METHODS We analyzed and identified technical solutions for legislation compliant automatic pseudonym generation and for the integration into EHR as well as electronic data capture (EDC) systems. The Mainzelliste was used as pseudonymization service, which is available as open source solution and compliant with the data privacy concept in Germany. Subject of the integration was the local EHR and an in-house developed EDC system. A time and motion study was conducted to evaluate the effects on the workflow. RESULTS Integration of EHR, pseudonymization service and EDC systems is technically feasible and leads to a less fragmented usage of all applications. Generated pseudonyms are obtained from the service hosted at a trusted third party and can now be used in the EDC as well as in the EHR system for direct access and re-identification. The evaluation of 90 registration iterations shows that the time for documentation has been significantly reduced in average by 39.6 s (56.3%) from 71 ± 8 s to 31 ± 5 s per registered study patient. CONCLUSIONS By incorporating EHR, EDC and pseudonymization systems, it is now feasible to support multicenter studies and registers out of an integrated system landscape within a hospital. Optimizing the workflow of patient registration for clinical research allows reduction of double data entry and transcription errors as well as a seamless transition from clinical routine to research data collection.
Background With the exponential expansion of clinical trials conducted in (Brazil, Russia, India, and China) and VISTA (Vietnam, Indonesia, South Africa, Turkey, and Argentina) countries, corresponding gains in cost and enrolment efficiency quickly outpace the consonant metrics in traditional countries in North America and European Union. However, questions still remain regarding the quality of data being collected in these countries. We used ethnographic, mapping and computer simulation studies to identify/address areas of threat to near miss events for data quality in two cancer trial sites in Brazil. Methodology/Principal Findings Two sites in Sao Paolo and Rio Janeiro were evaluated using ethnographic observations of workflow during subject enrolment and data collection. Emerging themes related to threats to near miss events for data quality were derived from observations. They were then transformed into workflows using UML-AD and modeled using System Dynamics. 139 tasks were observed and mapped through the ethnographic study. The UML-AD detected four major activities in the workflow evaluation of potential research subjects prior to signature of informed consent, visit to obtain subject́s informed consent, regular data collection sessions following study protocol and closure of study protocol for a given project. Field observations pointed to three major emerging themes: (a) lack of standardized process for data registration at source document, (b) multiplicity of data repositories and (c) scarcity of decision support systems at the point of research intervention. Simulation with policy model demonstrates a reduction of the rework problem. Conclusions/Significance Patterns of threats to data quality at the two sites were similar to the threats reported in the literature for American sites. The clinical trial site managers need to reorganize staff workflow by using information technology more efficiently, establish new standard procedures and manage professionals to reduce near miss events and save time/cost. Clinical trial sponsors should improve relevant support systems.
… coordinator is paramount to integrating clinical research into your private practice. … clinical workflow adjustments that enable research activities to run efficiently within a busy clinical …
We describe a clinical research visit scheduling system that can potentially coordinate clinical research visits with patient care visits and increase efficiency at clinical sites where clinical and research activities occur simultaneously. Participatory Design methods were applied to support requirements engineering and to create this software called Integrated Model for Patient Care and Clinical Trials (IMPACT). Using a multi-user constraint satisfaction and resource optimization algorithm, IMPACT automatically synthesizes temporal availability of various research resources and recommends the optimal dates and times for pending research visits. We conducted scenario-based evaluations with 10 clinical research coordinators (CRCs) from diverse clinical research settings to assess the usefulness, feasibility, and user acceptance of IMPACT. We obtained qualitative feedback using semi-structured interviews with the CRCs. Most CRCs acknowledged the usefulness of IMPACT features. Support for collaboration within research teams and interoperability with electronic health records and clinical trial management systems were highly requested features. Overall, IMPACT received satisfactory user acceptance and proves to be potentially useful for a variety of clinical research settings. Our future work includes comparing the effectiveness of IMPACT with that of existing scheduling solutions on the market and conducting field tests to formally assess user adoption.
The changing landscape of genomics research and clinical practice has created a need for computational pipelines capable of efficiently orchestrating complex analysis stages while handling large volumes of data across heterogeneous computational environments. Workflow Management Systems (WfMSs) are the software components employed to fill this gap. This work provides an approach and systematic evaluation of key features of popular bioinformatics WfMSs in use today: Nextflow, CWL, and WDL and some of their executors, along with Swift/T, a workflow manager commonly used in high-scale physics applications. We employed two use cases: a variant-calling genomic pipeline and a scalability-testing framework, where both were run locally, on an HPC cluster, and in the cloud. This allowed for evaluation of those four WfMSs in terms of language expressiveness, modularity, scalability, robustness, reproducibility, interoperability, ease of development, along with adoption and usage in research labs and healthcare settings. This article is trying to answer, which WfMS should be chosen for a given bioinformatics application regardless of analysis type?. The choice of a given WfMS is a function of both its intrinsic language and engine features. Within bioinformatics, where analysts are a mix of dry and wet lab scientists, the choice is also governed by collaborations and adoption within large consortia and technical support provided by the WfMS team/community. As the community and its needs continue to evolve along with computational infrastructure, WfMSs will also evolve, especially those with permissive licenses that allow commercial use. In much the same way as the dataflow paradigm and containerization are now well understood to be very useful in bioinformatics applications, we will continue to see innovations of tools and utilities for other purposes, like big data technologies, interoperability, and provenance.
Abstract Objective Integrating clinical research into routine clinical care workflows within electronic health record systems (EHRs) can be challenging, expensive, and labor-intensive. This case study presents a large-scale clinical research project conducted entirely within a commercial EHR during the COVID-19 pandemic. Case Report The UCSD and UCSDH COVID-19 NeutraliZing Antibody Project (ZAP) aimed to evaluate antibody levels to SARS-CoV-2 virus in a large population at an academic medical center and examine the association between antibody levels and subsequent infection diagnosis. Results The project rapidly and successfully enrolled and consented over 2000 participants, integrating the research trial with standing COVID-19 testing operations, staff, lab, and mobile applications. EHR-integration increased enrollment, ease of scheduling, survey distribution, and return of research results at a low cost by utilizing existing resources. Conclusion The case study highlights the potential benefits of EHR-integrated clinical research, expanding their reach across multiple health systems and facilitating rapid learning during a global health crisis.
Background With the globalization of clinical trials, a growing emphasis has been placed on the standardization of the workflow in order to ensure the reproducibility and reliability of the overall trial. Despite the importance of workflow evaluation, to our knowledge no previous studies have attempted to adapt existing modeling languages to standardize the representation of clinical trials. Unified Modeling Language (UML) is a computational language that can be used to model operational workflow, and a UML profile can be developed to standardize UML models within a given domain. This paper's objective is to develop a UML profile to extend the UML Activity Diagram schema into the clinical trials domain, defining a standard representation for clinical trial workflow diagrams in UML. Methods Two Brazilian clinical trial sites in rheumatology and oncology were examined to model their workflow and collect time-motion data. UML modeling was conducted in Eclipse, and a UML profile was developed to incorporate information used in discrete event simulation software. Results Ethnographic observation revealed bottlenecks in workflow: these included tasks requiring full commitment of CRCs, transferring notes from paper to computers, deviations from standard operating procedures, and conflicts between different IT systems. Time-motion analysis revealed that nurses' activities took up the most time in the workflow and contained a high frequency of shorter duration activities. Administrative assistants performed more activities near the beginning and end of the workflow. Overall, clinical trial tasks had a greater frequency than clinic routines or other general activities. Conclusions This paper describes a method for modeling clinical trial workflow in UML and standardizing these workflow diagrams through a UML profile. In the increasingly global environment of clinical trials, the standardization of workflow modeling is a necessary precursor to conducting a comparative analysis of international clinical trials workflows.
Graphical abstract
In this work, we present a conceptual framework to support clinical trial optimization and enrollment workflows and review the current state, limitations, and future trends in this space. This framework includes knowledge representation of clinical trials, clinical trial optimization, clinical trial design, enrollment workflows for prospective clinical trial matching, waitlist management, and, finally, evaluation strategies for assessing improvement.
BackgroundLarge multi-center clinical studies often involve the collection and analysis of biological samples. It is necessary to ensure timely, complete and accurate recording of analytical results and associated phenotypic and clinical information. The TRIBE-AKI Consortium http://www.yale.edu/tribeaki supports a network of multiple related studies and sample biorepository, thus allowing researchers to take advantage of a larger specimen collection than they might have at an individual institution.DescriptionWe describe a biospecimen data management system (BDMS) that supports TRIBE-AKI and is intended for multi-center collaborative clinical studies that involve shipment of biospecimens between sites. This system works in conjunction with a clinical research information system (CRIS) that stores the clinical data associated with the biospecimens, along with other patient-related parameters. Inter-operation between the two systems is mediated by an interactively invoked suite of Web Services, as well as by batch code. We discuss various challenges involved in integration.ConclusionsOur experience indicates that an approach that emphasizes inter-operability is reasonably optimal in allowing each system to be utilized for the tasks for which it is best suited.
This chapter aims to illustrate the methodologies of time and motion research, the observation of clinical care activities in the field and its limits, strengths and opportunities. We discuss how such studies can be used to address questions related to the quality of care and to examine the relationships between clinical workflow and safety. Further, the chapter provides specific examples of the application of time and motion studies, the practical challenges and results obtained.
BACKGROUND Automation of health care workflows has recently become a priority. This can be enabled and enhanced by a workflow monitoring tool (WMOT). OBJECTIVES We shared our experience in clinical workflow analysis via three cases studies in health care and summarized principles to design and develop such a WMOT. METHODS The case studies were conducted in different clinical settings with distinct goals. Each study used at least two types of workflow data to create a more comprehensive picture of work processes and identify bottlenecks, as well as quantify them. The case studies were synthesized using a data science process model with focuses on data input, analysis methods, and findings. RESULTS Three case studies were presented and synthesized to generate a system structure of a WMOT. When developing a WMOT, one needs to consider the following four aspects: (1) goal orientation, (2) comprehensive and resilient data collection, (3) integrated and extensible analysis, and (4) domain experts. DISCUSSION We encourage researchers to investigate the design and implementation of WMOTs and use the tools to create best practices to enable workflow automation and improve workflow efficiency and care quality.
OBJECTIVES: Clinical Research Informatics, an emerging sub-domain of Biomedical Informatics, is currently not well defined. A formal description of CRI including major challenges and opportunities is needed to direct progress in the field. DESIGN: Given the early stage of CRI knowledge and activity, we engaged in a series of qualitative studies with key stakeholders and opinion leaders to determine the range of challenges and opportunities facing CRI. These phases employed complimentary methods to triangulate upon our findings. MEASUREMENTS: Study phases included: 1) a group interview with key stakeholders, 2) an email follow-up survey with a larger group of self-identified CRI professionals, and 3) validation of our results via electronic peer-debriefing and member-checking with a group of CRI-related opinion leaders. Data were collected, transcribed, and organized for formal, independent content analyses by experienced qualitative investigators, followed by an iterative process to identify emergent categorizations and thematic descriptions of the data. RESULTS: We identified a range of challenges and opportunities facing the CRI domain. These included 13 distinct themes spanning academic, practical, and organizational aspects of CRI. These findings also informed the development of a formal definition of CRI and supported further representations that illustrate areas of emphasis critical to advancing the domain. CONCLUSIONS: CRI has emerged as a distinct discipline that faces multiple challenges and opportunities. The findings presented summarize those challenges and opportunities and provide a framework that should help inform next steps to advance this important new discipline.
Milk is a complex biological fluid composed mainly of water, carbohydrates, lipids, proteins, and diverse bioactive factors. Human milk represents a unique tailored source of nutrients that adapts during lactation to the specific needs of the developing infant. Proteins in milk have been studied for decades, and proteomics, peptidomics, and glycoproteomics are the main approaches previously deployed to decipher the proteome of human milk. In the present work, we aimed at implementing a highly automated pipeline for the proteomic analysis of human milk with liquid chromatography mass spectrometry (MS). Commercial human milk samples were used to evaluate and optimize workflows. Centrifugation for defatting milk samples was assessed before and after reduction, alkylation, and enzymatic digestion of proteins, without and with presence of surfactants. Skimmed milk samples were analyzed using isobaric labeling-based quantitative MS on an Orbitrap Tribrid mass spectrometer. Sample fractionation using isoelectric focusing was also evaluated to more deeply profile the human milk proteome. Finally, the most appropriate workflow was transferred to a liquid handling workstation for automated sample preparation. In conclusion, we have defined and describe herein an efficient highly automated proteomic workflow for human milk sample analysis. It is compatible with clinical research, possibly allowing the analysis of sufficiently large cohorts of samples.
Background The reporting of serious adverse events is a requirement when conducting a clinical trial involving human subjects, necessary for the protection of the participants. The reporting process is a multi-step procedure, involving a number of individuals from initiation to final review, and must be completed in a timely fashion. Purpose The purpose of this project was to automate the adverse event reporting process, replacing paper-based processes with computer-based processes, so that personnel effort and time required for serious adverse event reporting was reduced, and the monitoring of reporting performance and adverse event characteristics was facilitated. Methods Use case analysis was employed to understand the reporting workflow and generate software requirements. The automation of the workflow was then implemented, employing computer databases, web-based forms, electronic signatures, and email communication. Results In the initial year (2007) of full deployment, 588 SAE reports were processed by the automated system, eSAEyTM. The median time from initiation to Principal Investigator electronic signature was <2 days (mean 7 ± 0.7 days). This was a significant reduction from the prior paper-based system, which had a median time for signature of 24 days (mean of 45 ± 5.7 days). With eSAEyTM, reports on adverse event characteristics (type, grade, etc.) were easily obtained and had consistent values based on standard terminologies. Limitation The automated system described was designed specifically for the workflow at Thomas Jefferson University. While the methodology for system design, and the system requirements derived from common clinical trials adverse reporting procedures are applicable in general, specific workflow details may not be relevant at other institutions. Conclusion The system facilitated analysis of individual investigator reporting performance, as well as the aggregation and analysis of the nature of reported adverse events. Clinical Trials 2009; 6: 446—454. http://ctj.sagepub.com
Abstract Workflow is a commonly used term in human factors and informatics literature. We embraced a broad definition and used it as a concept to examine various work phenomena. This chapter aims to discuss how design can improve workflow in health settings and eventually make care delivery patient-centered and safer. In general, design studies can improve workflow in two different ways: (1) the overall workflow and (2) interventions that would improve workflow. We highlighted two design approaches: user-centered design and participatory. These are two closely related approaches that engage (to varying degrees) targeted users along a continuum of participation to improve usability. For each of the three informatics subfields (clinical, public health, and consumer health), we provide relevant examples from studies, where users were engaged in the design process. We have identified whether the notion of workflow was formally addressed and gave examples of how workflow methodology might complement user-centered and participatory design efforts in clinical, public health, and consumer health informatics research. Although user-centered design includes a variety of methods, we introduced contextual inquiry and participatory design. These designs demonstrated two key distinctive features of user-centered design: the importance of understanding the holistic context of users’ social, technical, and cultural environments; and engaging users as codesigners. We deepen our discussion through a case that focuses on supporting patient engagement for improved intra- and cross-institutional workflow in an emergency department. Design studies that involve multidisciplinary perspectives are necessary to improve workflow that contributes to safer, higher quality, and more accessible care delivery.
… simulations and virtual clinical trials, which could accelerate clinical research by facilitating valuable risk–reward inferences that could inform researchers about which studies are most …
The deployment of large language models (LLMs) within the healthcare sector has sparked both enthusiasm and apprehension. These models exhibit the remarkable ability to provide proficient responses to free-text queries, demonstrating a nuanced understanding of professional medical knowledge. This comprehensive survey delves into the functionalities of existing LLMs designed for healthcare applications and elucidates the trajectory of their development, starting with traditional Pretrained Language Models (PLMs) and then moving to the present state of LLMs in the healthcare sector. First, we explore the potential of LLMs to amplify the efficiency and effectiveness of diverse healthcare applications, particularly focusing on clinical language understanding tasks. These tasks encompass a wide spectrum, ranging from named entity recognition and relation extraction to natural language inference, multimodal medical applications, document classification, and question-answering. Additionally, we conduct an extensive comparison of the most recent state-of-the-art LLMs in the healthcare domain, while also assessing the utilization of various open-source LLMs and highlighting their significance in healthcare applications. Furthermore, we present the essential performance metrics employed to evaluate LLMs in the biomedical domain, shedding light on their effectiveness and limitations. Finally, we summarize the prominent challenges and constraints faced by large language models in the healthcare sector by offering a holistic perspective on their potential benefits and shortcomings. This review provides a comprehensive exploration of the current landscape of LLMs in healthcare, addressing their role in transforming medical applications and the areas that warrant further research and development.
Large language models (LLMs) are artificial intelligence (AI) tools specifically trained to process and generate text. LLMs attracted substantial public attention after OpenAI’s ChatGPT was made publicly available in November 2022. LLMs can often answer questions, summarize, paraphrase and translate text on a level that is nearly indistinguishable from human capabilities. The possibility to actively interact with models like ChatGPT makes LLMs attractive tools in various fields, including medicine. While these models have the potential to democratize medical knowledge and facilitate access to healthcare, they could equally distribute misinformation and exacerbate scientific misconduct due to a lack of accountability and transparency. In this article, we provide a systematic and comprehensive overview of the potentials and limitations of LLMs in clinical practice, medical research and medical education.
… decision-making, including diagnosis, prognosis, treatment suggestion, risk prediction and clinical trial matching, relies heavily on the synthesis and interpretation of vast amounts of …
As clinical trials scale up and grow more complex, researchers are facing mounting challenges, including inefficient participant recruitment, complex data management, and limited risk monitoring. These issues not only increase the workload for clinical researchers but also compromise trial reliability and safety, potentially elevating the risk of trial failure. Large language models (LLMs), as an emerging technology in natural language processing (NLP), exhibit notable advantages across various tasks, such as information extraction and relation classification. With domain-specific pre-training and fine-tuning, LLMs present promising potential in clinical trial tasks such as automated patient-trial matching and the extraction and processing of trial data, which are anticipated to reduce time and financial costs. Additionally, they offer valuable insights for scientific rationale, medical decision-making, and trial endpoint prediction. In this context, an increasing number of studies have begun to explore the applications of LLMs in the design and conduct of clinical trials. This paper provides a review of LLM applications in clinical trials with an emphasis on real-world integration. Comparative advantages over traditional NLP models, technical limitations, and future implementation challenges are also discussed. This narrative review aims to highlight the potential of LLMs in clinical trial workflows and clarify key challenges and future directions.
Med-PaLM, a state-of-the-art large language model for medicine, is introduced and evaluated across several medical question answering tasks, demonstrating the promise of these models in this domain. Large language models (LLMs) have demonstrated impressive capabilities, but the bar for clinical applications is high. Attempts to assess the clinical knowledge of models typically rely on automated evaluations based on limited benchmarks. Here, to address these limitations, we present MultiMedQA, a benchmark combining six existing medical question answering datasets spanning professional medicine, research and consumer queries and a new dataset of medical questions searched online, HealthSearchQA. We propose a human evaluation framework for model answers along multiple axes including factuality, comprehension, reasoning, possible harm and bias. In addition, we evaluate Pathways Language Model^ 1 (PaLM, a 540-billion parameter LLM) and its instruction-tuned variant, Flan-PaLM^ 2 on MultiMedQA. Using a combination of prompting strategies, Flan-PaLM achieves state-of-the-art accuracy on every MultiMedQA multiple-choice dataset (MedQA^ 3 , MedMCQA^ 4 , PubMedQA^ 5 and Measuring Massive Multitask Language Understanding (MMLU) clinical topics^ 6 ), including 67.6% accuracy on MedQA (US Medical Licensing Exam-style questions), surpassing the prior state of the art by more than 17%. However, human evaluation reveals key gaps. To resolve this, we introduce instruction prompt tuning, a parameter-efficient approach for aligning LLMs to new domains using a few exemplars. The resulting model, Med-PaLM, performs encouragingly, but remains inferior to clinicians. We show that comprehension, knowledge recall and reasoning improve with model scale and instruction prompt tuning, suggesting the potential utility of LLMs in medicine. Our human evaluations reveal limitations of today’s models, reinforcing the importance of both evaluation frameworks and method development in creating safe, helpful LLMs for clinical applications.
There are enormous enthusiasm and concerns in applying large language models (LLMs) to healthcare. Yet current assumptions are based on general-purpose LLMs such as ChatGPT, which are not developed for medical use. This study develops a generative clinical LLM, GatorTronGPT, using 277 billion words of text including (1) 82 billion words of clinical text from 126 clinical departments and approximately 2 million patients at the University of Florida Health and (2) 195 billion words of diverse general English text. We train GatorTronGPT using a GPT-3 architecture with up to 20 billion parameters and evaluate its utility for biomedical natural language processing (NLP) and healthcare text generation. GatorTronGPT improves biomedical natural language processing. We apply GatorTronGPT to generate 20 billion words of synthetic text. Synthetic NLP models trained using synthetic text generated by GatorTronGPT outperform models trained using real-world clinical text. Physicians’ Turing test using 1 (worst) to 9 (best) scale shows that there are no significant differences in linguistic readability ( p = 0.22; 6.57 of GatorTronGPT compared with 6.93 of human) and clinical relevance ( p = 0.91; 7.0 of GatorTronGPT compared with 6.97 of human) and that physicians cannot differentiate them ( p < 0.001). This study provides insights into the opportunities and challenges of LLMs for medical research and healthcare.
Large language models (LLMs) like OpenAI’s ChatGPT are powerful generative systems that rapidly synthesize natural language responses. Research on LLMs has revealed their potential and pitfalls, especially in clinical settings. However, the evolving landscape of LLM research in medicine has left several gaps regarding their evaluation, application, and evidence base. This scoping review aims to (1) summarize current research evidence on the accuracy and efficacy of LLMs in medical applications, (2) discuss the ethical, legal, logistical, and socioeconomic implications of LLM use in clinical settings, (3) explore barriers and facilitators to LLM implementation in healthcare, (4) propose a standardized evaluation framework for assessing LLMs’ clinical utility, and (5) identify evidence gaps and propose future research directions for LLMs in clinical applications. We screened 4,036 records from MEDLINE, EMBASE, CINAHL, medRxiv, bioRxiv, and arXiv from January 2023 (inception of the search) to June 26, 2023 for English-language papers and analyzed findings from 55 worldwide studies. Quality of evidence was reported based on the Oxford Centre for Evidence-based Medicine recommendations. Our results demonstrate that LLMs show promise in compiling patient notes, assisting patients in navigating the healthcare system, and to some extent, supporting clinical decision-making when combined with human oversight. However, their utilization is limited by biases in training data that may harm patients, the generation of inaccurate but convincing information, and ethical, legal, socioeconomic, and privacy concerns. We also identified a lack of standardized methods for evaluating LLMs’ effectiveness and feasibility. This review thus highlights potential future directions and questions to address these limitations and to further explore LLMs’ potential in enhancing healthcare delivery.
A knowledge gap persists between machine learning (ML) developers (e.g., data scientists) and practitioners (e.g., clinicians), hampering the full utilization of ML for clinical data analysis. We investigated the potential of the ChatGPT Advanced Data Analysis (ADA), an extension of GPT-4, to bridge this gap and perform ML analyses efficiently. Real-world clinical datasets and study details from large trials across various medical specialties were presented to ChatGPT ADA without specific guidance. ChatGPT ADA autonomously developed state-of-the-art ML models based on the original study’s training data to predict clinical outcomes such as cancer development, cancer progression, disease complications, or biomarkers such as pathogenic gene sequences. Following the re-implementation and optimization of the published models, the head-to-head comparison of the ChatGPT ADA-crafted ML models and their respective manually crafted counterparts revealed no significant differences in traditional performance metrics (p ≥ 0.072). Strikingly, the ChatGPT ADA-crafted ML models often outperformed their counterparts. In conclusion, ChatGPT ADA offers a promising avenue to democratize ML in medicine by simplifying complex data analyses, yet should enhance, not replace, specialized training and resources, to promote broader applications in medical research and practice. A knowledge gap persists between machine learning developers and clinicians. Here, the authors show that the Advanced Data Analysis extension of ChatGPT could bridge this gap and simplify complex data analyses, making them more accessible to clinicians.
This review analyzes current clinical trials investigating large language models’ (LLMs) applications in healthcare. We identified 27 trials (5 published and 22 ongoing) across 4 main clinical applications: patient care, data handling, decision support, and research assistance. Our analysis reveals diverse LLM uses, from clinical documentation to medical decision-making. Published trials show promise but highlight accuracy concerns. Ongoing studies explore novel applications like patient education and informed consent. Most trials occur in the United States of America and China. We discuss the challenges of evaluating rapidly evolving LLMs through clinical trials and identify gaps in current research. This review aims to inform future studies and guide the integration of LLMs into clinical practice.
… has been the quest of artificial intelligence (AI) research for decades. The seemingly abrupt advent of readily accessible, large language models (LLMs) has been greeted with great, …
Clinical trials are essential for medical research, but they often face challenges in matching patients to trials and planning. Large language models (LLMs) offer a promising solution, signaling a transformative shift in the field of clinical trials. This review explores the multifaceted applications of LLMs within clinical trials, focusing on five main areas expected to be implemented in the near future: enhancing patient-trial matching, streamlining clinical trial planning, analyzing free text narratives for coding and classification, assisting in technical writing tasks, and providing cognizant consent via LLM-powered chatbots. While the application of LLMs is promising, it poses challenges such as accuracy validation and legal concerns. The convergence of LLMs with clinical trials has the potential to revolutionize the efficiency of clinical trials, paving the way for innovative methodologies and enhancing patient engagement. However, this development requires careful consideration and investment to overcome potential hurdles.
Background: Large Language Models (LLMs) are reshaping medical research workflows. Objective: This narrative review synthesizes evidence on LLM applications across systematic reviews, scientific writing, and clinical research. Methods: We reviewed literature from 2023–2025 examining LLM applications in medical research, identified through PubMed, Scopus, Web of Science, arXiv, medRxiv, and Google Scholar. Studies reporting empirical findings, methodological evaluations, or systematic analyses of LLM applications were included; editorials and commentaries without empirical data were excluded. Results: In systematic reviews, LLMs achieve 80–94% data extraction accuracy and 40% reduction in screening workload, but show only slight-to-moderate agreement (κ = 0.16–0.43) in risk-of-bias assessment. In scientific writing, hallucination rates of 47–55% for fabricated references and over 90% prevalence of demographic bias require rigorous verification. For clinical research, LLMs assist with statistical coding and protocol development but require human validation. Critically, excessive reliance on automated tools may cause cognitive offloading that compromises analytical capabilities. Conclusions: LLMs are powerful but unstable tools requiring constant verification. Success depends on maintaining human-in-the-loop approaches that preserve critical thinking while leveraging AI efficiency.
Large language models (LLMs) are a specialized type of generative artificial intelligence (AI… recommend future research directions. Box 1 lists the relevant large language model terms …
… The use of a large language model has greatly assisted us in rephrasing and ensuring the clarity and effectiveness of our language, and highlights the potential of language models in …
OBJECTIVE: The objective of this study is to systematically examine the efficacy of both proprietary (GPT-3.5, GPT-4) and open-source large language models (LLMs) (LLAMA 7B, 13B, 70B) in the context of matching patients to clinical trials in healthcare. MATERIALS AND METHODS: The study employs a multifaceted evaluation framework, incorporating extensive automated and human-centric assessments along with a detailed error analysis for each model, and assesses LLMs' capabilities in analyzing patient eligibility against clinical trial's inclusion and exclusion criteria. To improve the adaptability of open-source LLMs, a specialized synthetic dataset was created using GPT-4, facilitating effective fine-tuning under constrained data conditions. RESULTS: The findings indicate that open-source LLMs, when fine-tuned on this limited and synthetic dataset, achieve performance parity with their proprietary counterparts, such as GPT-3.5. DISCUSSION: This study highlights the recent success of LLMs in the high-stakes domain of healthcare, specifically in patient-trial matching. The research demonstrates the potential of open-source models to match the performance of proprietary models when fine-tuned appropriately, addressing challenges like cost, privacy, and reproducibility concerns associated with closed-source proprietary LLMs. CONCLUSION: The study underscores the opportunity for open-source LLMs in patient-trial matching. To encourage further research and applications in this field, the annotated evaluation dataset and the fine-tuned LLM, Trial-LLAMA, are released for public use.
Background/Aims Clinical trials require numerous documents to be written: Protocols, consent forms, clinical study reports, and many others. Large language models offer the potential to rapidly generate first-draft versions of these documents; however, there are concerns about the quality of their output. Here, we report an evaluation of how good large language models are at generating sections of one such document, clinical trial protocols. Methods Using an off-the-shelf large language model, we generated protocol sections for a broad range of diseases and clinical trial phases. Each of these document sections we assessed across four dimensions: Clinical thinking and logic; Transparency and references; Medical and clinical terminology; and Content relevance and suitability. To improve performance, we used the retrieval-augmented generation method to enhance the large language model with accurate up-to-date information, including regulatory guidance documents and data from ClinicalTrials.gov. Using this retrieval-augmented generation large language model, we regenerated the same protocol sections and assessed them across the same four dimensions. Results We find that the off-the-shelf large language model delivers reasonable results, especially when assessing content relevance and the correct use of medical and clinical terminology, with scores of over 80%. However, the off-the-shelf large language model shows limited performance in clinical thinking and logic and transparency and references, with assessment scores of ≈40% or less. The use of retrieval-augmented generation substantially improves the writing quality of the large language model, with clinical thinking and logic and transparency and references scores increasing to ≈80%. The retrieval-augmented generation method thus greatly improves the practical usability of large language models for clinical trial-related writing. Discussion Our results suggest that hybrid large language model architectures, such as the retrieval-augmented generation method we utilized, offer strong potential for clinical trial-related writing, including a wide variety of documents. This is potentially transformative, since it addresses several major bottlenecks of drug development.
Abstract Objective Systematic reviews depend on rigorous risk-of-bias (RoB) assessments to ensure credibility, yet manual evaluation using the Cochrane RoB 2 tool is resource-intensive. While large language models (LLMs) offer potential for automation, their alignment with human judgment remains underexplored. This study evaluates the reliability of ChatGPT-4o, ChatGPT-5, and Claude 3.5 Sonnet in assessing RoB in randomized controlled trials (RCTs), comparing their agreement with human reviewers and internal consistency. Study Design We retrospectively analyzed 180 RCTs from systematic reviews published in the American Journal of Obstetrics and Gynecology (2021–2023) reporting complete human RoB 2 ratings. Each LLM processed full-text PDFs using a standardized prompt incorporating the complete RoB 2 algorithm. Model performance was evaluated against human benchmarks using Cohen's kappa and prevalence- and bias-adjusted kappa. Intramodel reliability was assessed across three independent runs to measure consistency. Results ChatGPT-5 consistently outperformed other models, achieving the highest agreement in randomization (Domain 1; 76%), missing outcome data (Domain 3; 80%), and outcome measurement (Domain 4; 76%). It showed moderate concordance for deviations from intended interventions (69%). However, all models struggled with selective reporting (Domain 5), where agreement dropped to 47 to 51%. For overall RoB judgments, ChatGPT-5 demonstrated superior concordance (60–62%, κ = 0.36–0.40) compared with ChatGPT-4o (45%) and Claude 3.5 Sonnet (43%). ChatGPT-5 also exhibited substantial to near-perfect internal consistency. Conclusion Among the evaluated models, ChatGPT-5 most closely approximated human RoB 2 assessments and achieved superior internal consistency, suggesting it could serve as a practical first-pass tool to reduce reviewer burden. However, persistent limitations in detecting selective reporting—likely due to the inability to cross-reference external trial registries—highlight that expert human oversight remains essential for accurate evidence synthesis. Key Points GPT-5, GPT-4o, and Claude evaluated 180 RCTs. GPT-5 outperformed GPT-4o and Claude models. Models struggled with selective reporting bias.
Alternatenobaric vertigo (ABV) develops when the middle ear pressure (MEP) is not equal at the same height in the sea or the air. This is possible when the altitude changes. Eustachian tube dysfunction (ETD) is a common cause of ABV. In this case report, we discuss a patient who experienced repeated bouts of ground-level alternobaric vertigo (GLABV) due to ETD. We also discuss how Conversational Generative Pre-trained Transformer (ChatGPT) might be used in the creation of this case report. A 41-year-old male patient complained of vertigo at ground level on several occasions. His medical history included chronic sinusitis, nasal congestion, and laryngopharyngeal reflux (LPR). During the physical exam, his tympanic membranes were dull and moved less. Tympanometry showed that he had an asymmetric type A and that both of his middle ears had negative pressure. The results of the audiometry test were normal, and the laryngoscopy revealed LPR. The patient was found to have GLABV because of ETD, and different treatment options, such as Eustachian tube catheterization (ETC), were thought about. This case study demonstrates how ChatGPT can be used to assist with medical documentation and the treatment of GLABV caused by ETD. Even though ChatGPT did not provide specific diagnostic or treatment recommendations for the patient's condition, it did assist the doctor in determining what was wrong and how to treat it while writing the case report. It also aided the doctor in writing the case report by allowing them to discuss it. The use of artificial intelligence (AI) tools such as ChatGPT has the potential to improve the accuracy and speed of medical documentation, thereby streamlining clinical workflows and improving patient care. Nonetheless, it is critical to consider the ethical implications of using AI in clinical practice This case study emphasizes the importance of understanding that ETD is a common cause of GLABV and how ChatGPT can aid in the diagnosis and treatment of this condition. More research is needed to fully understand how long-term AI interventions in medicine work and how reliable they are.
Clinical research and practice are generating important new findings at exponential rate which need to be readily available to clinicians. However, clinicians are confronted with serious challenges when they try to seek such information for their evidence-based decision making or to generate new clinical case report. One important challenge is the long time needed to browse, filter, summarize and compile information from different resources. The other important challenge is to identify relevant important evidence-based information resources required to answer clinical questions or support a clinical finding. Artificial intelligence can help in solving both challenges based on the automatic question answering (Q&A) and generative technologies. However, Q&A and generative techniques are not trained to answer clinical queries that can be used for evidence-based practice nor it can respond to structured clinical questioning protocol like PICO (Patient/Problem, Intervention, Comparison and Outcome). This article describes the use of deep learning techniques for Q&A that is based on generative models like BERT and GPT to answer PICO clinical questions that can be used for evidence-based practice extracted from sound medical research resources like PubMed. We are reporting acceptable clinical answers that are supported by findings from PubMed. Our generative methods are reaching state of the art performance based on two staged bootstrapping process involving filtering relevant articles followed by identifying articles that support the requested outcome expressed by the PICO question.
The purposes were to assess the efficacy of AI-generated radiology reports in terms of report summary, patient-friendliness, and recommendations and to evaluate the consistent performance of report quality and accuracy, contributing to the advancement of radiology workflow. Total 685 spine MRI reports were retrieved from our hospital database. AI-generated radiology reports were generated in three formats: (1) summary reports, (2) patient-friendly reports, and (3) recommendations. The occurrence of artificial hallucinations was evaluated in the AI-generated reports. Two radiologists conducted qualitative and quantitative assessments considering the original report as a standard reference. Two non-physician raters assessed their understanding of the content of original and patient-friendly reports using a 5-point Likert scale. The scoring of the AI-generated radiology reports were overall high average scores across all three formats. The average comprehension score for the original report was 2.71 ± 0.73, while the score for the patient-friendly reports significantly increased to 4.69 ± 0.48 (p < 0.001). There were 1.12% artificial hallucinations and 7.40% potentially harmful translations. In conclusion, the potential benefits of using generative AI assistants to generate these reports include improved report quality, greater efficiency in radiology workflow for producing summaries, patient-centered reports, and recommendations, and a move toward patient-centered radiology.
Manually converting unstructured text pathology reports into structured pathology reports is very time-consuming and prone to errors. This study demonstrates the transformative potential of generative AI in automating the analysis of free-text pathology reports. Employing the ChatGPT Large Language Model within a Streamlit web application, we automated the extraction and structuring of information from 33 unstructured breast cancer pathology reports from Taipei Medical University Hospital. Achieving a 99.61% accuracy rate, the AI system notably reduced the processing time compared to traditional methods. This not only underscores the efficacy of AI in converting unstructured medical text into structured data but also highlights its potential to enhance the efficiency and reliability of medical text analysis. However, this study is limited to breast cancer pathology reports and was conducted using data obtained from hospitals associated with a single institution. In the future, we plan to expand the scope of this research to include pathology reports for other cancer types incrementally and conduct external validation to further substantiate the robustness and generalizability of the proposed system. Through this technological integration, we aimed to substantiate the capabilities of generative AI in improving both the speed and reliability of data processing. The outcomes of this study affirm that generative AI can significantly transform the handling of pathology reports, promising substantial advancements in biomedical research by facilitating the structured analysis of complex medical data.
Generative artificial intelligence (GAI) can be broadly described as an artificial intelligence system capable of generating images, text, and other media types with human prompts. GAI models like ChatGPT, DALL-E, and Bard have recently caught the attention of industry and academia equally. GAI applications span various industries like art, gaming, fashion, and healthcare. In healthcare, GAI shows promise in medical research, diagnosis, treatment, and patient care and is already making strides in real-world deployments. There has yet to be any detailed study concerning the applications and scope of GAI in healthcare. Addressing this research gap, we explore several applications, real-world scenarios, and limitations of GAI in healthcare. We examine how GAI models like ChatGPT and DALL-E can be leveraged to aid in the applications of medical imaging, drug discovery, personalized patient treatment, medical simulation and training, clinical trial optimization, mental health support, healthcare operations and research, medical chatbots, human movement simulation, and a few more applications. Along with applications, we cover four real-world healthcare scenarios that employ GAI: visual snow syndrome diagnosis, molecular drug optimization, medical education, and dentistry. We also provide an elaborate discussion on seven healthcare-customized LLMs like Med-PaLM, BioGPT, DeepHealth, etc.,Since GAI is still evolving, it poses challenges like the lack of professional expertise in decision making, risk of patient data privacy, issues in integrating with existing healthcare systems, and the problem of data bias which are elaborated on in this work along with several other challenges. We also put forward multiple directions for future research in GAI for healthcare.
Background Multimodal generative artificial intelligence (AI) technologies can produce preliminary radiology reports, and validation with reader studies is crucial for understanding the clinical value of these technologies. Purpose To assess the clinical value of the use of a domain-specific multimodal generative AI tool for chest radiograph interpretation by means of a reader study. Materials and Methods A retrospective, sequential, multireader, multicase reader study was conducted using 758 chest radiographs from a publicly available dataset from 2009 to 2017. Five radiologists interpreted the chest radiographs in two sessions: without AI-generated reports and with AI-generated reports as preliminary reports. Reading times, reporting agreement (RADPEER), and quality scores (five-point scale) were evaluated by two experienced thoracic radiologists and compared between the first and second sessions from October to December 2023. Reading times, report agreement, and quality scores were analyzed using a generalized linear mixed model. Additionally, a subset of 258 chest radiographs was used to assess the factual correctness of the reports, and sensitivities and specificities were compared between the reports from the first and second sessions with use of the McNemar test. Results The introduction of AI-generated reports significantly reduced average reading times from 34.2 seconds ± 20.4 to 19.8 seconds ± 12.5 (P < .001). Report agreement scores shifted from a median of 5.0 (IQR, 4.0-5.0) without AI reports to 5.0 (IQR, 4.5-5.0) with AI reports (P < .001). Report quality scores changed from 4.5 (IQR, 4.0-5.0) without AI reports to 4.5 (IQR, 4.5-5.0) with AI reports (P < .001). From the subset analysis of factual correctness, the sensitivity for detecting various abnormalities increased significantly, including widened mediastinal silhouettes (84.3% to 90.8%; P < .001) and pleural lesions (77.7% to 87.4%; P < .001). While the overall diagnostic performance improved, variability among individual radiologists was noted. Conclusion The use of a domain-specific multimodal generative AI model increased the efficiency and quality of radiology report generation. © RSNA, 2025 Supplemental material is available for this article. See also the editorial by Babyn and Adams in this issue.
Large language models, specifically ChatGPT, are revolutionizing clinical research by improving content creation and providing specific useful features. These technologies can transform clinical research, including data collection, analysis, interpretation, and results sharing. However, integrating these technologies into the academic writing workflow poses significant challenges. In this review, I investigated the integration of large-language model-based AI tools into clinical research, focusing on practical implementation strategies and addressing the ethical considerations associated with their use. Additionally, I provide examples of the safe and sound use of generative AI in clinical research and emphasize the need to ensure that AI-generated outputs are reliable and valid in scholarly writing settings. In conclusion, large language models are a powerful tool for organizing and expressing ideas efficiently; however, they have limitations. Writing an academic paper requires critical analysis and intellectual input from the authors. Moreover, AI-generated text must be carefully reviewed to reflect the authors’ insights. These AI tools significantly enhance the efficiency of repetitive research tasks, although challenges related to plagiarism detection and ethical use persist.
Systematic reviews, a cornerstone of evidence-based medicine, are not produced quickly enough to support clinical practice. The cost of production, availability of the requisite expertise and timeliness are often quoted as major contributors for the delay. This detailed survey of the state of the art of information systems designed to support or automate individual tasks in the systematic review, and in particular systematic reviews of randomized controlled clinical trials, reveals trends that see the convergence of several parallel research projects.We surveyed literature describing informatics systems that support or automate the processes of systematic review or each of the tasks of the systematic review. Several projects focus on automating, simplifying and/or streamlining specific tasks of the systematic review. Some tasks are already fully automated while others are still largely manual. In this review, we describe each task and the effect that its automation would have on the entire systematic review process, summarize the existing information system support for each task, and highlight where further research is needed for realizing automation for the task. Integration of the systems that automate systematic review tasks may lead to a revised systematic review workflow. We envisage the optimized workflow will lead to system in which each systematic review is described as a computer program that automatically retrieves relevant trials, appraises them, extracts and synthesizes data, evaluates the risk of bias, performs meta-analysis calculations, and produces a report in real time.
OBJECTIVE We aim to investigate the application and accuracy of artificial intelligence (AI) methods for automated medical literature screening for systematic reviews. MATERIALS AND METHODS We systematically searched PubMed, Embase, and IEEE Xplore Digital Library to identify potentially relevant studies. We included studies in automated literature screening that reported study question, source of dataset, and developed algorithm models for literature screening. The literature screening results by human investigators were considered to be the reference standard. Quantitative synthesis of the accuracy was conducted using a bivariate model. RESULTS Eighty-six studies were included in our systematic review and 17 studies were further included for meta-analysis. The combined recall, specificity, and precision were 0.928 [95% confidence interval (CI), 0.878-0.958], 0.647 (95% CI, 0.442-0.809), and 0.200 (95% CI, 0.135-0.287) when achieving maximized recall, but were 0.708 (95% CI, 0.570-0.816), 0.921 (95% CI, 0.824-0.967), and 0.461 (95% CI, 0.375-0.549) when achieving maximized precision in the AI models. No significant difference was found in recall among subgroup analyses including the algorithms, the number of screened literatures, and the fraction of included literatures. DISCUSSION AND CONCLUSION This systematic review and meta-analysis study showed that the recall is more important than the specificity or precision in literature screening, and a recall over 0.95 should be prioritized. We recommend to report the effectiveness indices of automatic algorithms separately. At the current stage manual literature screening is still indispensable for medical systematic reviews.
Background Systematic review is an indispensable tool for optimal evidence collection and evaluation in evidence-based medicine. However, the explosive increase of the original literatures makes it difficult to accomplish critical appraisal and regular update. Artificial intelligence (AI) algorithms have been applied to automate the literature screening procedure in medical systematic reviews. In these studies, different algorithms were used and results with great variance were reported. It is therefore imperative to systematically review and analyse the developed automatic methods for literature screening and their effectiveness reported in current studies. Methods An electronic search will be conducted using PubMed, Embase, ACM Digital Library, and IEEE Xplore Digital Library databases, as well as literatures found through supplementary search in Google scholar, on automatic methods for literature screening in systematic reviews. Two reviewers will independently conduct the primary screening of the articles and data extraction, in which nonconformities will be solved by discussion with a methodologist. Data will be extracted from eligible studies, including the basic characteristics of study, the information of training set and validation set, and the function and performance of AI algorithms, and summarised in a table. The risk of bias and applicability of the eligible studies will be assessed by the two reviewers independently based on Quality Assessment of Diagnostic Accuracy Studies (QUADAS-2). Quantitative analyses, if appropriate, will also be performed. Discussion Automating systematic review process is of great help in reducing workload in evidence-based practice. Results from this systematic review will provide essential summary of the current development of AI algorithms for automatic literature screening in medical evidence synthesis and help to inspire further studies in this field. Systematic review registration PROSPERO CRD42020170815 (28 April 2020).
Background The demand for high-quality systematic literature reviews (SRs) for evidence-based medical decision-making is growing. SRs are costly and require the scarce resource of highly skilled reviewers. Automation technology has been proposed to save workload and expedite the SR workflow. We aimed to provide a comprehensive overview of SR automation studies indexed in PubMed, focusing on the applicability of these technologies in real world practice. Methods In November 2022, we extracted, combined, and ran an integrated PubMed search for SRs on SR automation. Full-text English peer-reviewed articles were included if they reported studies on SR automation methods (SSAM), or automated SRs (ASR). Bibliographic analyses and knowledge-discovery studies were excluded. Record screening was performed by single reviewers, and the selection of full text papers was performed in duplicate. We summarized the publication details, automated review stages, automation goals, applied tools, data sources, methods, results, and Google Scholar citations of SR automation studies. Results From 5321 records screened by title and abstract, we included 123 full text articles, of which 108 were SSAM and 15 ASR. Automation was applied for search (19/123, 15.4%), record screening (89/123, 72.4%), full-text selection (6/123, 4.9%), data extraction (13/123, 10.6%), risk of bias assessment (9/123, 7.3%), evidence synthesis (2/123, 1.6%), assessment of evidence quality (2/123, 1.6%), and reporting (2/123, 1.6%). Multiple SR stages were automated by 11 (8.9%) studies. The performance of automated record screening varied largely across SR topics. In published ASR, we found examples of automated search, record screening, full-text selection, and data extraction. In some ASRs, automation fully complemented manual reviews to increase sensitivity rather than to save workload. Reporting of automation details was often incomplete in ASRs. Conclusions Automation techniques are being developed for all SR stages, but with limited real-world adoption. Most SR automation tools target single SR stages, with modest time savings for the entire SR process and varying sensitivity and specificity across studies. Therefore, the real-world benefits of SR automation remain uncertain. Standardizing the terminology, reporting, and metrics of study reports could enhance the adoption of SR automation techniques in real-world practice.
Introduction Automated machine learning (autoML) removes technical and technological barriers to building artificial intelligence models. We aimed to summarise the clinical applications of autoML, assess the capabilities of utilised platforms, evaluate the quality of the evidence trialling autoML, and gauge the performance of autoML platforms relative to conventionally developed models, as well as each other. Methods This review adhered to a PROSPERO-registered protocol (CRD42022344427). The Cochrane Library, Embase, MEDLINE, and Scopus were searched from inception to 11 July 2022. Two researchers screened abstracts and full texts, extracted data and conducted quality assessment. Disagreement was resolved through discussion and as-required arbitration by a third researcher. Results In 82 studies, 26 distinct autoML platforms featured. Brain and lung disease were the most common fields of study of 22 specialties. AutoML exhibited variable performance: AUCROC 0.35-1.00, F1-score 0.16-0.99, AUCPR 0.51-1.00. AutoML exhibited the highest AUCROC in 75.6% trials; the highest F1-score in 42.3% trials; and the highest AUCPRC in 83.3% trials. In autoML platform comparisons, AutoPrognosis and Amazon Rekognition performed strongest with unstructured and structured data respectively. Quality of reporting was poor, with a median DECIDE-AI score of 14 of 27. Conclusions A myriad of autoML platforms have been applied in a variety of clinical contexts. The performance of autoML compares well to bespoke computational and clinical benchmarks. Further work is required to improve the quality of validation studies. AutoML may facilitate a transition to data-centric development, and integration with large language models may enable AI to build itself to fulfil user-defined goals.
Background Systematic reviews (SRs) are considered the highest level of evidence to answer research questions; however, they are time and resource intensive. Objective When comparing SR tasks done manually, using standard methods, versus those same SR tasks done using automated tools, (1) what is the difference in time to complete the SR task and (2) what is the impact on the error rate of the SR task? Methods A case study compared specific tasks done during the conduct of an SR on prebiotic, probiotic, and synbiotic supplementation in chronic kidney disease. Two participants (manual team) conducted the SR using current methods, comprising a total of 16 tasks. Another two participants (automation team) conducted the tasks where a systematic review automation (SRA) tool was available, comprising of a total of six tasks. The time taken and error rate of the six tasks that were completed by both teams were compared. Results The approximate time for the manual team to produce a draft of the background, methods, and results sections of the SR was 126 hours. For the six tasks in which times were compared, the manual team spent 2493 minutes (42 hours) on the tasks, compared to 708 minutes (12 hours) spent by the automation team. The manual team had a higher error rate in two of the six tasks—regarding Task 5: Run the systematic search, the manual team made eight errors versus three errors made by the automation team; regarding Task 12: Assess the risk of bias, 25 assessments differed from a reference standard for the manual team compared to 20 differences for the automation team. The manual team had a lower error rate in one of the six tasks—regarding Task 6: Deduplicate search results, the manual team removed one unique study and missed zero duplicates versus the automation team who removed two unique studies and missed seven duplicates. Error rates were similar for the two remaining compared tasks—regarding Task 7: Screen the titles and abstracts and Task 9: Screen the full text, zero relevant studies were excluded by both teams. One task could not be compared between groups—Task 8: Find the full text. Conclusions For the majority of SR tasks where an SRA tool was used, the time required to complete that task was reduced for novice researchers while methodological quality was maintained.
… screen medical studies. … automated method to classify relevant articles for inclusion or exclusion during the abstract triage stage for creating and updating systematic reviews of medical …
Background The exponential growth of the biomedical literature necessitates investigating strategies to reduce systematic reviewer burden while maintaining the high standards of systematic review validity and comprehensiveness. Methods We compared the traditional systematic review screening process with (1) a review-of-reviews (ROR) screening approach and (2) a semi-automation screening approach using two publicly available tools (RobotAnalyst and AbstrackR) and different types of training sets (randomly selected citations subjected to dual-review at the title-abstract stage, highly curated citations dually reviewed at the full-text stage, and a combination of the two). We evaluated performance measures of sensitivity, specificity, missed citations, and workload burden Results The ROR approach for treatments of early-stage prostate cancer had a poor sensitivity (0.54) and studies missed by the ROR approach tended to be of head-to-head comparisons of active treatments, observational studies, and outcomes of physical harms and quality of life. Title and abstract screening incorporating semi-automation only resulted in a sensitivity of 100% at high levels of reviewer burden (review of 99% of citations). A highly curated, smaller-sized, training set ( n = 125) performed similarly to a larger training set of random citations ( n = 938). Conclusion Two approaches to rapidly update SRs—review-of-reviews and semi-automation—failed to demonstrate reduced workload burden while maintaining an acceptable level of sensitivity. We suggest careful evaluation of the ROR approach through comparison of inclusion criteria and targeted searches to fill evidence gaps as well as further research of semi-automation use, including more study of highly curated training sets.
… Automated data collection has … an automated systematic review for medical practices. This study will look at how automating healthcare systematic reviews are perceived by the medical …
Abstract The daily operation of clinical laboratories will be drastically impacted by two disruptive technologies: automation and artificial intelligence (the development and use of computer systems able to perform tasks that normally require human intelligence). These technologies will also expand the scope of laboratory medicine. Automation will result in increased efficiency but will require changes to laboratory infrastructure and a shift in workforce training requirements. The application of artificial intelligence to large clinical datasets generated through increased automation will lead to the development of new diagnostic and prognostic models. Together, automation and artificial intelligence will support the move to personalized medicine. Changes in pathology and clinical doctoral scientist training will be necessary to fully participate in these changes. KEYWORDS: Automation; artificial intelligence; deep learning; laboratory medicine
Machine learning for healthcare researchers face challenges to progress and reproducibility due to a lack of standardized processing frameworks for public datasets. We present MIMIC-Extract, an open source pipeline for transforming the raw electronic health record (EHR) data of critical care patients from the publicly-available MIMIC-III database into data structures that are directly usable in common time-series prediction pipelines. MIMIC-Extract addresses three challenges in making complex EHR data accessible to the broader machine learning community. First, MIMIC-Extract transforms raw vital sign and laboratory measurements into usable hourly time series, performing essential steps such as unit conversion, outlier handling, and aggregation of semantically similar features to reduce missingness and improve robustness. Second, MIMIC-Extract extracts and makes prediction of clinically-relevant targets possible, including outcomes such as mortality and length-of-stay as well as comprehensive hourly intervention signals for ventilators, vasopressors, and fluid therapies. Finally, the pipeline emphasizes reproducibility and extensibility to future research questions. We demonstrate the pipeline's effectiveness by developing several benchmark tasks for outcome and intervention forecasting and assessing the performance of competitive models.
The accurate extraction of surgical data from electronic health records (EHRs), particularly operative notes through manual chart review (MCR), is complex, crucial, and time-intensive, limited by human error due to fatigue and the level of training. This study aimed to develop and validate a novel Natural Language Processing (NLP) algorithm integrated with a Large Language Model (LLM; GPT4-Turbo) to automate the extraction of spinal surgery data from EHRs. The algorithm employed a two-stage approach. Initially, a rule-based NLP framework reviewed and classified candidate segments from the text, preserving their reference segments. These segments were then verified in the second stage through the LLM. The primary outcomes of this study were the accurate extraction of surgical data, including the type of surgery, levels operated, number of disks removed, and presence of intraoperative incidental durotomies. Secondary objectives explored time efficiency, tokenization lengths, and costs. The performance of the algorithm was assessed across two validation databases, analyzing metrics such as accuracy, sensitivity, discrimination, F1-score, and precision, with 95% confidence intervals calculated using percentile-based bootstrapping. The NLP + LLM algorithm markedly outperformed all performance metrics, demonstrating significant improvements in time and cost efficiency. These results suggest the potential for widespread adoption of this technology.
… We present two applications designed to select data from medical documentation in … data. The evaluation of the implemented procedures shows their usability for clinical data extraction …
PURPOSE A substantial portion of medical data is unstructured. Extracting data from unstructured text presents a barrier to advancing clinical research and improving patient care. In addition, ongoing studies have been focused predominately on the English language, whereas inflected languages with non-Latin alphabets (such as Slavic languages with a Cyrillic alphabet) present numerous linguistic challenges. We developed deep-learning–based natural language processing algorithms for automatically extracting biomarker status of patients with breast cancer from three oncology centers in Bulgaria. METHODS We used dual embeddings for English and Bulgarian languages, encoding both syntactic and polarity information for the words. The embeddings were subsequently aligned so that they were in the same vector space. The embeddings were used as input to convolutional or recurrent neural networks to derive the biomarker status of estrogen receptor, progesterone receptor, and human epidermal growth factor receptor 2. RESULTS We showed that we can resolve ambiguity in highly variable medical text containing both Latin and Cyrillic text. Final models incorporating both English and Bulgarian syntax and polarity embeddings achieved F1 scores of 0.90 or higher for all estrogen receptor, progesterone receptor, and human epidermal growth factor receptor 2 biomarkers. The models were robust against human errors originally found in the training set. In addition, such models can be extended for analyzing text containing words not seen during training. CONCLUSION By using several techniques that incorporate dual-word embeddings encoding syntactic and polarity information in two languages followed by deep neural network architectures, we show that researchers can extract and normalize parameters within medical data. The principles described here can be used to analyze Cyrillic or Latin mixed medical text and extract other parameters.
BackgroundElectronic health records (EHRs) contain detailed clinical data stored in proprietary formats with non-standard codes and structures. Participating in multi-site clinical research networks requires EHR data to be restructured and transformed into a common format and standard terminologies, and optimally linked to other data sources. The expertise and scalable solutions needed to transform data to conform to network requirements are beyond the scope of many health care organizations and there is a need for practical tools that lower the barriers of data contribution to clinical research networks.MethodsWe designed and implemented a health data transformation and loading approach, which we refer to as Dynamic ETL (Extraction, Transformation and Loading) (D-ETL), that automates part of the process through use of scalable, reusable and customizable code, while retaining manual aspects of the process that requires knowledge of complex coding syntax. This approach provides the flexibility required for the ETL of heterogeneous data, variations in semantic expertise, and transparency of transformation logic that are essential to implement ETL conventions across clinical research sharing networks. Processing workflows are directed by the ETL specifications guideline, developed by ETL designers with extensive knowledge of the structure and semantics of health data (i.e., “health data domain experts”) and target common data model.ResultsD-ETL was implemented to perform ETL operations that load data from various sources with different database schema structures into the Observational Medical Outcome Partnership (OMOP) common data model. The results showed that ETL rule composition methods and the D-ETL engine offer a scalable solution for health data transformation via automatic query generation to harmonize source datasets.ConclusionsD-ETL supports a flexible and transparent process to transform and load health data into a target data model. This approach offers a solution that lowers technical barriers that prevent data partners from participating in research data networks, and therefore, promotes the advancement of comparative effectiveness research using secondary electronic health data.
The use of primary care electronic health records for research is abundant. The benefits gained from utilising such records lies in their size, longitudinal data collection and data quality. However, the use of such data to undertake high quality epidemiological studies, can lead to significant challenges particularly in dealing with misclassification, variation in coding and the significant effort required to pre-process the data in a meaningful format for statistical analysis. In this paper, we describe a methodology to aid with the extraction and processing of such databases, delivered by a novel software programme; the “Data extraction for epidemiological research” (DExtER). The basis of DExtER relies on principles of extract, transform and load processes. The tool initially provides the ability for the healthcare dataset to be extracted, then transformed in a format whereby data is normalised, converted and reformatted. DExtER has a user interface designed to obtain data extracts specific to each research question and observational study design. There are facilities to input the requirements for; eligible study period, definition of exposed and unexposed groups, outcome measures and important baseline covariates. To date the tool has been utilised and validated in a multitude of settings. There have been over 35 peer-reviewed publications using the tool, and DExtER has been implemented as a validated public health surveillance tool for obtaining accurate statistics on epidemiology of key morbidities. Future direction of this work will be the application of the framework to linked as well as international datasets and the development of standardised methods for conducting electronic pre-processing and extraction from datasets for research purposes.
Despite significant strides in big data technology, extracting information from unstructured clinical data remains a formidable challenge. This study investigated the utility of large language models (LLMs) for extracting clinical data from unstructured radiological reports without additional training. In this retrospective study, 1800 radiologic reports, 600 from each of the three university hospitals, were collected, with seven pulmonary outcomes defined. Three pulmonology-trained specialists discerned the presence or absence of diseases. Data extraction from the reports was executed using Google Gemini Pro 1.0, OpenAI’s GPT-3.5, and GPT-4. The gold standard was predicated on agreement between at least two pulmonologists. This study evaluated the performance of the three LLMs in diagnosing seven pulmonary diseases (active tuberculosis, emphysema, interstitial lung disease, lung cancer, pleural effusion, pneumonia, and pulmonary edema) utilizing chest radiography and computed tomography scans. All models exhibited high accuracy (0.85–1.00) for most conditions. GPT-4 consistently outperformed its counterparts, demonstrating a sensitivity of 0.71–1.00; specificity of 0.89–1.00; and accuracy of 0.89 and 0.99 across both modalities, thus underscoring its superior capability in interpreting radiological reports. Notably, the accuracy of pleural effusion and emphysema on chest radiographs and pulmonary edema on chest computed tomography scans reached 0.99. The proficiency of LLMs, particularly GPT-4, in accurately classifying unstructured radiological data hints at their potential as alternatives to the traditional manual chart reviews conducted by clinicians.
… , automated extraction of data from … into clinical practice, with evidence that clinical workflow is not affected negatively. Additionally, the resultant benefits in readily available clinical data …
Summary Background: Clinical Data Warehouses (CDW) reuse Electronic health records (EHR) to make their data retrievable for research purposes or patient recruitment for clinical trials. However, much information are hidden in unstructured data like discharge letters. They can be preprocessed and converted to structured data via information extraction (IE), which is unfortunately a laborious task and therefore usually not available for most of the text data in CDW. Objectives: The goal of our work is to provide an ad hoc IE service that allows users to query text data ad hoc in a manner similar to querying structured data in a CDW. While search engines just return text snippets, our systems also returns frequencies (e.g. how many patients exist with “heart failure” including textual synonyms or how many patients have an LVEF < 45) based on the content of discharge letters or textual reports for special investigations like heart echo. Three subtasks are addressed: (1) To recognize and to exclude negations and their scopes, (2) to extract concepts, i.e. Boolean values and (3) to extract numerical values. Methods: We implemented an extended version of the NegEx-algorithm for German texts that detects negations and determines their scope. Furthermore, our document oriented CDW PaDaWaN was extended with query functions, e.g. context sensitive queries and regex queries, and an extraction mode for computing the frequencies for Boolean and numerical values. Results: Evaluations in chest X-ray reports and in discharge letters showed high F1-scores for the three subtasks: Detection of negated concepts in chest X-ray reports with an F1-score of 0.99 and in discharge letters with 0.97; of Boolean values in chest X-ray reports about 0.99, and of numerical values in chest X-ray reports and discharge letters also around 0.99 with the exception of the concept age. Discussion: The advantages of an ad hoc IE over a standard IE are the low development effort (just entering the concept with its variants), the promptness of the results and the adaptability by the user to his or her particular question. Disadvantage are usually lower accuracy and confidence. This ad hoc information extraction approach is novel and exceeds existing systems: Roogle [1] extracts predefined concepts from texts at preprocessing and makes them retrievable at runtime. Dr. Warehouse [2] applies negation detection and indexes the produced subtexts which include affirmed findings. Our approach combines negation detection and the extraction of concepts. But the extraction does not take place during preprocessing, but at runtime. That provides an ad hoc, dynamic, interactive and adjustable information extraction of random concepts and even their values on the fly at runtime. Conclusions: We developed an ad hoc information extraction query feature for Boolean and numerical values within a CDW with high recall and precision based on a pipeline that detects and removes negations and their scope in clinical texts.
Breast cancer is the leading cause of cancer mortality in women between the ages of 15 and 54. During mammography screening, radiologists use a strict lexicon (BI-RADS) to describe and report their findings. Mammography records are then stored in a well-defined database format (NMD). Lately, researchers have applied data mining and machine learning techniques to these databases. They successfully built breast cancer classifiers that can help in early detection of malignancy. However, the validity of these models depends on the quality of the underlying databases. Unfortunately, most databases suffer from inconsistencies, missing data, inter-observer variability and inappropriate term usage. In addition, many databases are not compliant with the NMD format and/or solely consist of text reports. BI-RADS feature extraction from free text and consistency checks between recorded predictive variables and text reports are crucial to addressing this problem. We describe a general scheme for concept information retrieval from free text given a lexicon, and present a BI-RADS features extraction algorithm for clinical data mining. It consists of a syntax analyzer, a concept finder and a negation detector. The syntax analyzer preprocesses the input into individual sentences. The concept finder uses a semantic grammar based on the BI-RADS lexicon and the experts’ input. It parses sentences detecting BI-RADS concepts. Once a concept is located, a lexical scanner checks for negation. Our method can handle multiple latent concepts within the text, filtering out ultrasound concepts. On our dataset, our algorithm achieves 97.7% precision, 95.5% recall and an F1-score of 0.97. It outperforms manual feature extraction at the 5% statistical significance level.
Background Electronic medical records (EMRs) are adopted for storing patient-related healthcare information. Using data mining techniques, it is possible to make use of and derive benefit from this massive amount of data effectively. We aimed to evaluate validity of data extracted by the Customized eXtraction Program (CXP). Methods The CXP extracts and structures data in rapid standardised processes. The CXP was programmed to extract TNFα-native active ulcerative colitis (UC) patients from EMRs using defined International Classification of Disease-10 (ICD-10) codes. Extracted data were read in parallel with manual assessment of the EMR to compare with CXP-extracted data. Results From the complete EMR set, 2,802 patients with code K51 (UC) were extracted. Then, CXP extracted 332 patients according to inclusion and exclusion criteria. Of these, 97.5% were correctly identified, resulting in a final set of 320 cases eligible for the study. When comparing CXP-extracted data against manually assessed EMRs, the recovery rate was 95.6–101.1% over the years with 96.1% weighted average sensitivity. Conclusion Utilisation of the CXP software can be considered as an effective way to extract relevant EMR data without significant errors. Hence, by extracting from EMRs, CXP accurately identifies patients and has the capacity to facilitate research studies and clinical trials by finding patients with the requested code as well as funnel down itemised individuals according to specified inclusion and exclusion criteria. Beyond this, medical procedures and laboratory data can rapidly be retrieved from the EMRs to create tailored databases of extracted material for immediate use in clinical trials.
BackgroundTranslational research typically requires data abstracted from medical records as well as data collected specifically for research. Unfortunately, many data within electronic health records are represented as text that is not amenable to aggregation for analyses. We present a scalable open source SQL Server Integration Services package, called Regextractor, for including regular expression parsers into a classic extract, transform, and load workflow. We have used Regextractor to abstract discrete data from textual reports from a number of ‘machine generated’ sources. To validate this package, we created a pulmonary function test data mart and analyzed the quality of the data mart versus manual chart review.MethodsEleven variables from pulmonary function tests performed closest to the initial clinical evaluation date were studied for 100 randomly selected subjects with scleroderma. One research assistant manually reviewed, abstracted, and entered relevant data into a database. Correlation with data obtained from the automated pulmonary function test data mart within the Northwestern Medical Enterprise Data Warehouse was determined.ResultsThere was a near perfect (99.5%) agreement between results generated from the Regextractor package and those obtained via manual chart abstraction. The pulmonary function test data mart has been used subsequently to monitor disease progression of patients in the Northwestern Scleroderma Registry. In addition to the pulmonary function test example presented in this manuscript, the Regextractor package has been used to create cardiac catheterization and echocardiography data marts. The Regextractor package was released as open source software in October 2009 and has been downloaded 552 times as of 6/1/2012.ConclusionsCollaboration between clinical researchers and biomedical informatics experts enabled the development and validation of a tool (Regextractor) to parse, abstract and assemble structured data from text data contained in the electronic health record. Regextractor has been successfully used to create additional data marts in other medical domains and is available to the public.
… patient data for clinical decision support. In contrast, a clinical data warehouse focuses exclusively on data … The basic clinical warehouse operation involves identifying a set of patients …
合并后形成一条由“技术背景—病例规范—工作流基础设施—数据抽取与工程—循证检索—自动化分析—LLM科研应用—医学写作—治理与人员能力”组成的完整链路。分组将病例报告规范、临床数据准备、研究流程整合、证据综合、计算分析及大语言模型应用分别处理,避免把不同环节过度合并;同时以可解释性、伦理责任、人工复核和医生科研能力作为贯穿全流程的质量保障。