Classifying free-text triage chief complaints into syndromic categories with natural language processing

作者:

Highlights:

摘要

Objective: Develop and evaluate a natural language processing application for classifying chief complaints into syndromic categories for syndromic surveillance. Introduction: Much of the input data for artificial intelligence applications in the medical field are free-text patient medical records, including dictated medical reports and triage chief complaints. To be useful for automated systems, the free-text must be translated into encoded form. Methods: We implemented a biosurveillance detection system from Pennsylvania to monitor the 2002 Winter Olympic Games. Because input data was in free-text format, we used a natural language processing text classifier to automatically classify free-text triage chief complaints into syndromic categories used by the biosurveillance system. The classifier was trained on 4700 chief complaints from Pennsylvania. We evaluated the ability of the classifier to classify free-text chief complaints into syndromic categories with a test set of 800 chief complaints from Utah. Results: The classifier produced the following areas under the ROC curve: Constitutional = 0.95; Gastrointestinal = 0.97; Hemorrhagic = 0.99; Neurological = 0.96; Rash = 1.0; Respiratory = 0.99; Other = 0.96. Using information stored in the system’s semantic model, we extracted from the Respiratory classifications lower respiratory complaints and lower respiratory complaints with fever with a precision of 0.97 and 0.96, respectively. Conclusion: Results suggest that a trainable natural language processing text classifier can accurately extract data from free-text chief complaints for biosurveillance.

论文关键词:Natural language processing,Text classification,Syndromic surveillance,Chief complaints

论文评审过程:Received 22 January 2004, Revised 26 March 2004, Accepted 3 April 2004, Available online 20 June 2004.

论文官网地址:https://doi.org/10.1016/j.artmed.2004.04.001