Medical record retrieval and extraction for professional information access
Journal
CEUR
Journal Volume
968
Pages
17-25
Date Issued
2013
Author(s)
Abstract
This paper analyzes the linguistic phenomena in medical records in different departments, including average record size, vocabulary, entropy of medical languages, grammaticality, and so on. Five retrieval models with six pre-processing strategies on different parts of medical records are explored on an NTUH medical record dataset. Both coarse-grained relevance evaluation on department level and fine-grained relevance evaluation on course and treatment level are conducted. Query accesses to the medical records in medical languages of smaller entropy tend to have better performance. The departments related to generic parts of body such as Departments of Internal Medicine and Surgery may confuse the retrieval, in particular, for Departments of Oncology and Neurology. Okapi model with stemming achieves the best performance on both department and course and treatment levels. Copyright ? by the paper's authors.
Type
conference paper
