Special issue: Text mining and information analysis; Retrieving and clustering keywords in neurosurgery operation reports using text mining techniques
Journal
ACM International Conference Proceeding Series
Pages
88-100
Date Issued
2018
Author(s)
Abstract
Background: To develop a more practical and reasonable classification of surgical procedures, we applied text mining techniques to retrieve and categorize keywords in operation reports. Materials and Methods: Based on neurosurgical operation reports performed in a Taiwan medical center between 2009 and 2012, a corpus containing 3,657 documents was built. A total of 9,906 words were extracted. Initially, we applied term frequency-inverse document frequency (TF-IDF) weighting to automatically select pertinent keywords but the results were unsatisfactory. Then, we manually chose 45 keywords that belong to 3 categories: brain, spine and others. All documents were checked in an automated fashion for the presence of these keywords, producing a binary data matrix, which was used to compute the cosine similarity matrix. Then, we applied 6 variants of agglomerative clustering to build the dendrograms. Results: The document frequencies (DFs) of these 45 keywords ranged from 12 to 1,250, with an average of 444±342. The number of distinctive keywords per document ranged from 0-15, with an average of 5.5±2.5. The similarities between DF vectors are higher between keywords in the same category (brain or spine). The shortest link method and the unweighted pair-group method using the centroid (UPGMC) methods performed best on external and internal evaluation, respectively. Conclusion: The distributions of important keywords in neurosurgery operation reports reveal the localized nature of surgical procedures.
SDGs
Publisher
Association for Computing Machinery
Type
conference paper
