https://scholars.lib.ntu.edu.tw/handle/123456789/309065
標題: | DOMISA: DOM-based information space adsorption for web information hierarchy mining | 作者: | Kao, H.-Y. Ho, J.-M. MING-SYAN CHEN |
公開日期: | 2004 | 起(迄)頁: | 312-320 | 來源出版物: | SIAM Proceedings Series | 摘要: | Due to the growth of dynamic page generation techniques, the amount and the complexity of Web pages has been increasing explosively, as has the information contained within Web pages. Redundant and irrelevant information is distributed and mixed throughout a page, making it difficult to automatically identify the useful information in that page. Consequently, we propose an information hierarchy in this paper, and, from that hierarchy, we can extract the significance and the relationship value of information contained within a Web page. We can then use this hierarchical structure to create a new browsing process. Our DOM-based Information Space Adsorption (DOMISA) system applies information theory to map information in a page into an information space, and our gradient tree adsorption (GTA) process uses the document object model (DOM) trees of pages to build information hierarchies. Experiments on several commercial news Web sites show high precision and recall rates achieved by DOMISA in determining information clusters of pages which validates its practical applicability to Web sites. |
URI: | http://www.scopus.com/inward/record.url?eid=2-s2.0-2942525697&partnerID=MN8TOARS http://scholars.lib.ntu.edu.tw/handle/123456789/309065 |
DOI: | 10.1137/1.9781611972740.29 |
顯示於: | 電機工程學系 |
在 IR 系統中的文件,除了特別指名其著作權條款之外,均受到著作權保護,並且保留所有的權利。