Efficient Data Mining for Path Traversal Patterns

MING-SYAN CHEN; Park, Jong Soo; Yu, P.S.; Chen, Ming-Syan; Park, Jong Soo; Yu, P.S.

doi:10.1109/69.683753

Efficient Data Mining for Path Traversal Patterns

Journal

IEEE Transactions on Knowledge and Data Engineering

Journal Volume

10

Journal Issue

2

Pages

209-221

Date Issued

1998

Date

1998

Author(s)

MING-SYAN CHEN

Park, Jong Soo

Yu, P.S.

DOI

10.1109/69.683753

URI

http://ntur.lib.ntu.edu.tw//handle/246246/141942

http://ntur.lib.ntu.edu.tw/bitstream/246246/141942/1/15.pdf

https://www.scopus.com/inward/record.uri?eid=2-s2.0-0032028932&doi=10.1109%2f69.683753&partnerID=40&md5=87af610575e1b92ecdb605ed38558904

Abstract

In this paper, we explore a new data mining capability that involves mining path traversal patterns in a distributed information-providing environment where documents or objects are linked together to facilitate interactive access. Our solution procedure consists of two steps. First, we derive an algorithm to convert the original sequence of log data into a set of maximal forward references. By doing so, we can filter out the effect of some backward references, which are mainly made for ease of traveling and concentrate on mining meaningful user access sequences. Second, we derive algorithms to determine the frequent traversal patterns - i.e., large reference sequences - from the maximal forward references obtained. Two algorithms are devised for determining large reference sequences; one is based on some hashing and pruning techniques, and the other is further improved with the option of determining large reference sequences in batch so as to reduce the number of database scans required. Performance of these two methods is comparatively analyzed. It is shown that the option of selective scan is very advantageous and can lead to prominent performance improvement. Sensitivity analysis on various parameters is conducted. © 1998 IEEE.

Subjects

Data mining; Distributed information system; Performance analysis; Traversal patterns; World Wide Web

Other Subjects

Algorithms; Database systems; Distributed computer systems; Interactive computer systems; Online systems; Sensitivity analysis; Wide area networks; Data mining; Distributed information system; Traversal patterns; World wide web; Data acquisition

Type

journal article

File(s)

Name

15.pdf

Size

426.7 KB

Format

Adobe PDF

Checksum

(MD5):b2ca5443a53ddd7fdb5a47d77671be3b

Efficient Data Mining for Path Traversal Patterns

關於 (About)

聯絡資訊 (Contact Us)

相關網站 (Useful Links)

關於開放取用 (Open Access, OA)

出版社期刊論文授權政策 (Copyright)

使用說明 (Instructions)

登入說明 (Sign-in)

匯入著作 (Submission)