An Efficient Filtration Method in Biological Sequence Databases
Date Issued
2004
Date
2004
Author(s)
Lo, Wen-Hsing
DOI
en-US
Abstract
Sequence comparison is one of the most important primitive operations in bioinformatics. Roughly speaking, this operation finds which parts of sequences are alike and which parts are different. As the size of a sequence database scales to millions of base pairs, it becomes impractical to search the whole database with sequence alignment methods based on the dynamic programming approach which yields quadratic time complexity. Filtration methods are thus proposed in order to screen out most unrelated data sequences in the preprocessing stage. However, existing filtration methods either incurs false negatives or retains too many candidates.
In this thesis, we proposed a filtration method called Transformation-based Database Filtration method (TDF) which consists of two phases. First, we divide the data sequences into several blocks, each of which is transformed into a feature vector by Haar wavelet transform. Then, we build an index for them. In the second phase, we search the index and extract those candidate blocks whose distance to the feature vector of the query sequence is less than a predefined threshold. Finally, for each candidate block, we calculate the edit distance between the corresponding data sequence and the query sequence. Experimental results show that our method prunes a large portion of the database and guarantees no false negative.
Subjects
序列比對
查詢篩選
生物序列資料庫
sequence comparison
filtration method
biological sequence database
Type
other
File(s)![Thumbnail Image]()
Loading...
Name
ntu-93-R91725034-1.pdf
Size
23.31 KB
Format
Adobe PDF
Checksum
(MD5):d361d520233e9e05719278ca32e10dd2
