Exploiting Semantic Concepts for Keyword-based Photo and Video Search
Date Issued
2009
Date
2009
Author(s)
Wu, Po-Tun
Abstract
There appears explosive growth of photos and videos due to the proliferation of capture devices and numerous easy-to-use image and video sharing services such as Flickr, MySpace, YouTube, etc. How to effectively index and retrieve these large databases still remains an open problem. Traditional content-based and keyword-based multimedia retrieval methods often fail to meet users’ expectation due to the semantic gap. In response to the strong demands of semantic search over large-scale consumer photos, which generally lack reliable user-provided annotations, we investigate the feasibility and challenges entailed by the new multimedia search paradigm – “concept search,” namely, retrieving visual objects by large-scale automatic concept detectors with keywords.hough concept search is promising, several important issues must be addressed. We investigate the problem in several folds: (1) the effective query-to-concept mapping and concept selection methods over large-scale concept ontology; (2) a comprehensive performance study of the fundamental factors of keyword-based concept search, including concept selection strategy, lexicon size, detector accuracy, to name a few; (3) the quality and feasibility of the pre-trained concept detectors applying to cross-domain consumer-generated data (i.e., Flickr photos); (4) the search quality by fusing automatic concepts and user-generated data (tags); (5) the demand of efficient (query-time) indexing techniques over large-scale multimedia instead of off-line indexing (where query information is ignored); (6) the study of leveraging both semantic meaning and visual co-occurrence to fulfill user’s information needs.xperimenting over two large-scale benchmarks, TRECVID (broadcast news videos) and Flickr550 (consumer photos), we have confirmed the effectiveness of concept search via the semantic mapping by Google-based semantic expansion methods and demonstrated its superiority over conventional WordNet-like methods both in effectiveness and efficiency. Most of the parameterized factors in concept search are investigated – leading to the conclusion that concept search is indeed more effective than text-based search. We further illustrate the potential of pre-trained concept detectors applying on cross-domain consumer photos. We point that the user-contributed tags are somehow inaccurate or ambiguous and can be improved by semantic concepts in the applications of keyword-based search. Auxiliary visual information mining from large-scale image database is utilized to improve effectiveness of user queries. Most of all, the proposed novel query-time indexing method, FRANK-TAAT, not only reduces indexing overhead for large-scale database but also solves commonly observed “low recall” problem, which is seldom addressed in the prior work.
Subjects
Content-based Image Retrieval
Semantic Concept Search
Large-scale Multimedia Indexing and Retrieval
Type
thesis
File(s)![Thumbnail Image]()
Loading...
Name
ntu-98-R96922091-1.pdf
Size
23.32 KB
Format
Adobe PDF
Checksum
(MD5):5450d76366d106b5e392fb07cf39215e
