Exploiting Captions for Web Data Mining

View Sample PDF

Author(s): Neil C. Rowe (U.S. Naval Postgraduate School, USA)
Copyright: 2008
Pages: 25
Source title: Data Warehousing and Mining: Concepts, Methodologies, Tools, and Applications
Source Author(s)/Editor(s): John Wang (Montclair State University, USA)
DOI: 10.4018/978-1-59904-951-9.ch084

Keywords: Data Mining and Databases / Data Warehousing / Information Science Reference / Library & Information Science

Purchase

View Exploiting Captions for Web Data Mining on the publisher's website for pricing and purchasing information.

Abstract

We survey research on using captions in data mining from the Web. Captions are text that describes some other information (typically, multimedia). Since text is considerably easier to analyze than non-text, a good way to support access to non-text is to index the words of its captions. However, captions vary considerably in form and content on the Web. We discuss the range of syntactic clues (such as HTML tags) and semantic clues (such as particular words). We discuss how to quantify clue strength and combine clues for a consensus. We then discuss the problem of mapping information in captions to information in media objects. While it is hard, classes of mapping schemes are distinguishable, and a segmentation of the media can be matched to a parse of the caption.

The IRMA Community

Research IRM

Exploiting Captions for Web Data Mining

Purchase

Abstract

Related Content

IRMA Sponsors