The IRMA Community
Newsletters
Research IRM
Click a keyword to search titles using our InfoSci-OnDemand powered search:
|
Exploiting Captions for Web Data Mining
Abstract
We survey research on using captions in data mining from the Web. Captions are text that describes some other information (typically, multimedia). Since text is considerably easier to analyze than non-text, a good way to support access to non-text is to index the words of its captions. However, captions vary considerably in form and content on the Web. We discuss the range of syntactic clues (such as HTML tags) and semantic clues (such as particular words). We discuss how to quantify clue strength and combine clues for a consensus. We then discuss the problem of mapping information in captions to information in media objects. While it is hard, classes of mapping schemes are distinguishable, and a segmentation of the media can be matched to a parse of the caption.
Related Content
Md Sakir Ahmed, Abhijit Bora.
© 2024.
15 pages.
|
Lakshmi Haritha Medida, Kumar.
© 2024.
18 pages.
|
Gypsy Nandi, Yadika Prasad.
© 2024.
16 pages.
|
Saurav Bhattacharjee, Sabiha Raiyesha.
© 2024.
14 pages.
|
Naren Kathirvel, Kathirvel Ayyaswamy, B. Santhoshi.
© 2024.
26 pages.
|
K. Sudha, C. Balakrishnan, T. P. Anish, T. Nithya, B. Yamini, R. Siva Subramanian, M. Nalini.
© 2024.
25 pages.
|
Sabiha Raiyesha, Papul Changmai.
© 2024.
28 pages.
|
|
|