Developing a Dataset for Evaluating Approaches for Document Expansion with Images

Debasis Ganguly, Iacer Calixto, Gareth Jones


Abstract
Motivated by the adage that a “picture is worth a thousand words” it can be reasoned that automatically enriching the textual content of a document with relevant images can increase the readability of a document. Moreover, features extracted from the additional image data inserted into the textual content of a document may, in principle, be also be used by a retrieval engine to better match the topic of a document with that of a given query. In this paper, we describe our approach of building a ground truth dataset to enable further research into automatic addition of relevant images to text documents. The dataset is comprised of the official ImageCLEF 2010 collection (a collection of images with textual metadata) to serve as the images available for automatic enrichment of text, a set of 25 benchmark documents that are to be enriched, which in this case are children’s short stories, and a set of manually judged relevant images for each query story obtained by the standard procedure of depth pooling. We use this benchmark dataset to evaluate the effectiveness of standard information retrieval methods as simple baselines for this task. The results indicate that using the whole story as a weighted query, where the weight of each query term is its tf-idf value, achieves an precision of 0:1714 within the top 5 retrieved images on an average.
Anthology ID:
L16-1299
Volume:
Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC'16)
Month:
May
Year:
2016
Address:
Portorož, Slovenia
Editors:
Nicoletta Calzolari, Khalid Choukri, Thierry Declerck, Sara Goggi, Marko Grobelnik, Bente Maegaard, Joseph Mariani, Helene Mazo, Asuncion Moreno, Jan Odijk, Stelios Piperidis
Venue:
LREC
SIG:
Publisher:
European Language Resources Association (ELRA)
Note:
Pages:
1898–1901
Language:
URL:
https://preview.aclanthology.org/build-pipeline-with-new-library/L16-1299/
DOI:
Bibkey:
Cite (ACL):
Debasis Ganguly, Iacer Calixto, and Gareth Jones. 2016. Developing a Dataset for Evaluating Approaches for Document Expansion with Images. In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC'16), pages 1898–1901, Portorož, Slovenia. European Language Resources Association (ELRA).
Cite (Informal):
Developing a Dataset for Evaluating Approaches for Document Expansion with Images (Ganguly et al., LREC 2016)
Copy Citation:
PDF:
https://preview.aclanthology.org/build-pipeline-with-new-library/L16-1299.pdf