TY - GEN
T1 - Any Information Is Just Worth One Single Screenshot
T2 - 63rd Annual Meeting of the Association for Computational Linguistics, ACL 2025
AU - Liu, Zheng
AU - Liu, Ze
AU - Liang, Zhengyang
AU - Zhou, Junjie
AU - Xiao, Shitao
AU - Gao, Chao
AU - Zhang, Chen Jason
AU - Lian, Defu
N1 - Publisher Copyright:
© 2025 Association for Computational Linguistics.
PY - 2025/7
Y1 - 2025/7
N2 - With the popularity of multimodal techniques, it receives growing interests to acquire useful information in visual forms. In this work, we formulate an emerging IR paradigm called Visualized Information Retrieval, or Vis-IR, where multimodal information, such as texts, images, tables and charts, is jointly represented by a unified visual format called Screenshots, for various retrieval applications. We further make three key contributions for Vis-IR. First, we create VIRA (Vis-IR Aggregation), a large-scale dataset comprising a vast collection of screenshots from diverse sources, carefully curated into captioned and question-answer formats. Second, we develop UniSE (Universal Screenshot Embeddings), a family of retrieval models that enable screenshots to query or be queried across arbitrary data modalities. Finally, we construct MVRB (Massive Visualized IR Benchmark), a comprehensive benchmark covering a variety of task forms and application scenarios. Through extensive evaluations on MVRB, we highlight the deficiency from existing multimodal retrievers and the substantial improvements made by UniSE. Our data, model and benchmark have been made publicly available, which lays a solid foundation for this emerging field.
AB - With the popularity of multimodal techniques, it receives growing interests to acquire useful information in visual forms. In this work, we formulate an emerging IR paradigm called Visualized Information Retrieval, or Vis-IR, where multimodal information, such as texts, images, tables and charts, is jointly represented by a unified visual format called Screenshots, for various retrieval applications. We further make three key contributions for Vis-IR. First, we create VIRA (Vis-IR Aggregation), a large-scale dataset comprising a vast collection of screenshots from diverse sources, carefully curated into captioned and question-answer formats. Second, we develop UniSE (Universal Screenshot Embeddings), a family of retrieval models that enable screenshots to query or be queried across arbitrary data modalities. Finally, we construct MVRB (Massive Visualized IR Benchmark), a comprehensive benchmark covering a variety of task forms and application scenarios. Through extensive evaluations on MVRB, we highlight the deficiency from existing multimodal retrievers and the substantial improvements made by UniSE. Our data, model and benchmark have been made publicly available, which lays a solid foundation for this emerging field.
UR - https://www.scopus.com/pages/publications/105021011943
U2 - 10.18653/v1/2025.acl-long.943
DO - 10.18653/v1/2025.acl-long.943
M3 - Conference article published in proceeding or book
AN - SCOPUS:105021011943
T3 - Proceedings of the Annual Meeting of the Association for Computational Linguistics
SP - 19238
EP - 19261
BT - Long Papers
A2 - Che, Wanxiang
A2 - Nabende, Joyce
A2 - Shutova, Ekaterina
A2 - Pilehvar, Mohammad Taher
PB - Association for Computational Linguistics (ACL)
Y2 - 27 July 2025 through 1 August 2025
ER -