VLDB 2026 Research / reviewers in the wild / expert
Rubèn Tito
dblp:257/5708 · also Rubèn Pérez Tito
· DBLP profile ↗
12ranked-venue papers
4as first author
10since 2021 · last 2024
0000-0002-5657-9790ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 4 first-author · 9 since 2021Databases, data management, data science and information retrieval · 6 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Multi-page Document Visual Question Answering Using Self-attention Scoring Mechanism
Lei Kang 0002, Rubèn Tito, Ernest Valveny, Dimosthenis Karatzas |
ICDAR (6) | 2 |
| 2024 | Privacy-Aware Document Visual Question Answering
Rubèn Tito, Marlon Tobaben, Raouf Kerkouche, Mohamed Ali Souibgui, Kangsoo Jung, Joonas Jälkö, Vincent Poulain D'Andecy, Aurélie Joseph, Lei Kang 0002, Ernest Valveny, Antti Honkela, Mario Fritz, Dimosthenis Karatzas |
ICDAR (6) | 1 |
| 2023 | Document Understanding Dataset and Evaluation (DUDE)abstractWe call on the Document AI (DocAI) community to reevaluate current methodologies and embrace the challenge of creating more practically-oriented benchmarks. Document Understanding Dataset and Evaluation (DUDE) seeks to remediate the halted research progress in understanding visually-rich documents (VRDs). We present a new dataset1with novelties related to types of questions, answers, and document layouts based on multi-industry, multi-domain, and multi-page VRDs of various origins, and dates. Moreover, we are pushing the boundaries of current methods by creating multi-task and multi-domain evaluation setups that more accurately simulate real-world situations where powerful generalization and adaptation under low-resource settings are desired. DUDE aims to set a new standard as a more practical, long-standing benchmark for the community, and we hope that it will lead to future extensions and contributions that address real-world challenges. Finally, our work illustrates the importance of finding more efficient ways to model language, images, and layout in DocAI. Jordy Van Landeghem, Rafal Powalski, Rubèn Tito, Dawid Jurkiewicz, Matthew B. Blaschko, Lukasz Borchmann, Mickaël Coustaty, Marie-Francine Moens, Michal Pietruszka, Bertrand Anckaert, Tomasz Stanislawek, Pawel Józiak, Ernest Valveny |
ICCV | 3 |
| 2023 | ICDAR 2023 Competition on Document UnderstanDing of Everything (DUDE)
Jordy Van Landeghem, Rubèn Tito, Lukasz Borchmann, Michal Pietruszka, Dawid Jurkiewicz, Rafal Powalski, Pawel Józiak, Sanket Biswas, Mickaël Coustaty, Tomasz Stanislawek |
ICDAR (2) | 2 |
| 2023 | Hierarchical multimodal transformers for Multipage DocVQA
Rubèn Tito, Dimosthenis Karatzas, Ernest Valveny |
Pattern Recognit. | 1 |
| 2022 | InfographicVQAabstractInfographics communicate information using a combination of textual, graphical and visual elements. This work explores the automatic understanding of infographic images by using a Visual Question Answering technique. To this end, we present InfographicVQA, a new dataset comprising a diverse collection of infographics and question-answer annotations. The questions require methods that jointly reason over the document layout, textual content, graphical elements, and data visualizations. We curate the dataset with an emphasis on questions that require elementary reasoning and basic arithmetic skills. For VQA on the dataset, we evaluate two Transformer-based strong baselines. Both the baselines yield unsatisfactory results compared to near perfect human performance on the dataset. The results suggest that VQA on infographics—images that are designed to communicate information quickly and clearly to human brain—is ideal for benchmarking machine understanding of complex document images. The dataset is available for download at docvqa.org Minesh Mathew, Viraj Bagal, Rubèn Tito, Dimosthenis Karatzas, Ernest Valveny, C. V. Jawahar |
WACV | 3 |
| 2021 | Document Collection Visual Question Answering
Rubèn Tito, Dimosthenis Karatzas, Ernest Valveny |
ICDAR (2) | 1 |
| 2021 | ICDAR 2021 Competition on Document Visual Question Answering
Rubèn Tito, Minesh Mathew, C. V. Jawahar, Ernest Valveny, Dimosthenis Karatzas |
ICDAR (4) | 1 |
| 2021 | Real-time Lexicon-free Scene Text Retrieval
Andrés Mafla, Rubèn Tito, Sounak Dey, Lluís Gómez i Bigorda, Marçal Rusiñol, Ernest Valveny, Dimosthenis Karatzas |
Pattern Recognit. | 2 |
| 2021 | Multimodal grid features and cell pointers for scene text visual question answering
Lluís Gómez i Bigorda, Ali Furkan Biten, Rubèn Tito, Andrés Mafla, Marçal Rusiñol, Ernest Valveny, Dimosthenis Karatzas |
Pattern Recognit. Lett. | 3 |
| 2019 | Scene Text Visual Question AnsweringabstractCurrent visual question answering datasets do not consider the rich semantic information conveyed by text within an image. In this work, we present a new dataset, ST-VQA, that aims to highlight the importance of exploiting high-level semantic information present in images as textual cues in the Visual Question Answering process. We use this dataset to define a series of tasks of increasing difficulty for which reading the scene text in the context provided by the visual information is necessary to reason and generate an appropriate answer. We propose a new evaluation metric for these tasks to account both for reasoning errors as well as shortcomings of the text recognition module. In addition we put forward a series of baseline methods, which provide further insight to the newly released dataset, and set the scene for further research. Ali Furkan Biten, Rubèn Tito, Andrés Mafla, Lluís Gómez i Bigorda, Marçal Rusiñol, C. V. Jawahar, Ernest Valveny, Dimosthenis Karatzas |
ICCV | 2 |
| 2019 | ICDAR 2019 Competition on Scene Text Visual Question AnsweringabstractThis paper presents final results of ICDAR 2019 Scene Text Visual Question Answering competition (ST-VQA). ST-VQA introduces an important aspect that is not addressed by any Visual Question Answering system up to date, namely the incorporation of scene text to answer questions asked about an image. The competition introduces a new dataset comprising 23,038 images annotated with 31,791 question / answer pairs where the answer is always grounded on text instances present in the image. The images are taken from 7 different public computer vision datasets, covering a wide range of scenarios. The competition was structured in three tasks of increasing difficulty, that require reading the text in a scene and understanding it in the context of the scene, to correctly answer a given question. A novel evaluation metric is presented, which elegantly assesses both key capabilities expected from an optimal model: text recognition and image understanding. A detailed analysis of results from different participants is showcased, which provides insight into the current capabilities of VQA systems that can read. We firmly believe the dataset proposed in this challenge will be an important milestone to consider towards a path of more robust and general models that can exploit scene text to achieve holistic image understanding. Ali Furkan Biten, Rubèn Tito, Andrés Mafla, Lluís Gómez i Bigorda, Marçal Rusiñol, Minesh Mathew, C. V. Jawahar, Ernest Valveny, Dimosthenis Karatzas |
ICDAR | 2 |