VLDB 2026 Research / reviewers in the wild / expert
Lluís Gómez i Bigorda
dblp:136/4068 · also Lluis Gomez i Bigorda, Lluis Gomez-Bigorda, Lluís Gómez
· DBLP profile ↗
54ranked-venue papers
14as first author
24since 2021 · last 2026
0000-0003-1408-9803ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 35 · 11 first-author · 13 since 2021Databases, data management, data science and information retrieval · 26 · 6 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 4 first-author · 10 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Rethinking Membership Inference Attacks for CLIPabstractMembership Inference Attacks (MIAs) test whether a model has memorized training data, and are a key tool for auditing privacy risks in machine learning. Recent papers report near-perfect MIA success against large vision-language models such as CLIP, but almost all evaluations train on one web-scale corpus (e.g. LAION-400M) and treat samples from a different corpus (e.g. COCO or CC12M) as non-members - thereby turning the task into out-of-distribution (OOD) detection rather than true membership testing, introducing spurious signals unrelated to true memorization. We revisit the problem with a distribution-matched benchmark built from the CommonPool-L corpus of DataComp. A ViT-B/16 CLIP trained on 400M pairs is accompanied by two 26-shard, i.i.d. splits that serve as member and non-member sets, sharing the exact same acquisition and preprocessing pipeline. Under this strictly in-distribution setting, every published MIA baseline collapses to chance (~51% AUC). To explain this collapse, we derive a scaling-law upper bound for similarity-based attacks showing that the expected member vs. non-member similarity gap decays as O(T/N) for contrastive learning with T epochs over N samples. Empirically, as we vary the training set size while holding all hyper-parameters fixed, the gap follows the predicted linear trend in log–log space, and Cosine Similarity Attack AUC drops from 94% to 51%. Finally, we propose a simple, white-box, gradient-based MIA that outperforms prior attacks for CLIP without relying on OOD cues. We release code, checkpoints, and data to foster comprehensive and reproducible privacy research on multimodal CLIP-like foundation models. Lluís Gómez i Bigorda |
AAAI | 1 |
| 2026 | Debiasing CLIP with Neural Interventions
Amelia Gómez Grabowska, Jordi Gonzàlez 0001, Lluís Gómez i Bigorda |
ECIR (3) | 3 |
| 2026 | Revisiting How We Access Historical Archives: Auditing Gender Stereotypes and the Division of Labour in the Analysis of Historical Photography Collections
Francesc Net, Adrià Molina, Sofia Llacer-Caro, Lluís Gómez i Bigorda |
ICDAR (3) | 4 |
| 2026 | Preserving privacy without compromising accuracy: Machine unlearning for handwritten text recognitionabstract• Encoder-only Transformer baseline for handwriting text recognition (HTR). • Added writer-style head tracks & controls user-identifiable memorization. • Neural-activation pruning removes writer clues while preserving HTR accuracy. • Writer-ID Confusion (WIC) enforces uniform IDs, boosting efficient machine unlearning. • Membership-inference tests confirm privacy gains without costly retraining. Handwritten Text Recognition (HTR) is crucial for document digitization, but handwritten data can contain user-identifiable features, like unique writing styles, posing privacy risks. Regulations such as the “right to be forgotten” require models to remove these sensitive traces without full retraining. We introduce a practical encoder-only transformer baseline as a robust reference for future HTR research. Building on this, we propose a two-stage unlearning framework for multihead transformer HTR models. Our method combines neural pruning with machine unlearning applied to a writer classification head, ensuring sensitive information is removed while preserving the recognition head. We also present Writer-ID Confusion (WIC), a method that forces the forget set to follow a uniform distribution over writer identities, unlearning user-specific cues while maintaining text recognition performance. We compare WIC to Random Labeling, Fisher Forgetting, Amnesiac Unlearning, and DELETE within our prune-unlearn pipeline and consistently achieve better privacy and accuracy trade-offs. This is the first systematic study of machine unlearning for HTR. Using metrics such as Accuracy, Character Error Rate (CER), Word Error Rate (WER), and Membership Inference Attacks (MIA) on the IAM and CVL datasets, we demonstrate that our method achieves state-of-the-art or superior performance for effective unlearning. These experiments show that our approach effectively safeguards privacy without compromising accuracy, opening new directions for document analysis research. Our code is publicly available at https://github.com/leitro/WIC-WriterIDConfusion-MachineUnlearning . Lei Kang 0002, Xuanshuo Fu, Lluís Gómez i Bigorda, Alicia Fornés, Ernest Valveny, Dimosthenis Karatzas |
Pattern Recognit. | 3 |
| 2025 | Measuring Text-Image Retrieval Fairness with Synthetic DataabstractIn this paper, we study social bias in cross-modal text-image retrieval systems, focusing on the interaction between textual queries and image responses. Despite the significant advancements in cross-modal retrieval models, the potential for social bias in their responses remains a pressing concern, necessitating a comprehensive framework for assessment and mitigation. We introduce a novel framework for evaluating social bias in cross-modal retrieval systems, leveraging a new dataset and appropriate metrics specifically designed for this purpose. Our dataset, Social Inclusive Synthetic Professionals Images (SISPI), comprises 49K images generated using state-of-the-art text-to-image models, ensuring a balanced representation of demographic groups across various professional roles. We use this dataset to conduct an extensive analysis of social bias (gender and ethnic) in state of the art cross-modal retrieval deep models, including CLIP, ALIGN, BLIP, FLAVA, COCA, and many others. Using diversity metrics, grounded in the distribution of different demographic groups' images in the retrieval rankings, we provide a quantitative measure of fairness, facilitating a detailed analysis of models' behavior. Our work sheds light on biases present in current cross-modal retrieval systems and emphasizes the importance of training data curation, providing a foundation for future research and development towards more equitable and unbiased models. The dataset and code of our framework is publicly available at https://sispi-benchmark.github.io/sispi-benchmark/. Lluís Gómez i Bigorda |
SIGIR | 1 |
| 2025 | EUFCC-340K: A faceted hierarchical dataset for metadata annotation in GLAM collectionsabstractAbstract In this paper, we address the challenges of automatic metadata annotation in the domain of Galleries, Libraries, Archives, and Museums (GLAMs) by introducing a novel dataset, EUFCC-340K, collected from the Europeana portal. Comprising over 340,000 images, the EUFCC-340K dataset is organized across multiple facets – Materials, Object Types, Disciplines, and Subjects – following a hierarchical structure based on the Art & Architecture Thesaurus (AAT). We developed several baseline models, incorporating multiple heads on a ConvNeXT backbone for multi-label image tagging on these facets, and fine-tuning a CLIP model with our image-text pairs. Our experiments to evaluate model robustness and generalization capabilities in two different test scenarios demonstrate the dataset’s utility in improving multi-label classification tools that have the potential to alleviate cataloging tasks in the cultural heritage sector. The EUFCC-340K dataset is publicly available at https://github.com/cesc47/EUFCC-340K . Francesc Net, Marc Folia, Pep Casals, Andrew D. Bagdanov, Lluís Gómez i Bigorda |
Multim. Tools Appl. | 5 |
| 2024 | GRIF-DM: Generation of Rich Impression Fonts Using Diffusion ModelsabstractFonts are integral to creative endeavors, design processes, and artistic productions. The appropriate selection of a font can significantly enhance artwork and endow advertisements with a higher level of expressivity. Despite the availability of numerous diverse font designs online, traditional retrieval-based methods for font selection are increasingly being supplanted by generation-based approaches. These newer methods offer enhanced flexibility, catering to specific user preferences and capturing unique stylistic impressions. However, current impression font techniques based on Generative Adversarial Networks (GANs) necessitate the utilization of multiple auxiliary losses to provide guidance during generation. Furthermore, these methods commonly employ weighted summation for the fusion of impression-related keywords. This leads to generic vectors with the addition of more impression keywords, ultimately lacking in detail generation capacity. In this paper, we introduce a diffusion-based method, termed GRIF-DM, to generate fonts that vividly embody specific impressions, utilizing an input consisting of a single letter and a set of descriptive impression keywords. The core innovation of GRIF-DM lies in the development of dual cross-attention modules, which process the characteristics of the letters and impression keywords independently but synergistically, ensuring effective integration of both types of information. Our experimental results, conducted on the MyFonts dataset, affirm that this method is capable of producing realistic, vibrant, and high-fidelity fonts that are closely aligned with user specifications. This confirms the potential of our approach to revolutionize font generation by accommodating a broad spectrum of user-driven design requirements. Our code is publicly available at https://github.com/leitro/GRIF-DM. Lei Kang 0002, Fei Yang 0004, Kai Wang 0060, Mohamed Ali Souibgui, Lluís Gómez i Bigorda, Alicia Fornés, Ernest Valveny, Dimosthenis Karatzas |
ECAI | 5 |
| 2024 | A Transformer-Based Object-Centric Approach for Date Estimation of Historical Photographs
Francesc Net, Núria Hernández, Adrià Molina, Lluís Gómez i Bigorda |
ECIR (3) | 4 |
| 2024 | Machine Unlearning for Document Classification
Lei Kang 0002, Mohamed Ali Souibgui, Fei Yang 0004, Lluís Gómez i Bigorda, Ernest Valveny, Dimosthenis Karatzas |
ICDAR (4) | 4 |
| 2023 | Show, Interpret and Tell: Entity-Aware Contextualised Image Captioning in WikipediaabstractHumans exploit prior knowledge to describe images, and are able to adapt their explanation to specific contextual information given, even to the extent of inventing plausible explanations when contextual information and images do not match. In this work, we propose the novel task of captioning Wikipedia images by integrating contextual knowledge. Specifically, we produce models that jointly reason over Wikipedia articles, Wikimedia images and their associated descriptions to produce contextualized captions. The same Wikimedia image can be used to illustrate different articles, and the produced caption needs to be adapted to the specific context allowing us to explore the limits of the model to adjust captions to different contextual information. Dealing with out-of-dictionary words and Named Entities is a challenging task in this domain. To address this, we propose a pre-training objective, Masked Named Entity Modeling (MNEM), and show that this pretext task results to significantly improved models. Furthermore, we verify that a model pre-trained in Wikipedia generalizes well to News Captioning datasets. We further define two different test splits according to the difficulty of the captioning task. We offer insights on the role and the importance of each modality and highlight the limitations of our model. Ali Furkan Biten, Andrés Mafla, Lluís Gómez i Bigorda, Dimosthenis Karatzas |
AAAI | 4 |
| 2023 | Text-DIAE: A Self-Supervised Degradation Invariant Autoencoder for Text Recognition and Document EnhancementabstractIn this paper, we propose a Text-Degradation Invariant Auto Encoder (Text-DIAE), a self-supervised model designed to tackle two tasks, text recognition (handwritten or scene-text) and document image enhancement. We start by employing a transformer-based architecture that incorporates three pretext tasks as learning objectives to be optimized during pre-training without the usage of labelled data. Each of the pretext objectives is specifically tailored for the final downstream tasks. We conduct several ablation experiments that confirm the design choice of the selected pretext tasks. Importantly, the proposed model does not exhibit limitations of previous state-of-the-art methods based on contrastive losses, while at the same time requiring substantially fewer data samples to converge. Finally, we demonstrate that our method surpasses the state-of-the-art in existing supervised and self-supervised settings in handwritten and scene text recognition and document image enhancement. Our code and trained models will be made publicly available at https://github.com/dali92002/SSL-OCR Mohamed Ali Souibgui, Sanket Biswas, Andrés Mafla, Ali Furkan Biten, Alicia Fornés, Yousri Kessentini, Josep Lladós 0001, Lluís Gómez i Bigorda, Dimosthenis Karatzas |
AAAI | 8 |
| 2023 | Transductive Learning for Near-Duplicate Image Detection in Scanned Photo Collections
Francesc Net, Marc Folia, Pep Casals, Lluís Gómez i Bigorda |
ICDAR (5) | 4 |
| 2022 | A Generic Image Retrieval Method for Date Estimation of Historical Document Collections
Adrià Molina, Lluís Gómez i Bigorda, Oriol Ramos Terrades, Josep Lladós 0001 |
DAS | 2 |
| 2022 | A Multilingual Approach to Scene Text Visual Question Answering
Josep Brugués i Pujolràs, Lluís Gómez i Bigorda, Dimosthenis Karatzas |
DAS | 2 |
| 2022 | Let there be a clock on the beach: Reducing Object Hallucination in Image CaptioningabstractExplaining an image with missing or non-existent objects is known as object bias (hallucination) in image captioning. This behaviour is quite common in the state-of-the-art captioning models which is not desirable by humans. To decrease the object hallucination in captioning, we propose three simple yet efficient training augmentation method for sentences which requires no new training data or increase in the model size. By extensive analysis, we show that the proposed methods can significantly diminish our models’ object bias on hallucination metrics. Moreover, we experimentally demonstrate that our methods decrease the dependency on the visual features. All of our code, configuration files and model weights are available online1. Ali Furkan Biten, Lluís Gómez i Bigorda, Dimosthenis Karatzas |
WACV | 2 |
| 2022 | Is An Image Worth Five Sentences? A New Look into Semantics for Image-Text MatchingabstractThe task of image-text matching aims to map representations from different modalities into a common joint visual-textual embedding. However, the most widely used datasets for this task, MSCOCO and Flickr30K, are actually image captioning datasets that offer a very limited set of relation-ships between images and sentences in their ground-truth annotations. This limited ground truth information forces us to use evaluation metrics based on binary relevance: given a sentence query we consider only one image as relevant. However, many other relevant images or captions may be present in the dataset. In this work, we propose two metrics that evaluate the degree of semantic relevance of retrieved items, independently of their annotated binary relevance. Additionally, we incorporate a novel strategy that uses an image captioning metric, CIDEr, to define a Semantic Adaptive Margin (SAM) to be optimized in a standard triplet loss. By incorporating our formulation to existing models, a large improvement is obtained in scenarios where available training data is limited. We also demonstrate that the performance on the annotated image-caption pairs is maintained while improving on other non-annotated relevant items when employing the full training set. The code for our new metric can be found at github.com/furkanbiten/ncs_metric and the model implementation at github.com/andrespmd/semantic_adaptive_margin. Ali Furkan Biten, Andrés Mafla, Lluís Gómez i Bigorda, Dimosthenis Karatzas |
WACV | 3 |
| 2022 | One-shot Compositional Data Generation for Low Resource Handwritten Text RecognitionabstractLow resource Handwritten Text Recognition (HTR) is a hard problem due to the scarce annotated data and the very limited linguistic information (dictionaries and language models). For example, in the case of historical ciphered manuscripts, which are usually written with invented alphabets to hide the message contents. Thus, in this paper we address this problem through a data generation technique based on Bayesian Program Learning (BPL). Contrary to traditional generation approaches, which require a huge amount of annotated images, our method is able to generate human-like handwriting using only one sample of each symbol in the alphabet. After generating symbols, we create synthetic lines to train state-of-the-art HTR architectures in a segmentation free fashion. Quantitative and qualitative analyses were carried out and confirm the effectiveness of the proposed method. Mohamed Ali Souibgui, Ali Furkan Biten, Sounak Dey, Alicia Fornés, Yousri Kessentini, Lluís Gómez i Bigorda, Dimosthenis Karatzas, Josep Lladós 0001 |
WACV | 6 |
| 2021 | Date Estimation in the Wild of Scanned Historical Photos: An Image Retrieval Approach
Adrià Molina, Pau Riba, Lluís Gómez i Bigorda, Oriol Ramos Terrades, Josep Lladós 0001 |
ICDAR (2) | 3 |
| 2021 | Learning to Rank Words: Optimizing Ranking Metrics for Word Spotting
Pau Riba, Adrià Molina, Lluís Gómez i Bigorda, Oriol Ramos Terrades, Josep Lladós 0001 |
ICDAR (2) | 3 |
| 2021 | Multi-Modal Reasoning Graph for Scene-Text Based Fine-Grained Image Classification and RetrievalabstractScene text instances found in natural images carry explicit semantic information that can provide important cues to solve a wide array of computer vision problems. In this paper, we focus on leveraging multi-modal content in the form of visual and textual cues to tackle the task of fine-grained image classification and retrieval. First, we obtain the text instances from images by employing a text reading system. Then, we combine textual features with salient image regions to exploit the complementary information carried by the two sources. Specifically, we employ a Graph Convolutional Network to perform multi-modal reasoning and obtain relationship-enhanced features by learning a common semantic space between salient objects and text found in an image. By obtaining an enhanced set of visual and textual features, the proposed model greatly outperforms previous state-of-the-art in two different tasks, fine-grained classification and image retrieval in the Con-Text[23] and Drink Bottle[4] datasets. Andrés Mafla, Sounak Dey, Ali Furkan Biten, Lluís Gómez i Bigorda, Dimosthenis Karatzas |
WACV | 4 |
| 2021 | StacMR: Scene-Text Aware Cross-Modal RetrievalabstractRecent models for cross-modal retrieval have benefited from an increasingly rich understanding of visual scenes, afforded by scene graphs and object interactions to mention a few. This has resulted in an improved matching between the visual representation of an image and the textual representation of its caption. Yet, current visual representations overlook a key aspect: the text appearing in images, which may contain crucial information for retrieval. In this paper, we first propose a new dataset that allows exploration of cross-modal retrieval where images contain scene-text instances. Then, armed with this dataset, we describe several approaches which leverage scene text, including a better scene-text aware cross-modal retrieval method which uses specialized representations for text from the captions and text from the visual scene, and reconcile them in a common embedding space. Extensive experiments confirm that cross-modal retrieval approaches benefit from scene text and highlight interesting research questions worth exploring further. Dataset and code are available at europe.naverlabs.com/stacmr. Andrés Mafla, Rafael S. Rezende, Lluís Gómez i Bigorda, Diane Larlus, Dimosthenis Karatzas |
WACV | 3 |
| 2021 | Asking questions on handwritten document collections
Minesh Mathew, Lluís Gómez i Bigorda, Dimosthenis Karatzas, C. V. Jawahar |
Int. J. Document Anal. Recognit. | 2 |
| 2021 | Real-time Lexicon-free Scene Text Retrieval
Andrés Mafla, Rubèn Tito, Sounak Dey, Lluís Gómez i Bigorda, Marçal Rusiñol, Ernest Valveny, Dimosthenis Karatzas |
Pattern Recognit. | 4 |
| 2021 | Multimodal grid features and cell pointers for scene text visual question answering
Lluís Gómez i Bigorda, Ali Furkan Biten, Rubèn Tito, Andrés Mafla, Marçal Rusiñol, Ernest Valveny, Dimosthenis Karatzas |
Pattern Recognit. Lett. | 1 |
| 2020 | Location Sensitive Image Retrieval and Tagging
Raul Gomez, Jaume Gibert, Lluís Gómez i Bigorda, Dimosthenis Karatzas |
ECCV (16) | 3 |
| 2020 | Text Recognition - Real World Data and Where to Find ThemabstractWe present a method for exploiting weakly annotated images to improve text extraction pipelines. The approach uses an arbitrary end-to-end text recognition system to obtain text region proposals and their, possibly erroneous, transcriptions. The method includes matching of imprecise transcriptions to weak annotations and an edit distance guided neighbourhood search. It produces nearly error-free, localised instances of scene text, which we treat as “pseudo ground truth” (PGT). The method is applied to two weakly-annotated datasets. Training with the extracted PGT consistently improves the accuracy of a state of the art recognition model, by 3.7% on average, across different benchmark datasets (image domains) and 24.5% on one of the weakly annotated datasets11Acknowledgements. The authors were supported by Czech Technical University student grant SGS20/171/0HK3/3TJ13, the MEYS VVV project CZ.02.1.01/0.010.0J16 019/0000765 Research Center for Informatics, the Spanish Research project TIN2017-89779-P and the CERCA Programme / Generalitat de Catalunya. Klára Janousková, Jiri Matas, Lluís Gómez i Bigorda, Dimosthenis Karatzas |
ICPR | 3 |
| 2020 | RoadText-1K: Text Detection & Recognition Dataset for Driving VideosabstractPerceiving text is crucial to understand semantics of outdoor scenes and hence is a critical requirement to build intelligent systems for driver assistance and self-driving. Most of the existing datasets for text detection and recognition comprise still images and are mostly compiled keeping text in mind. This paper introduces a new "RoadText-1K" dataset for text in driving videos. The dataset is 20 times larger than the existing largest dataset for text in videos. Our dataset comprises 1000 video clips of driving without any bias towards text and with annotations for text bounding boxes and transcriptions in every frame. State of the art methods for text detection, recognition and tracking are evaluated on the new dataset and the results signify the challenges in unconstrained driving videos compared to existing datasets. This suggests that RoadText-1K is suited for research and development of reading systems, robust enough to be incorporated into more complex downstream tasks like driver assistance and self-driving. The dataset can be found at http://cvit.iiit.ac.in/research/projects/cvit-projects/roadtext-1k. Sangeeth Reddy, Minesh Mathew, Lluís Gómez i Bigorda, Marçal Rusiñol, Dimosthenis Karatzas, C. V. Jawahar |
ICRA | 3 |
| 2020 | Exploring Hate Speech Detection in Multimodal PublicationsabstractIn this work we target the problem of hate speech detection in multimodal publications formed by a text and an image. We gather and annotate a large scale dataset from Twitter, MMHS150K, and propose different models that jointly analyze textual and visual information for hate speech detection, comparing them with unimodal detection. We provide quantitative and qualitative results and analyze the challenges of the proposed task. We find that, even though images are useful for the hate speech detection task, current multimodal models cannot outperform models analyzing only text. We discuss why and open the field and the dataset for further research. Raul Gomez, Jaume Gibert, Lluís Gómez i Bigorda, Dimosthenis Karatzas |
WACV | 3 |
| 2020 | Fine-grained Image Classification and Retrieval by Combining Visual and Locally Pooled Textual FeaturesabstractText contained in an image carries high-level semantics that can be exploited to achieve richer image understanding. In particular, the mere presence of text provides strong guiding content that should be employed to tackle a diversity of computer vision tasks such as image retrieval, fine-grained classification, and visual question answering. In this paper, we address the problem of fine-grained classification and image retrieval by leveraging textual information along with visual cues to comprehend the existing intrinsic relation between the two modalities. The novelty of the proposed model consists of the usage of a PHOC descriptor to construct a bag of textual words along with a Fisher Vector Encoding that captures the morphology of text. This approach provides a stronger multimodal representation for this task and as our experiments demonstrate, it achieves state-of-the-art results on two different tasks, fine-grained classification and image retrieval. The code of this model will be publicly available at1. Andrés Mafla, Sounak Dey, Ali Furkan Biten, Lluís Gómez i Bigorda, Dimosthenis Karatzas |
WACV | 4 |
| 2019 | Good News, Everyone! Context Driven Entity-Aware Captioning for News ImagesabstractCurrent image captioning systems perform at a merely descriptive level, essentially enumerating the objects in the scene and their relations. Humans, on the contrary, interpret images by integrating several sources of prior knowledge of the world. In this work, we aim to take a step closer to producing captions that offer a plausible interpretation of the scene, by integrating such contextual information into the captioning pipeline. For this we focus on the captioning of images used to illustrate news articles. We propose a novel captioning method that is able to leverage contextual information provided by the text of news articles associated with an image. Our model is able to selectively draw information from the article guided by visual cues, and to dynamically extend the output dictionary to out-of-vocabulary named entities that appear in the context source. Furthermore we introduce ``GoodNews'', the largest news image captioning dataset in the literature and demonstrate state-of-the-art results. Ali Furkan Biten, Lluís Gómez i Bigorda, Marçal Rusiñol, Dimosthenis Karatzas |
CVPR | 2 |
| 2019 | Scene Text Visual Question AnsweringabstractCurrent visual question answering datasets do not consider the rich semantic information conveyed by text within an image. In this work, we present a new dataset, ST-VQA, that aims to highlight the importance of exploiting high-level semantic information present in images as textual cues in the Visual Question Answering process. We use this dataset to define a series of tasks of increasing difficulty for which reading the scene text in the context provided by the visual information is necessary to reason and generate an appropriate answer. We propose a new evaluation metric for these tasks to account both for reasoning errors as well as shortcomings of the text recognition module. In addition we put forward a series of baseline methods, which provide further insight to the newly released dataset, and set the scene for further research. Ali Furkan Biten, Rubèn Tito, Andrés Mafla, Lluís Gómez i Bigorda, Marçal Rusiñol, C. V. Jawahar, Ernest Valveny, Dimosthenis Karatzas |
ICCV | 4 |
| 2019 | ICDAR 2019 Competition on Scene Text Visual Question AnsweringabstractThis paper presents final results of ICDAR 2019 Scene Text Visual Question Answering competition (ST-VQA). ST-VQA introduces an important aspect that is not addressed by any Visual Question Answering system up to date, namely the incorporation of scene text to answer questions asked about an image. The competition introduces a new dataset comprising 23,038 images annotated with 31,791 question / answer pairs where the answer is always grounded on text instances present in the image. The images are taken from 7 different public computer vision datasets, covering a wide range of scenarios. The competition was structured in three tasks of increasing difficulty, that require reading the text in a scene and understanding it in the context of the scene, to correctly answer a given question. A novel evaluation metric is presented, which elegantly assesses both key capabilities expected from an optimal model: text recognition and image understanding. A detailed analysis of results from different participants is showcased, which provides insight into the current capabilities of VQA systems that can read. We firmly believe the dataset proposed in this challenge will be an important milestone to consider towards a path of more robust and general models that can exploit scene text to achieve holistic image understanding. Ali Furkan Biten, Rubèn Tito, Andrés Mafla, Lluís Gómez i Bigorda, Marçal Rusiñol, Minesh Mathew, C. V. Jawahar, Ernest Valveny, Dimosthenis Karatzas |
ICDAR | 4 |
| 2019 | Selective Style Transfer for TextabstractThis paper explores the possibilities of image style transfer applied to text maintaining the original transcriptions. Results on different text domains (scene text, machine printed text and handwritten text) and cross-modal results demonstrate that this is feasible, and open different research lines. Furthermore, two architectures for selective style transfer, which means transferring style to only desired image pixels, are proposed. Finally, scene text selective style transfer is evaluated as a data augmentation technique to expand scene text detection datasets, resulting in a boost of text detectors performance. Our implementation of the described models is publicly available. Raul Gomez, Ali Furkan Biten, Lluís Gómez i Bigorda, Jaume Gibert, Dimosthenis Karatzas, Marçal Rusiñol |
ICDAR | 3 |
| 2019 | Self-Supervised Visual Representations for Cross-Modal RetrievalabstractCross-modal retrieval methods have been significantly improved in last years with the use of deep neural networks and large-scale annotated datasets such as ImageNet and Places. However, collecting and annotating such datasets requires a tremendous amount of human effort and, besides, their annotations are limited to discrete sets of popular visual classes that may not be representative of the richer semantics found on large-scale cross-modal retrieval datasets. In this paper, we present a self-supervised cross-modal retrieval framework that leverages as training data the correlations between images and text on the entire set of Wikipedia articles. Our method consists in training a CNN to predict: (1) the semantic context of the article in which an image is more probable to appear as an illustration, and (2) the semantic context of its caption. Our experiments demonstrate that the proposed method is not only capable of learning discriminative visual representations for solving vision tasks like classification, but that the learned representations are better for cross-modal retrieval when compared to supervised pre-training of the network on the ImageNet dataset. Lluís Gómez i Bigorda, Marçal Rusiñol, Dimosthenis Karatzas, C. V. Jawahar |
ICMR | 2 |
| 2019 | FAST: Facilitated and Accurate Scene Text Proposals through FCN Guided Pruning
Dena Bazazian, Raul Gomez, Anguelos Nicolaou, Lluís Gómez i Bigorda, Dimosthenis Karatzas, Andrew D. Bagdanov |
Pattern Recognit. Lett. | 4 |
| 2018 | Cutting Sayre's Knot: Reading Scene Text without Segmentation. Application to Utility MetersabstractIn this paper we present a segmentation-free system for reading text in natural scenes. A CNN architecture is trained in an end-to-end manner, and is able to directly output readings without any explicit text localization step. In order to validate our proposal, we focus on the specific case of reading utility meters. We present our results in a large dataset of images acquired by different users and devices, so text appears in any location, with different sizes, fonts and lengths, and the images present several distortions such as dirt, illumination highlights or blur. Lluís Gómez i Bigorda, Marçal Rusiñol, Dimosthenis Karatzas |
DAS | 1 |
| 2018 | The Robust Reading Competition Annotation and Evaluation PlatformabstractThe ICDAR Robust Reading Competition (RRC), initiated in 2003 and re-established in 2011, has become a de-facto evaluation standard for robust reading systems and algorithms. Concurrent with its second incarnation in 2011, a continuous effort started to develop an on-line framework to facilitate the hosting and management of competitions. This paper outlines the Robust Reading Competition Annotation and Evaluation Platform, the backbone of the competitions. The RRC Annotation and Evaluation Platform is a modular framework, fully accessible through on-line interfaces. It comprises a collection of tools and services for managing all processes involved with defining and evaluating a research task, from dataset definition to annotation management, evaluation specification and results analysis. Although the framework has been designed with robust reading research in mind, many of the provided tools are generic by design. All aspects of the RRC Annotation and Evaluation Framework are available for research use. Dimosthenis Karatzas, Lluís Gómez i Bigorda, Anguelos Nicolaou, Marçal Rusiñol |
DAS | 2 |
| 2018 | Single Shot Scene Text Retrieval
Lluís Gómez i Bigorda, Andrés Mafla, Marçal Rusiñol, Dimosthenis Karatzas |
ECCV (14) | 1 |
| 2017 | Self-Supervised Learning of Visual Features through Embedding Images into Text Topic SpacesabstractEnd-to-end training from scratch of current deep architectures for new computer vision problems would require Imagenet-scale datasets, and this is not always possible. In this paper we present a method that is able to take advantage of freely available multi-modal content to train computer vision algorithms without human supervision. We put forward the idea of performing self-supervised learning of visual features by mining a large scale corpus of multi-modal (text and image) documents. We show that discriminative visual features can be learnt efficiently by training a CNN to predict the semantic context in which a particular image is more probable to appear as an illustration. For this we leverage the hidden semantic structures discovered in the text corpus with a well-known topic modeling technique. Our experiments demonstrate state of the art performance in image classification, object detection, and multi-modal retrieval compared to recent self-supervised or natural-supervised approaches. Lluís Gómez i Bigorda, Marçal Rusiñol, Dimosthenis Karatzas, C. V. Jawahar |
CVPR | 1 |
| 2017 | LSDE: Levenshtein Space Deep Embedding for Query-by-String Word SpottingabstractIn this paper we present the LSDE string representation and its application to handwritten word spotting. LSDE is a novel embedding approach for representing strings that learns a space in which distances between projected points are correlated with the Levenshtein edit distance between the original strings. We show how such a representation produces a more semantically interpretable retrieval from the user's perspective than other state of the art ones such as PHOC and DCToW. We also conduct a preliminary handwritten word spotting experiment on the George Washington dataset. Lluís Gómez i Bigorda, Marçal Rusiñol, Dimosthenis Karatzas |
ICDAR | 1 |
| 2017 | ICDAR2017 Robust Reading Challenge on COCO-TextabstractThis report presents the final results of the ICDAR 2017 Robust Reading Challenge on COCO-Text. A challenge on scene text detection and recognition based on the largest real scene text dataset currently available: the COCO-Text dataset. The competition is structured around three tasks: Text Localization, Cropped Word Recognition and End-To-End Recognition. The competition received a total of 27 submissions over the different opened tasks. This report describes the datasets and the ground truth, details the performance evaluation protocols used and presents the final results along with a brief summary of the participating methods. Raul Gomez, Baoguang Shi, Lluís Gómez i Bigorda, Lukás Neumann, Andreas Veit, Jiri Matas, Serge J. Belongie, Dimosthenis Karatzas |
ICDAR | 3 |
| 2017 | ICDAR2017 Robust Reading Challenge on Omnidirectional VideoabstractResults of ICDAR 2017 Robust Reading Challenge on Omnidirectional Video are presented. This competition uses Downtown Osaka Scene Text (DOST) Dataset that was captured in Osaka, Japan with an omnidirectional camera. Hence, it consists of sequential images (videos) of different view angles. Regarding the sequential images as videos (video mode), two tasks of localisation and end-to-end recognition are prepared. Regarding them as a set of still images (still image mode), three tasks of localisation, cropped word recognition and end-to-end recognition are prepared. As the dataset has been captured in Japan, the dataset contains Japanese text but also include text consisting of alphanumeric characters (Latin text). Hence, a submitted result for each task is evaluated in three ways: using Japanese only ground truth (GT), using Latin only GT and using combined GTs of both. Finally, by the submission deadline, we have received two submissions in the text localisation task of the still image mode. We intend to continue the competition in the open mode. Expecting further submissions, in this report we provide baseline results in all the tasks in addition to the submissions from the community. Masakazu Iwamura, Naoyuki Morimoto, Keishi Tainaka, Dena Bazazian, Lluís Gómez i Bigorda, Dimosthenis Karatzas |
ICDAR | 5 |
| 2017 | Improving patch-based scene text script identification with ensembles of conjoined networks
Lluís Gómez i Bigorda, Anguelos Nicolaou, Dimosthenis Karatzas |
Pattern Recognit. | 1 |
| 2017 | TextProposals: A text-specific selective search algorithm for word spotting in the wild
Lluís Gómez i Bigorda, Dimosthenis Karatzas |
Pattern Recognit. | 1 |
| 2016 | A Fine-Grained Approach to Scene Text Script IdentificationabstractThis paper focuses on the problem of script identification in unconstrained scenarios. Script identification is an important prerequisite to recognition, and an indispensable condition for automatic text understanding systems designed for multi-language environments. Although widely studied for document images and handwritten documents, it remains an almost unexplored territory for scene text images. We detail a novel method for script identification in natural images that combines convolutional features and the Naive-Bayes Nearest Neighbor classifier. The proposed framework efficiently exploits the discriminative power of small stroke-parts, in a fine-grained classification framework. In addition, we propose a new public benchmark dataset for the evaluation of joint text detection and script identification in natural scenes. Experiments done in this new dataset demonstrate that the proposed method yields state of the art results, while it generalizes well to different datasets and variable number of scripts. The evidence provided shows that multi-lingual scene text recognition in the wild is a viable proposition. Source code of the proposed method is made available online. Lluís Gómez i Bigorda, Dimosthenis Karatzas |
DAS | 1 |
| 2016 | Visual Script and Language IdentificationabstractIn this paper we introduce a script identification method based on hand-crafted texture features and an artificial neural network. The proposed pipeline achieves near state-of-the-art performance for script identification of video-text and state-of-the-art performance on visual language identification of handwritten text. More than using the deep network as a classifier, the use of its intermediary activations as a learned metric demonstrates remarkable results and allows the use of discriminative models on unknown classes. Comparative experiments in video-text and text in the wild datasets provide insights on the internals of the proposed deep network. Anguelos Nicolaou, Andrew D. Bagdanov, Lluís Gómez i Bigorda, Dimosthenis Karatzas |
DAS | 3 |
| 2016 | A fast hierarchical method for multi-script and arbitrary oriented scene text extraction
Lluís Gómez i Bigorda, Dimosthenis Karatzas |
Int. J. Document Anal. Recognit. | 1 |
| 2015 | Object proposals for text extraction in the wildabstractObject Proposals is a recent computer vision technique receiving increasing interest from the research community. Its main objective is to generate a relatively small set of bounding box proposals that are most likely to contain objects of interest. The use of Object Proposals techniques in the scene text understanding field is innovative. Motivated by the success of powerful while expensive techniques to recognize words in a holistic way, Object Proposals techniques emerge as an alternative to the traditional text detectors. In this paper we study to what extent the existing generic Object Proposals methods may be useful for scene text understanding. Also, we propose a new Object Proposals algorithm that is specifically designed for text and compare it with other generic methods in the state of the art. Experiments show that our proposal is superior in its ability of producing good quality word proposals in an efficient way. The source code of our method is made publicly available1. Lluís Gómez i Bigorda, Dimosthenis Karatzas |
ICDAR | 1 |
| 2015 | Efficient indexing for Query By String text retrievalabstractThis paper deals with Query By String word spotting in scene images. A hierarchical text segmentation algorithm based on text specific selective search is used to find text regions. These regions are indexed per character n-grams present in the text region. An attribute representation based on Pyramidal Histogram of Characters (PHOC) is used to compare text regions with the query text. For generation of the index a similar attribute space based Pyramidal Histogram of character n-grams is used. These attribute models are learned using linear SVMs over the Fisher Vector [1] representation of the images along with the PHOC labels of the corresponding strings. Suman K. Ghosh, Lluís Gómez i Bigorda, Dimosthenis Karatzas, Ernest Valveny |
ICDAR | 2 |
| 2015 | ICDAR 2015 competition on Robust ReadingabstractResults of the ICDAR 2015 Robust Reading Competition are presented. A new Challenge 4 on Incidental Scene Text has been added to the Challenges on Born-Digital Images, Focused Scene Images and Video Text. Challenge 4 is run on a newly acquired dataset of 1,670 images evaluating Text Localisation, Word Recognition and End-to-End pipelines. In addition, the dataset for Challenge 3 on Video Text has been substantially updated with more video sequences and more accurate ground truth data. Finally, tasks assessing End-to-End system performance have been introduced to all Challenges. The competition took place in the first quarter of 2015, and received a total of 44 submissions. Only the tasks newly introduced in 2015 are reported on. The datasets, the ground truth specification and the evaluation protocols are presented together with the results and a brief summary of the participating methods. Dimosthenis Karatzas, Lluís Gómez i Bigorda, Anguelos Nicolaou, Suman K. Ghosh, Andrew D. Bagdanov, Masakazu Iwamura, Jiri Matas, Lukás Neumann, Vijay Chandrasekhar 0001, Shijian Lu, Faisal Shafait, Seiichi Uchida, Ernest Valveny |
ICDAR | 2 |
| 2014 | An On-line Platform for Ground Truthing and Performance Evaluation of Text Extraction SystemsabstractThis work presents a set of on-line software tools for creating ground truth and calculating performance evaluation metrics for text extraction tasks such as localization, segmentation and recognition. The platform supports the definition of comprehensive ground truth information at different text representation levels while it offers centralised management and quality control of the ground truthing effort. It implements a range of state of the art performance evaluation algorithms and offers functionality for the definition of evaluation scenarios, on-line calculation of various performance metrics and visualisation of the results. The presented platform, which comprises the backbone of the ICDAR 2011 (challenge 1) and 2013 (challenges 1 and 2) Robust Reading competitions, is now made available for public use. Dimosthenis Karatzas, Sergi Robles, Lluís Gómez i Bigorda |
Document Analysis Systems | 3 |
| 2014 | MSER-Based Real-Time Text Detection and TrackingabstractWe present a hybrid algorithm for detection and tracking of text in natural scenes that goes beyond the full-detection approaches in terms of time performance optimization. A state-of-the-art scene text detection module based on Maximally Stable Extremal Regions (MSER) is used to detect text asynchronously, while on a separate thread detected text objects are tracked by MSER propagation. The cooperation of these two modules yields real time video processing at high frame rates even on low-resource devices. Lluís Gómez i Bigorda, Dimosthenis Karatzas |
ICPR | 1 |
| 2013 | Multi-script Text Extraction from Natural ScenesabstractScene text extraction methodologies are usually based in classification of individual regions or patches, using a priori knowledge for a given script or language. Human perception of text, on the other hand, is based on perceptual organisation through which text emerges as a perceptually significant group of atomic objects. Therefore humans are able to detect text even in languages and scripts never seen before. In this paper, we argue that the text extraction problem could be posed as the detection of meaningful groups of regions. We present a method built around a perceptual organisation framework that exploits collaboration of proximity and similarity laws to create text-group hypotheses. Experiments demonstrate that our algorithm is competitive with state of the art approaches on a standard dataset covering text in variable orientations and two languages. Lluís Gómez i Bigorda, Dimosthenis Karatzas |
ICDAR | 1 |
| 2013 | ICDAR 2013 Robust Reading CompetitionabstractThis report presents the final results of the ICDAR 2013 Robust Reading Competition. The competition is structured in three Challenges addressing text extraction in different application domains, namely born-digital images, real scene images and real-scene videos. The Challenges are organised around specific tasks covering text localisation, text segmentation and word recognition. The competition took place in the first quarter of 2013, and received a total of 42 submissions over the different tasks offered. This report describes the datasets and ground truth specification, details the performance evaluation protocols used and presents the final results along with a brief summary of the participating methods. Dimosthenis Karatzas, Faisal Shafait, Seiichi Uchida, Masakazu Iwamura, Lluís Gómez i Bigorda, Sergi Robles, Joan Mas Romeu, David Fernández Mota, Jon Almazán, Lluís-Pere de las Heras |
ICDAR | 5 |