EDBT 2026 Demo / reviewers in the wild / expert
Dimosthenis Karatzas
dblp:03/6509 · also Dimosthenis A. Karatzas
· DBLP profile ↗
70ranked-venue papers in the field
8as first author
24since 2021 · last 2025
0000-0001-8762-4454ORCID · verified
Domains — venue-derived; a paper can count in several
Other / Interdisciplinary · 67 (8 first)Information Retrieval & Web Search · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Position-Aware Stamp-Like Adversarial Attack for Document Classification
Lei Kang 0002, Maura Pintor, Dimosthenis Karatzas |
ICDAR (4) | 4 |
| 2025 | LLM-Driven Medical Document Analysis: Enhancing Trustworthy Pathology and Differential Diagnosis
Lei Kang 0002, Xuanshuo Fu, Oriol Ramos Terrades, Javier Vazquez-Corral, Ernest Valveny, Dimosthenis Karatzas |
ICDAR (3) | 6 |
| 2025 | ICDAR 2025 Handwritten Notes Understanding Challenge
Aniket Pal, Sanket Biswas, Alloy Das, Ayush Lodh, Priyanka Banerjee, Soumitri Chattopadhyay, Ajoy Mondal, Dimosthenis Karatzas, Josep Lladós 0001, C. V. Jawahar |
ICDAR (5) | 8 |
| 2025 | ComicsPAP: Understanding Comic Strips by Picking the Correct Panel
Emanuele Vivoli, Artemis Llabrés, Mohamed Ali Souibgui, Marco Bertini 0001, Ernest Valveny, Dimosthenis Karatzas |
ICDAR (1) | 6 |
| 2025 | AVIR: Adaptive Visual In-Document Retrieval for Efficient Multi-Page Document Question AnsweringabstractMulti‑page Document Visual Question Answering (MP‑DocVQA) remains challenging because long documents not only strain computational resources but also reduce the effectiveness of the attention mechanism in large vision–language models (LVLMs). We tackle these issues with an Adaptive Visual In‑document Retrieval (AVIR) framework. A lightweight retrieval model first scores each page for question relevance. Pages are then clustered according to the score distribution to adaptively select relevant content. The clustered pages are screened again by Top-K to keep the context compact. However, for short documents, clustering reliability decreases, so we use a relevance probability threshold to select pages. The selected pages alone are fed to a frozen LVLM for answer generation, eliminating the need for model fine‑tuning. The proposed AVIR framework reduces the average page count required for question answering by 70%, while achieving an ANLS of 84.58% on the MP-DocVQA dataset—surpassing previous methods with significantly lower computational cost. The effectiveness of the proposed AVIR is also verified on the SlideVQA and DUDE benchmarks. Our code will be made publicly available upon acceptance. Zongmin Li, Yachuan Li, Lei Kang 0002, Dimosthenis Karatzas, Wenkang Ma |
MMAsia | 4 |
| 2024 | Multi-page Document VQA with Recurrent Memory Transformer
Lei Kang 0002, Dimosthenis Karatzas |
DAS | 3 |
| 2024 | Image-Text Matching for Large-Scale Book Collections
Artemis Llabrés, Arka Ujjal Dey, Dimosthenis Karatzas, Ernest Valveny |
DAS | 3 |
| 2024 | Machine Unlearning for Document Classification
Lei Kang 0002, Mohamed Ali Souibgui, Fei Yang 0004, Lluís Gómez i Bigorda, Ernest Valveny, Dimosthenis Karatzas |
ICDAR (4) | 6 |
| 2024 | Multi-page Document Visual Question Answering Using Self-attention Scoring Mechanism
Lei Kang 0002, Rubèn Tito, Ernest Valveny, Dimosthenis Karatzas |
ICDAR (6) | 4 |
| 2024 | Federated Document Visual Question Answering: A Pilot Study
Dimosthenis Karatzas |
ICDAR (6) | 2 |
| 2024 | Privacy-Aware Document Visual Question Answering
Rubèn Tito, Marlon Tobaben, Raouf Kerkouche, Mohamed Ali Souibgui, Kangsoo Jung, Joonas Jälkö, Vincent Poulain D'Andecy, Aurélie Joseph, Lei Kang 0002, Ernest Valveny, Antti Honkela, Mario Fritz, Dimosthenis Karatzas |
ICDAR (6) | 14 |
| 2024 | Multimodal Transformer for Comics Text-Cloze
Emanuele Vivoli, Joan Lafuente Baeza, Ernest Valveny, Dimosthenis Karatzas |
ICDAR (6) | 4 |
| 2024 | Counting the Corner Cases: Revisiting Robust Reading Challenge Data Sets, Evaluation Protocols, and Metrics
Jerod J. Weinman, Amelia Gómez Grabowska, Dimosthenis Karatzas |
ICDAR (4) | 3 |
| 2023 | Accelerating Transformer-Based Scene Text Detection and Recognition via Token Pruning
Sergi Garcia-Bordils, Dimosthenis Karatzas, Marçal Rusiñol |
ICDAR (6) | 2 |
| 2023 | DocILE Benchmark for Document Information Localization and Extraction
Stepán Simsa, Milan Sulc, Michal Uricár, Ahmed Hamdi, Matej Kocián, Matyás Skalický, Jiri Matas, Antoine Doucet, Mickaël Coustaty, Dimosthenis Karatzas |
ICDAR (2) | 11 |
| 2023 | ICDAR 2023 Competition on RoadText Video Text Detection, Tracking and Recognition
George Tom, Minesh Mathew, Sergi Garcia-Bordils, Dimosthenis Karatzas, C. V. Jawahar |
ICDAR (2) | 4 |
| 2023 | Reading Between the Lanes: Text VideoQA on the Road
George Tom, Minesh Mathew, Sergi Garcia-Bordils, Dimosthenis Karatzas, C. V. Jawahar |
ICDAR (6) | 4 |
| 2023 | ICDAR 2023 Competition on Video Text Reading for Dense and Small Text
Weijia Wu 0001, Yuzhong Zhao, Zhuang Li 0002, Zheng Shou 0001, Umapada Pal 0001, Dimosthenis Karatzas, Xiang Bai |
ICDAR (2) | 7 |
| 2023 | ICDAR 2023 Competition on Reading the Seal Title
Wenwen Yu, Mingrui Chen 0001, Ning Lu 0003, Yinlong Wen, Dimosthenis Karatzas, Xiang Bai |
ICDAR (2) | 7 |
| 2023 | ICDAR 2023 Competition on Structured Text Extraction from Visually-Rich Document Images
Wenwen Yu, Chengquan Zhang, Haoyu Cao 0001, Wei Hua 0005, Bohan Li 0010, Mingrui Chen 0001, Jianfeng Kuang, Mengjun Cheng, Yuning Du, Shikun Feng, Xiaoguang Hu, Pengyuan Lv, Yuechen Yu, Wanxiang Che, Errui Ding, Cheng-Lin Liu 0001, Jiebo Luo 0001, Shuicheng Yan, Min Zhang 0005, Dimosthenis Karatzas, Xing Sun 0001, Jingdong Wang 0001, Xiang Bai |
ICDAR (2) | 24 |
| 2022 | Read While You Drive - Multilingual Text Tracking on the Road
Sergi Garcia-Bordils, George Tom, Sangeeth Reddy, Minesh Mathew, Marçal Rusiñol, C. V. Jawahar, Dimosthenis Karatzas |
DAS | 7 |
| 2022 | A Multilingual Approach to Scene Text Visual Question Answering
Josep Brugués i Pujolràs, Lluís Gómez i Bigorda, Dimosthenis Karatzas |
DAS | 3 |
| 2021 | Document Collection Visual Question Answering
Rubèn Tito, Dimosthenis Karatzas, Ernest Valveny |
ICDAR (2) | 2 |
| 2021 | ICDAR 2021 Competition on Document Visual Question Answering
Rubèn Tito, Minesh Mathew, C. V. Jawahar, Ernest Valveny, Dimosthenis Karatzas |
ICDAR (4) | 5 |
| 2019 | Can One Deep Learning Model Learn Script-Independent Multilingual Word-Spotting?abstractWord spotting has gained increased attention lately as it can be used to extract textual information from handwritten documents and scene-text images. Current word spotting approaches are designed to work on a single language and/or script. Building intelligent models that learn script-independent multilingual word-spotting is challenging due to the large variability of multilingual alphabets and symbols. We used ResNet-152 and the Pyramidal Histogram of Characters (PHOC) embedding to build a one-model script-independent multilingual word-spotting and we tested it on Latin, Arabic, and Bangla (Indian) languages. The one-model we propose performs on par with the multi-model language-specific word-spotting system, and thus, reduces the number of models needed for each script and/or language. Mohammed Al-Rawi, Ernest Valveny, Dimosthenis Karatzas |
ICDAR | 3 |
| 2019 | ICDAR 2019 Competition on Scene Text Visual Question AnsweringabstractThis paper presents final results of ICDAR 2019 Scene Text Visual Question Answering competition (ST-VQA). ST-VQA introduces an important aspect that is not addressed by any Visual Question Answering system up to date, namely the incorporation of scene text to answer questions asked about an image. The competition introduces a new dataset comprising 23,038 images annotated with 31,791 question / answer pairs where the answer is always grounded on text instances present in the image. The images are taken from 7 different public computer vision datasets, covering a wide range of scenarios. The competition was structured in three tasks of increasing difficulty, that require reading the text in a scene and understanding it in the context of the scene, to correctly answer a given question. A novel evaluation metric is presented, which elegantly assesses both key capabilities expected from an optimal model: text recognition and image understanding. A detailed analysis of results from different participants is showcased, which provides insight into the current capabilities of VQA systems that can read. We firmly believe the dataset proposed in this challenge will be an important milestone to consider towards a path of more robust and general models that can exploit scene text to achieve holistic image understanding. Ali Furkan Biten, Rubèn Tito, Andrés Mafla, Lluís Gómez i Bigorda, Marçal Rusiñol, Minesh Mathew, C. V. Jawahar, Ernest Valveny, Dimosthenis Karatzas |
ICDAR | 9 |
| 2019 | ICDAR2019 Robust Reading Challenge on Arbitrary-Shaped Text - RRC-ArTabstractThis paper reports the ICDAR2019 Robust Reading Challenge on Arbitrary-Shaped Text - RRC-ArT that consists of three major challenges: i) scene text detection, ii) scene text recognition, and iii) scene text spotting. A total of 78 submissions from 46 unique teams/individuals were received for this competition. The top performing score of each challenge is as follows: i) T1 - 82.65%, ii) T2.1 - 74.3%, iii) T2.2 - 85.32%, iv) T3.1 - 53.86%, and v) T3.2 - 54.91%. Apart from the results, this paper also details the ArT dataset, tasks description, evaluation metrics and participants' methods. The dataset, the evaluation kit as well as the results are publicly available at the challenge website. Chee Kheng Chng, Errui Ding, Jingtuo Liu, Dimosthenis Karatzas, Chee Seng Chan, Yipeng Sun, Chun Chet Ng, Canjie Luo, Zihan Ni, ChuanMing Fang, Shuaitao Zhang, Junyu Han |
ICDAR | 4 |
| 2019 | Selective Style Transfer for TextabstractThis paper explores the possibilities of image style transfer applied to text maintaining the original transcriptions. Results on different text domains (scene text, machine printed text and handwritten text) and cross-modal results demonstrate that this is feasible, and open different research lines. Furthermore, two architectures for selective style transfer, which means transferring style to only desired image pixels, are proposed. Finally, scene text selective style transfer is evaluated as a data augmentation technique to expand scene text detection datasets, resulting in a boost of text detectors performance. Our implementation of the described models is publicly available. Raul Gomez, Ali Furkan Biten, Lluís Gómez i Bigorda, Jaume Gibert, Dimosthenis Karatzas, Marçal Rusiñol |
ICDAR | 5 |
| 2019 | ICDAR2019 Competition on Scanned Receipt OCR and Information ExtractionabstractThe ICDAR 2019 Challenge on "Scanned receipts OCR and key information extraction" (SROIE) covers important aspects related to the automated analysis of scanned receipts. The SROIE tasks play a key role in many document analysis systems and hold significant commercial potential. Although a lot of work has been published over the years on administrative document analysis, the community has advanced relatively slowly, as most datasets have been kept private. One of the key contributions of SROIE to the document analysis community is to offer a first, standardized dataset of 1000 whole scanned receipt images and annotations, as well as an evaluation procedure for such tasks. The Challenge is structured around three tasks, namely Scanned Receipt Text Localization (Task 1), Scanned Receipt OCR (Task 2) and Key Information Extraction from Scanned Receipts (Task 3). The competition opened on 10th February, 2019 and closed on 5th May, 2019. We received 29, 24 and 18 valid submissions received for the three competition tasks, respectively. This report presents the competition datasets, define the tasks and the evaluation protocols, offer detailed submission statistics, as well as an analysis of the submitted performance. While the tasks of text localization and recognition seem to be relatively easy to tackle, it is interesting to observe the variety of ideas and approaches proposed for the information extraction task. According to the submissions' performance we believe there is still margin for improving information extraction performance, although the current dataset would have to grow substantially in following editions. Given the success of the SROIE competition evidenced by the wide interest generated and the healthy number of submissions from academic, research institutes and industry over different countries, we consider that the SROIE competition can evolve into a useful resource for the community, drawing further attention and promoting research and development efforts in this field. Kai Chen 0006, Jianhua He 0001, Xiang Bai, Dimosthenis Karatzas, Shijian Lu, C. V. Jawahar |
ICDAR | 5 |
| 2019 | ICDAR2019 Robust Reading Challenge on Multi-lingual Scene Text Detection and Recognition - RRC-MLT-2019abstractWith the growing cosmopolitan culture of modern cities, the need of robust Multi-Lingual scene Text (MLT) detection and recognition systems has never been more immense. With the goal to systematically benchmark and push the state-of-the-art forward, the proposed competition builds on top of the RRC-MLT-2017 with an additional end-to-end task, an additional language in the real images dataset, a large scale multi-lingual synthetic dataset to assist the training, and a baseline End-to-End recognition method. The real dataset consists of 20,000 images containing text from 10 languages. The challenge has 4 tasks covering various aspects of multi-lingual scene text: (a) text detection, (b) cropped word script classification, (c) joint text detection and script classification and (d) end-to-end detection and recognition. In total, the competition received 60 submissions from the research and industrial communities. This paper presents the dataset, the tasks and the findings of the presented RRC-MLT-2019 challenge. Nibal Nayef, Cheng-Lin Liu 0001, Jean-Marc Ogier, Michal Busta, Pinaki Nath Chowdhury, Dimosthenis Karatzas, Wafa Khlif, Jiri Matas, Umapada Pal 0001, Jean-Christophe Burie |
ICDAR | 7 |
| 2019 | ICDAR 2019 Competition on Large-Scale Street View Text with Partial Labeling - RRC-LSVTabstractRobust text reading from street view images provides valuable information for various applications. Performance improvement of existing methods in such a challenging scenario heavily relies on the amount of fully annotated training data, which is costly and in-efficient to obtain. To scale up the amount of training data while keeping the labeling procedure cost-effective, this competition introduces a new challenge on Large-scale Street View Text with Partial Labeling (LSVT), providing 5,0000 and 400,000 images in full and weak annotations, respectively. This competition aims to explore the abilities of state-of-the-art methods to detect and recognize text instances from large-scale street view images, closing gaps between research benchmarks and real applications. During the competition period, a total number of 41 teams participate in the two tasks with 132 valid submissions, i.e., text detection and end-to-end text spotting. This paper includes dataset descriptions, task definitions, evaluation protocols and results summaries of ICDAR 2019-LSVT challenge. Yipeng Sun, Dimosthenis Karatzas, Chee Seng Chan, Zihan Ni, Chee Kheng Chng, Canjie Luo, Chun Chet Ng, Junyu Han, Errui Ding, Jingtuo Liu |
ICDAR | 2 |
| 2019 | ICDAR 2019 Robust Reading Challenge on Reading Chinese Text on SignboardabstractChinese scene text reading is one of the most challenging problems in computer vision and has attracted great interest. Different from English text, Chinese has more than 6000 commonly used characters and Chinese characters can be arranged in various layouts with numerous fonts. The Chinese signboards in street view are a good choice for Chinese scene text images since they have different backgrounds, fonts and layouts. We organized a competition called ICDAR2019-ReCTS, which mainly focuses on reading Chinese text on signboard. This report presents the final results of the competition. A large-scale dataset of 25,000 annotated signboard images, in which all the text lines and characters are annotated with locations and transcriptions, were released. Four tasks, namely character recognition, text line recognition, text line detection and end-to-end recognition were set up. Besides, considering the Chinese text ambiguity issue, we proposed a multi ground truth (multi-GT) evaluation method to make evaluation fairer. The competition started on March 1, 2019 and ended on April 30, 2019. 262 submissions from 46 teams are received. Most of the participants come from universities, research institutes, and tech companies in China. There are also some participants from the United States, Australia, Singapore, and Korea. 21 teams submit results for Task 1, 23 teams submit results for Task 2, 24 teams submit results for Task 3, and 13 teams submit results for Task 4. The official website for the competition is http://rrc.cvc.uab.es/?ch=12. Rui Zhang 0056, Xiang Bai, Baoguang Shi, Dimosthenis Karatzas, Shijian Lu, C. V. Jawahar, Yongsheng Zhou, Qianyi Jiang, Nan Li 0071, Dong Wang 0004, Minghui Liao |
ICDAR | 5 |
| 2019 | Self-Supervised Visual Representations for Cross-Modal RetrievalabstractCross-modal retrieval methods have been significantly improved in last years with the use of deep neural networks and large-scale annotated datasets such as ImageNet and Places. However, collecting and annotating such datasets requires a tremendous amount of human effort and, besides, their annotations are limited to discrete sets of popular visual classes that may not be representative of the richer semantics found on large-scale cross-modal retrieval datasets. In this paper, we present a self-supervised cross-modal retrieval framework that leverages as training data the correlations between images and text on the entire set of Wikipedia articles. Our method consists in training a CNN to predict: (1) the semantic context of the article in which an image is more probable to appear as an illustration, and (2) the semantic context of its caption. Our experiments demonstrate that the proposed method is not only capable of learning discriminative visual representations for solving vision tasks like classification, but that the learned representations are better for cross-modal retrieval when compared to supervised pre-training of the network on the ImageNet dataset. Lluís Gómez i Bigorda, Marçal Rusiñol, Dimosthenis Karatzas, C. V. Jawahar |
ICMR | 4 |
| 2018 | Cutting Sayre's Knot: Reading Scene Text without Segmentation. Application to Utility MetersabstractIn this paper we present a segmentation-free system for reading text in natural scenes. A CNN architecture is trained in an end-to-end manner, and is able to directly output readings without any explicit text localization step. In order to validate our proposal, we focus on the specific case of reading utility meters. We present our results in a large dataset of images acquired by different users and devices, so text appears in any location, with different sizes, fonts and lengths, and the images present several distortions such as dirt, illumination highlights or blur. Lluís Gómez i Bigorda, Marçal Rusiñol, Dimosthenis Karatzas |
DAS | 3 |
| 2018 | The Robust Reading Competition Annotation and Evaluation PlatformabstractThe ICDAR Robust Reading Competition (RRC), initiated in 2003 and re-established in 2011, has become a de-facto evaluation standard for robust reading systems and algorithms. Concurrent with its second incarnation in 2011, a continuous effort started to develop an on-line framework to facilitate the hosting and management of competitions. This paper outlines the Robust Reading Competition Annotation and Evaluation Platform, the backbone of the competitions. The RRC Annotation and Evaluation Platform is a modular framework, fully accessible through on-line interfaces. It comprises a collection of tools and services for managing all processes involved with defining and evaluating a research task, from dataset definition to annotation management, evaluation specification and results analysis. Although the framework has been designed with robust reading research in mind, many of the provided tools are generic by design. All aspects of the RRC Annotation and Evaluation Framework are available for research use. Dimosthenis Karatzas, Lluís Gómez i Bigorda, Anguelos Nicolaou, Marçal Rusiñol |
DAS | 1 |
| 2017 | LSDE: Levenshtein Space Deep Embedding for Query-by-String Word SpottingabstractIn this paper we present the LSDE string representation and its application to handwritten word spotting. LSDE is a novel embedding approach for representing strings that learns a space in which distances between projected points are correlated with the Levenshtein edit distance between the original strings. We show how such a representation produces a more semantically interpretable retrieval from the user's perspective than other state of the art ones such as PHOC and DCToW. We also conduct a preliminary handwritten word spotting experiment on the George Washington dataset. Lluís Gómez i Bigorda, Marçal Rusiñol, Dimosthenis Karatzas |
ICDAR | 3 |
| 2017 | ICDAR2017 Robust Reading Challenge on COCO-TextabstractThis report presents the final results of the ICDAR 2017 Robust Reading Challenge on COCO-Text. A challenge on scene text detection and recognition based on the largest real scene text dataset currently available: the COCO-Text dataset. The competition is structured around three tasks: Text Localization, Cropped Word Recognition and End-To-End Recognition. The competition received a total of 27 submissions over the different opened tasks. This report describes the datasets and the ground truth, details the performance evaluation protocols used and presents the final results along with a brief summary of the participating methods. Raul Gomez, Baoguang Shi, Lluís Gómez i Bigorda, Lukás Neumann, Andreas Veit, Jiri Matas, Serge J. Belongie, Dimosthenis Karatzas |
ICDAR | 8 |
| 2017 | ICDAR2017 Robust Reading Challenge on Omnidirectional VideoabstractResults of ICDAR 2017 Robust Reading Challenge on Omnidirectional Video are presented. This competition uses Downtown Osaka Scene Text (DOST) Dataset that was captured in Osaka, Japan with an omnidirectional camera. Hence, it consists of sequential images (videos) of different view angles. Regarding the sequential images as videos (video mode), two tasks of localisation and end-to-end recognition are prepared. Regarding them as a set of still images (still image mode), three tasks of localisation, cropped word recognition and end-to-end recognition are prepared. As the dataset has been captured in Japan, the dataset contains Japanese text but also include text consisting of alphanumeric characters (Latin text). Hence, a submitted result for each task is evaluated in three ways: using Japanese only ground truth (GT), using Latin only GT and using combined GTs of both. Finally, by the submission deadline, we have received two submissions in the text localisation task of the still image mode. We intend to continue the competition in the open mode. Expecting further submissions, in this report we provide baseline results in all the tasks in addition to the submissions from the community. Masakazu Iwamura, Naoyuki Morimoto, Keishi Tainaka, Dena Bazazian, Lluís Gómez i Bigorda, Dimosthenis Karatzas |
ICDAR | 6 |
| 2017 | ICDAR2017 Robust Reading Challenge on Multi-Lingual Scene Text Detection and Script Identification - RRC-MLTabstractText detection and recognition in a natural environment are key components of many applications, ranging from business card digitization to shop indexation in a street. This competition aims at assessing the ability of state-of-the-art methods to detect Multi-Lingual Text (MLT) in scene images, such as in contents gathered from the Internet media and in modern cities where multiple cultures live and communicate together. This competition is an extension of the Robust Reading Competition (RRC) which has been held since 2003 both in ICDAR and in an online context. The proposed competition is presented as a new challenge of the RRC. The dataset built for this challenge largely extends the previous RRC editions in many aspects: the multi-lingual text, the size of the dataset, the multi-oriented text, the wide variety of scenes. The dataset is comprised of 18,000 images which contain text belonging to 9 languages. The challenge is comprised of three tasks related to text detection and script classification. We have received a total of 16 participations from the research and industrial communities. This paper presents the dataset, the tasks and the findings of this RRC-MLT challenge. Nibal Nayef, Imen Bizid, Hyunsoo Choi, Dimosthenis Karatzas, Zhenbo Luo, Umapada Pal 0001, Christophe Rigaud, Joseph Chazalon, Wafa Khlif, Muhammad Muzzamil Luqman, Jean-Christophe Burie, Cheng-Lin Liu 0001, Jean-Marc Ogier |
ICDAR | 6 |
| 2017 | ICDAR2017 Robust Reading Challenge on Text Extraction from Biomedical Literature Figures (DeTEXT)abstractHundreds of millions of figures are available in the biomedical literature, representing important biomedical experimental evidence. Since text is a rich source of information in figures, automatically extracting such text may assist in the task of mining figure information and understanding biomedical documents. Unlike images in the open domain, biomedical figures present a variety of unique challenges. For example, biomedical figures typically have complex layouts, small font sizes, short text, specific text, complex symbols and irregular text arrangements. This paper presents the final results of the ICDAR 2017 Competition on Text Extraction from Biomedical Literature Figures (ICDAR2017 DeTEXT Competition), which aims at extracting (detecting and recognizing) text from biomedical literature figures. Similar to text extraction from scene images and web pictures, ICDAR2017 DeTEXT Competition includes three major tasks, i.e., text detection, cropped word recognition and end-to-end text recognition. Here, we describe in detail the data set, tasks, evaluation protocols and participants of this competition, and report the performance of the participating methods. Xu-Cheng Yin, Dimosthenis Karatzas |
ICDAR | 4 |
| 2016 | A Fine-Grained Approach to Scene Text Script IdentificationabstractThis paper focuses on the problem of script identification in unconstrained scenarios. Script identification is an important prerequisite to recognition, and an indispensable condition for automatic text understanding systems designed for multi-language environments. Although widely studied for document images and handwritten documents, it remains an almost unexplored territory for scene text images. We detail a novel method for script identification in natural images that combines convolutional features and the Naive-Bayes Nearest Neighbor classifier. The proposed framework efficiently exploits the discriminative power of small stroke-parts, in a fine-grained classification framework. In addition, we propose a new public benchmark dataset for the evaluation of joint text detection and script identification in natural scenes. Experiments done in this new dataset demonstrate that the proposed method yields state of the art results, while it generalizes well to different datasets and variable number of scripts. The evidence provided shows that multi-lingual scene text recognition in the wild is a viable proposition. Source code of the proposed method is made available online. Lluís Gómez i Bigorda, Dimosthenis Karatzas |
DAS | 2 |
| 2016 | Human-Document Interaction Systems - A New Frontier for Document Image AnalysisabstractAll indications show that paper documents will not cede in favour of their digital counterparts, but will instead be used increasingly in conjunction with digital information. An open challenge is how to seamlessly link the physical with the digital -- how to continue taking advantage of the important affordances of paper, without missing out on digital functionality. This paper presents the authors' experience with developing systems for Human-Document Interaction based on augmented document interfaces and examines new challenges and opportunities arising for the document image analysis field in this area. The system presented combines state of the art camera-based document image analysis techniques with a range of complementary technologies to offer fluid Human-Document Interaction. Both fixed and nomadic setups are discussed that have gone through user testing in real-life environments, and use cases are presented that span the spectrum from business to educational applications. Dimosthenis Karatzas, Vincent Poulain D'Andecy, Marçal Rusiñol, Antoni Chica, Pere-Pau Vázquez |
DAS | 1 |
| 2016 | Visual Script and Language IdentificationabstractIn this paper we introduce a script identification method based on hand-crafted texture features and an artificial neural network. The proposed pipeline achieves near state-of-the-art performance for script identification of video-text and state-of-the-art performance on visual language identification of handwritten text. More than using the deep network as a classifier, the use of its intermediary activations as a learned metric demonstrates remarkable results and allows the use of discriminative models on unknown classes. Comparative experiments in video-text and text in the wild datasets provide insights on the internals of the proposed deep network. Anguelos Nicolaou, Andrew D. Bagdanov, Lluís Gómez i Bigorda, Dimosthenis Karatzas |
DAS | 4 |
| 2015 | Object proposals for text extraction in the wildabstractObject Proposals is a recent computer vision technique receiving increasing interest from the research community. Its main objective is to generate a relatively small set of bounding box proposals that are most likely to contain objects of interest. The use of Object Proposals techniques in the scene text understanding field is innovative. Motivated by the success of powerful while expensive techniques to recognize words in a holistic way, Object Proposals techniques emerge as an alternative to the traditional text detectors. In this paper we study to what extent the existing generic Object Proposals methods may be useful for scene text understanding. Also, we propose a new Object Proposals algorithm that is specifically designed for text and compare it with other generic methods in the state of the art. Experiments show that our proposal is superior in its ability of producing good quality word proposals in an efficient way. The source code of our method is made publicly available1. Lluís Gómez i Bigorda, Dimosthenis Karatzas |
ICDAR | 2 |
| 2015 | Novel line verification for multiple instance focused retrieval in document collectionsabstractSpatial verification is typically employed to check the spatial consistency among matched local features and to remove outliers. However, when looking for multiple instances of the query within a target image, RANSAC algorithms which are widely applied in many one-to-one matching applications might fail due to the large proportion of “outliers” - correct matches corresponding to other instances. On the other hand, geometrical verification methods are more robust to outliers but usually suffer from high computational costs. In this paper, we introduce a novel two-step line verification method which is more flexible than existing methods and leads to lower computational complexity especially when multiple instances of a query are sought. We study this approach within an information extraction scenario, where the objective is to locate document structures indicative of certain type of information (e.g. different records on invoices). Hongxing Gao, Marçal Rusiñol, Dimosthenis Karatzas, Josep Lladós 0001, Rajiv Jain, David S. Doermann |
ICDAR | 3 |
| 2015 | Efficient indexing for Query By String text retrievalabstractThis paper deals with Query By String word spotting in scene images. A hierarchical text segmentation algorithm based on text specific selective search is used to find text regions. These regions are indexed per character n-grams present in the text region. An attribute representation based on Pyramidal Histogram of Characters (PHOC) is used to compare text regions with the query text. For generation of the index a similar attribute space based Pyramidal Histogram of character n-grams is used. These attribute models are learned using linear SVMs over the Fisher Vector [1] representation of the images along with the PHOC labels of the corresponding strings. Suman K. Ghosh, Lluís Gómez i Bigorda, Dimosthenis Karatzas, Ernest Valveny |
ICDAR | 3 |
| 2015 | ICDAR 2015 competition on Robust ReadingabstractResults of the ICDAR 2015 Robust Reading Competition are presented. A new Challenge 4 on Incidental Scene Text has been added to the Challenges on Born-Digital Images, Focused Scene Images and Video Text. Challenge 4 is run on a newly acquired dataset of 1,670 images evaluating Text Localisation, Word Recognition and End-to-End pipelines. In addition, the dataset for Challenge 3 on Video Text has been substantially updated with more video sequences and more accurate ground truth data. Finally, tasks assessing End-to-End system performance have been introduced to all Challenges. The competition took place in the first quarter of 2015, and received a total of 44 submissions. Only the tasks newly introduced in 2015 are reported on. The datasets, the ground truth specification and the evaluation protocols are presented together with the results and a brief summary of the participating methods. Dimosthenis Karatzas, Lluís Gómez i Bigorda, Anguelos Nicolaou, Suman K. Ghosh, Andrew D. Bagdanov, Masakazu Iwamura, Jiri Matas, Lukás Neumann, Vijay Chandrasekhar 0001, Shijian Lu, Faisal Shafait, Seiichi Uchida, Ernest Valveny |
ICDAR | 1 |
| 2015 | Sparse radial sampling LBP for writer identificationabstractSampling Local Binary Patterns, a variant of Local Binary Patterns (LBP) for text-as-texture classification. By adapting and extending the standard LBP operator to the particularities of text we get a generic text-as-texture classification scheme and apply it to writer identification. In experiments on CVL and ICDAR 2013 datasets, the proposed feature-set and a simple end-to-end pipeline demonstrate State-Of-the-Art (SOA) performance. Among the SOA, the proposed method is the only one that is based on dense extraction of a single local feature descriptor. This makes it fast and applicable at the earliest stages in a DIA pipeline without the need for segmentation, binarization, or extraction of multiple features. Anguelos Nicolaou, Andrew D. Bagdanov, Marcus Liwicki, Dimosthenis Karatzas |
ICDAR | 4 |
| 2014 | A Cache Language Model for Whole Document Handwriting RecognitionabstractWith increasing computational power, the trend in unconstrained text recognition is going towards whole document processing. For this task, more sophisticated language models can be employed. One approach is to take advantage the fact that the text of a document normally deals with a specific topic and hence the word occurrence probability is biased. Cache language models combine information about recent words, the cache, with a general statistical language model to increase the recognition rate. In this work we introduce a modified version of the cache language model to the task of handwriting recognition, where the N-best recognition output of the entire document is used to refine the language model for a consecutive recognition pass. An experimental evaluation on the IAM database demonstrates that the word error rate can be reduced with the proposed cache language model. Volkmar Frinken, Dimosthenis Karatzas, Andreas Fischer 0002 |
Document Analysis Systems | 2 |
| 2014 | An On-line Platform for Ground Truthing and Performance Evaluation of Text Extraction SystemsabstractThis work presents a set of on-line software tools for creating ground truth and calculating performance evaluation metrics for text extraction tasks such as localization, segmentation and recognition. The platform supports the definition of comprehensive ground truth information at different text representation levels while it offers centralised management and quality control of the ground truthing effort. It implements a range of state of the art performance evaluation algorithms and offers functionality for the definition of evaluation scenarios, on-line calculation of various performance metrics and visualisation of the results. The presented platform, which comprises the backbone of the ICDAR 2011 (challenge 1) and 2013 (challenges 1 and 2) Robust Reading competitions, is now made available for public use. Dimosthenis Karatzas, Sergi Robles, Lluís Gómez i Bigorda |
Document Analysis Systems | 1 |
| 2014 | Color Descriptor for Content-Based Drawing RetrievalabstractHuman detection in computer vision field is an active field of research. Extending this to human-like drawings such as the main characters in comic book stories is not trivial. Comics analysis is a very recent field of research at the intersection of graphics, texts, objects and people recognition. The detection of the main comic characters is an essential step towards a fully automatic comic book understanding. This paper presents a color-based approach for comics character retrieval using content-based drawing retrieval and color palette. Christophe Rigaud, Dimosthenis Karatzas, Jean-Christophe Burie, Jean-Marc Ogier |
Document Analysis Systems | 2 |
| 2013 | Key-Region Detection for Document Images - Application to Administrative Document RetrievalabstractIn this paper we argue that a key-region detector designed to take into account the special characteristics of document images can result in the detection of less and more meaningful key-regions. We propose a fast key-region detector able to capture aspects of the structural information of the document, and demonstrate its efficiency by comparing against standard detectors in an administrative document retrieval scenario. We show that using the proposed detector results to a smaller number of detected key-regions and higher performance without any drop in speed compared to standard state of the art detectors. Hongxing Gao, Marçal Rusiñol, Dimosthenis Karatzas, Josep Lladós 0001, Tomokazu Sato, Masakazu Iwamura, Koichi Kise |
ICDAR | 3 |
| 2013 | Multi-script Text Extraction from Natural ScenesabstractScene text extraction methodologies are usually based in classification of individual regions or patches, using a priori knowledge for a given script or language. Human perception of text, on the other hand, is based on perceptual organisation through which text emerges as a perceptually significant group of atomic objects. Therefore humans are able to detect text even in languages and scripts never seen before. In this paper, we argue that the text extraction problem could be posed as the detection of meaningful groups of regions. We present a method built around a perceptual organisation framework that exploits collaboration of proximity and similarity laws to create text-group hypotheses. Experiments demonstrate that our algorithm is competitive with state of the art approaches on a standard dataset covering text in variable orientations and two languages. Lluís Gómez i Bigorda, Dimosthenis Karatzas |
ICDAR | 2 |
| 2013 | Document Classification and Page Stream Segmentation for Digital Mailroom ApplicationsabstractIn this paper we present a method for the segmentation of continuous page streams into multipage documents and the simultaneous classification of the resulting documents. We first present an approach to combine the multiple pages of a document into a single feature vector that represents the whole document. Despite its simplicity and low computational cost, the proposed representation yields results comparable to more complex methods in multipage document classification tasks. We then exploit this representation in the context of page stream segmentation. The most plausible segmentation of a page stream into a sequence of multipage documents is obtained by optimizing a statistical model that represents the probability of each segmented multipage document belonging to a particular class. Experimental results are reported on a large sample of real administrative multipage documents. Albert Gordo, Marçal Rusiñol, Dimosthenis Karatzas, Andrew D. Bagdanov |
ICDAR | 3 |
| 2013 | ICDAR 2013 Robust Reading CompetitionabstractThis report presents the final results of the ICDAR 2013 Robust Reading Competition. The competition is structured in three Challenges addressing text extraction in different application domains, namely born-digital images, real scene images and real-scene videos. The Challenges are organised around specific tasks covering text localisation, text segmentation and word recognition. The competition took place in the first quarter of 2013, and received a total of 42 submissions over the different tasks offered. This report describes the datasets and ground truth specification, details the performance evaluation protocols used and presents the final results along with a brief summary of the participating methods. Dimosthenis Karatzas, Faisal Shafait, Seiichi Uchida, Masakazu Iwamura, Lluís Gómez i Bigorda, Sergi Robles, Joan Mas Romeu, David Fernández Mota, Jon Almazán, Lluís-Pere de las Heras |
ICDAR | 1 |
| 2013 | An Active Contour Model for Speech Balloon Detection in ComicsabstractComic books constitute an important cultural heritage asset in many countries. Digitization combined with subsequent comic book understanding would enable a variety of new applications, including content-based retrieval and content retargeting. Document understanding in this domain is challenging as comics are semi-structured documents, combining semantically important graphical and textual parts. Few studies have been done in this direction. In this work we detail a novel approach for closed and non-closed speech balloon localization in scanned comic book pages, an essential step towards a fully automatic comic book understanding. The approach is compared with existing methods for closed balloon localization found in the literature and results are presented. Christophe Rigaud, Jean-Christophe Burie, Jean-Marc Ogier, Dimosthenis Karatzas, Joost van de Weijer 0001 |
ICDAR | 4 |
| 2011 | Interactive Trademark Image Retrieval by Fusing Semantic and Visual Content
Marçal Rusiñol, David Aldavert, Dimosthenis Karatzas, Ricardo Toledo, Josep Lladós 0001 |
ECIR | 3 |
| 2011 | ICDAR 2011 Robust Reading Competition - Challenge 1: Reading Text in Born-Digital Images (Web and Email)abstractThis paper presents the results of the first Challenge of ICDAR 2011 Robust Reading Competition. Challenge 1 is focused on the extraction of text from born-digital images, specifically from images found in Web pages and emails. The challenge was organized in terms of three tasks that look at different stages of the process: text localization, text segmentation and word recognition. In this paper we present the results of the challenge for all three tasks, and make an open call for continuous participation outside the context of ICDAR 2011. Dimosthenis Karatzas, Sergi Robles, Joan Mas Romeu, Farshad Nourbakhsh, Partha Pratim Roy 0001 |
ICDAR | 1 |
| 2010 | A framework for the assessment of text extraction algorithms on complex colour imagesabstractThe availability of open, ground-truthed datasets and clear performance metrics is a crucial factor in the development of an application domain. The domain of colour text image analysis (real scenes, Web and spam images, scanned colour documents) has traditionally suffered from a lack of a comprehensive performance evaluation framework. Such a framework is extremely difficult to specify, and corresponding pixel-level accurate information tedious to define. In this paper we discuss the challenges and technical issues associated with developing such a framework. Then, we describe a complete framework for the evaluation of text extraction methods at multiple levels, provide a detailed ground-truth specification and present a case study on how this framework can be used in a real-life situation. Antonio Clavelli, Dimosthenis Karatzas, Josep Lladós 0001 |
Document Analysis Systems | 2 |
| 2010 | A polar-based logo representation based on topological and colour featuresabstractIn this paper, we propose a novel rotation and scale invariant method for colour logo retrieval and classification, which involves performing a simple colour segmentation and subsequently describing each of the resultant colour components based on a set of topological and colour features. A polar representation is used to represent the logo and the subsequent logo matching is based on Cyclic Dynamic Time Warping (CDTW). We also show how combining information about the global distribution of the logo components and their local neighbourhood using the Delaunay triangulation allows to improve the results. All experiments are performed on a dataset of 2500 instances of 100 colour logo images in different rotations and scales. Farshad Nourbakhsh, Dimosthenis Karatzas, Ernest Valveny |
Document Analysis Systems | 2 |
| 2009 | Text Segmentation in Colour Posters from the Spanish Civil War EraabstractThe extraction of textual content from colour documents of a graphical nature is a complicated task. The text can be rendered in any colour, size and orientation while the existence of complex background graphics with repetitive patterns can make its localization and segmentation extremely difficult. Here, we propose a new method for extracting textual content from such colour images that makes no assumption as to the size of the characters, their orientation or colour, while it is tolerant to characters that do not follow a straight baseline. We evaluate this method on a collection of documents with historical connotations: the posters from the Spanish Civil War. Antonio Clavelli, Dimosthenis Karatzas |
ICDAR | 2 |
| 2008 | Detecting Gradients in Text Images Using the Hough TransformabstractThe use of gradients in text images is nowadays quite frequent. Existing segmentation methods encounter serious problems when it comes to modern text images where gradients might appear in the background or the foreground or both at the same time. This paper presents an approach for lightness gradient areas detection based on the Hough Transform. The issues arising are discussed, and results are presented on a dataset comprising Web images, logos and scanned documents. Dimosthenis Karatzas |
Document Analysis Systems | 1 |
| 2008 | HistoSketch: A Semi-Automatic Annotation Tool for Archival DocumentsabstractThis article describes a sketch-based framework for semi-automatic annotation of historical document collections. It is motivated by the fact that fully automatic methods, while helpful for extracting metadata from large collections, have two main drawbacks in a real-world application: (i) they are error-prone and (ii) they only capture a subset of all the knowledge in the document base, both meaning that manual intervention is always required. Therefore, we have developed a practical framework for allowing experts to extract knowledge from document collections in a sketch-based scenario. The main possibilities of the proposed framework are: (a) browsing the collection efficiently, (b) providing gestures for metadata input, (c) supporting handwritten notes and (d) providing gestures for launching automatic extraction processes such as OCR or word spotting. Joan Mas Romeu, José A. Rodríguez 0001, Dimosthenis Karatzas, Gemma Sánchez, Josep Lladós 0001 |
Document Analysis Systems | 3 |
| 2006 | Ground Truth for Layout Analysis Performance Evaluation
Apostolos Antonacopoulos, Dimosthenis Karatzas, David Bridson |
Document Analysis Systems | 2 |
| 2005 | Semantics-Based Content Extraction in Typewritten Historical DocumentsabstractThis paper presents a flexible approach to extracting content from scanned historical documents using semantic information. The final electronic document is the result of a "digital historical document lifecycle" process, where the expert knowledge of the historian/archivist user is incorporated at different stages. Results show that such a conversion strategy aided by (expert) user-specified semantic information and which enables the processing of individual parts of the document in a specialised way, produces superior (in a variety of significant ways) results than document analysis and understanding techniques devised for contemporary documents. Apostolos Antonacopoulos, Dimosthenis Karatzas |
ICDAR | 2 |
| 2004 | A Complete Approach to the Conversion of Typewritten Historical Documents for Digital Archives
Apostolos Antonacopoulos, Dimosthenis Karatzas |
Document Analysis Systems | 2 |
| 2004 | The lifecycle of a digital historical document: structure and contentabstractThis paper describes the lifecycle of a digital historical document, from template-based structure definition through to content extraction from the scanned pages and its final reconstitution as an electronic document (combining content and semantic information) along with the tools that have been created to realise each stage in the lifecycle. The whole approach is described in the context of different types of typewritten documents relating to prisoners in World-War II concentration camps and is the result of a multinational collaboration under the MEMORIAL project funded (€1.5M) by the European Union (www.memorial-project.info). Extensive tests with historians/archivists and evaluation of the content extraction results indicate the superior performance of the whole semantics-driven approach both over manual transcription and over the semi-automated application of off-the-shelf OCR and the use of a conventional (text and layout) document format. Apostolos Antonacopoulos, Dimosthenis Karatzas, Henryk Krawczyk, Bogdan Wiszniewski |
ACM Symposium on Document Engineering | 2 |
| 2003 | ICDAR 2003 Page Segmentation Competition
Apostolos Antonacopoulos, Basilios Gatos, Dimosthenis Karatzas |
ICDAR | 3 |
| 2003 | Two Approaches for Text Segmentation in Web ImagesabstractThere is a significant need to recognise the text in images on Web pages, both for effective indexing and for presentation by non-visual means (e.g., audio). This paper presents and compares two novel methods for the segmentation of characters for subsequent extraction and recognition. The novelty of both approaches is the combination of (different in each case) topological features of characters with an anthropocentric perspective of colour perception - in preference to RGB space analysis. Both approaches enable the extraction of text in complex situations such as in the presence of varying colour and texture (characters and background). Dimosthenis Karatzas, Apostolos Antonacopoulos |
ICDAR | 1 |
| 2002 | Fuzzy Segmentation of Characters in Web Images Based on Human Colour Perception
Apostolos Antonacopoulos, Dimosthenis Karatzas |
Document Analysis Systems | 2 |