Ernest Valveny

dblp:78/6032 · also Ernest Valveny Llobet · DBLP profile ↗
← Back
36ranked-venue papers in the field
4as first author
9since 2021 · last 2025
0000-0002-0368-9697ORCID · verified

Domains — venue-derived; a paper can count in several

Other / Interdisciplinary · 36 (4 first)
YearPublicationVenuePosition
2025 LLM-Driven Medical Document Analysis: Enhancing Trustworthy Pathology and Differential Diagnosis
Lei Kang 0002, Xuanshuo Fu, Oriol Ramos Terrades, Javier Vazquez-Corral, Ernest Valveny, Dimosthenis Karatzas
ICDAR (3)5
2025 ComicsPAP: Understanding Comic Strips by Picking the Correct Panel
Emanuele Vivoli, Artemis Llabrés, Mohamed Ali Souibgui, Marco Bertini 0001, Ernest Valveny, Dimosthenis Karatzas
ICDAR (1)5
2024 Image-Text Matching for Large-Scale Book Collections
Artemis Llabrés, Arka Ujjal Dey, Dimosthenis Karatzas, Ernest Valveny
DAS4
2024 Machine Unlearning for Document Classification
Lei Kang 0002, Mohamed Ali Souibgui, Fei Yang 0004, Lluís Gómez i Bigorda, Ernest Valveny, Dimosthenis Karatzas
ICDAR (4)5
2024 Multi-page Document Visual Question Answering Using Self-attention Scoring Mechanism
Lei Kang 0002, Rubèn Tito, Ernest Valveny, Dimosthenis Karatzas
ICDAR (6)3
2024 Privacy-Aware Document Visual Question Answering
Rubèn Tito, Marlon Tobaben, Raouf Kerkouche, Mohamed Ali Souibgui, Kangsoo Jung, Joonas Jälkö, Vincent Poulain D'Andecy, Aurélie Joseph, Lei Kang 0002, Ernest Valveny, Antti Honkela, Mario Fritz, Dimosthenis Karatzas
ICDAR (6)11
2024 Multimodal Transformer for Comics Text-Cloze
Emanuele Vivoli, Joan Lafuente Baeza, Ernest Valveny, Dimosthenis Karatzas
ICDAR (6)3
2021 Document Collection Visual Question Answering
Rubèn Tito, Dimosthenis Karatzas, Ernest Valveny
ICDAR (2)3
2021 ICDAR 2021 Competition on Document Visual Question Answering
Rubèn Tito, Minesh Mathew, C. V. Jawahar, Ernest Valveny, Dimosthenis Karatzas
ICDAR (4)4
2019 Can One Deep Learning Model Learn Script-Independent Multilingual Word-Spotting?
abstract
Word spotting has gained increased attention lately as it can be used to extract textual information from handwritten documents and scene-text images. Current word spotting approaches are designed to work on a single language and/or script. Building intelligent models that learn script-independent multilingual word-spotting is challenging due to the large variability of multilingual alphabets and symbols. We used ResNet-152 and the Pyramidal Histogram of Characters (PHOC) embedding to build a one-model script-independent multilingual word-spotting and we tested it on Latin, Arabic, and Bangla (Indian) languages. The one-model we propose performs on par with the multi-model language-specific word-spotting system, and thus, reduces the number of models needed for each script and/or language.
Mohammed Al-Rawi, Ernest Valveny, Dimosthenis Karatzas
ICDAR2
2019 ICDAR 2019 Competition on Scene Text Visual Question Answering
abstract
This paper presents final results of ICDAR 2019 Scene Text Visual Question Answering competition (ST-VQA). ST-VQA introduces an important aspect that is not addressed by any Visual Question Answering system up to date, namely the incorporation of scene text to answer questions asked about an image. The competition introduces a new dataset comprising 23,038 images annotated with 31,791 question / answer pairs where the answer is always grounded on text instances present in the image. The images are taken from 7 different public computer vision datasets, covering a wide range of scenarios. The competition was structured in three tasks of increasing difficulty, that require reading the text in a scene and understanding it in the context of the scene, to correctly answer a given question. A novel evaluation metric is presented, which elegantly assesses both key capabilities expected from an optimal model: text recognition and image understanding. A detailed analysis of results from different participants is showcased, which provides insight into the current capabilities of VQA systems that can read. We firmly believe the dataset proposed in this challenge will be an important milestone to consider towards a path of more robust and general models that can exploit scene text to achieve holistic image understanding.
Ali Furkan Biten, Rubèn Tito, Andrés Mafla, Lluís Gómez i Bigorda, Marçal Rusiñol, Minesh Mathew, C. V. Jawahar, Ernest Valveny, Dimosthenis Karatzas
ICDAR8
2017 R-PHOC: Segmentation-Free Word Spotting Using CNN
abstract
This paper proposes a region based convolutional neural network for segmentation-free word spotting. Our network takes as input an image and a set of word candidate bounding boxes and embeds all bounding boxes into an embedding space, where word spotting can be casted as a simple nearest neighbour search between the query representation and each of the candidate bounding boxes. We make use of PHOC embedding as it has previously achieved significant success in segmentation-based word spotting. Word candidates are generated using a simple procedure based on grouping connected components using some spatial constraints. %For all images in the dataset, we first generate a set of word candidate bounding boxes and then use our R-PHOC network to generate PHOC embeddings for all the bounding boxes using a single forward pass. Experiments show that R-PHOC which operates on images directly can improve the current state-of-the-art in the standard GW dataset and performs as good as PHOCNET in some cases designed for segmentation based word spotting.
Suman K. Ghosh, Ernest Valveny
ICDAR2
2017 Visual Attention Models for Scene Text Recognition
abstract
In this paper we propose an approach to lexicon-free recognition of text in scene images. Our approach relies on a LSTM-based soft visual attention model learned from convolutional features. A set of feature vectors are derived from an intermediate convolutional layer corresponding to different areas of the image. This permits encoding of spatial information into the image representation. In this way, the framework is able to learn how to selectively focus on different parts of the image. At every time step the recognizer emits one character using a weighted combination of the convolutional feature vectors according to the learned attention model. Training can be done end-to-end using only word level annotations. In addition, we show that modifying the beam search algorithm by integrating an explicit language model leads to significantly better recognition results. We validate the performance of our approach on standard SVT and ICDAR'03 scene text datasets, showing state-of-the-art performance in unconstrained text recognition.
Suman K. Ghosh, Ernest Valveny, Andrew D. Bagdanov
ICDAR2
2015 Efficient indexing for Query By String text retrieval
abstract
This paper deals with Query By String word spotting in scene images. A hierarchical text segmentation algorithm based on text specific selective search is used to find text regions. These regions are indexed per character n-grams present in the text region. An attribute representation based on Pyramidal Histogram of Characters (PHOC) is used to compare text regions with the query text. For generation of the index a similar attribute space based Pyramidal Histogram of character n-grams is used. These attribute models are learned using linear SVMs over the Fisher Vector [1] representation of the images along with the PHOC labels of the corresponding strings.
Suman K. Ghosh, Lluís Gómez i Bigorda, Dimosthenis Karatzas, Ernest Valveny
ICDAR4
2015 Query by string word spotting based on character bi-gram indexing
abstract
In this paper we propose a segmentation-free query by string word spotting method. Both the documents and query strings are encoded using a recently proposed word representation that projects images and strings into a common attribute space based on a Pyramidal Histogram of Characters (PHOC). These attribute models are learned using linear SVMs over the Fisher Vector [8] representation of the images along with the PHOC labels of the corresponding strings. In order to search through the whole page, document regions are indexed per character bi-gram using a similar attribute representation. On top of that, we propose an integral image representation of the document using a simplified version of the attribute model for efficient computation. Finally we introduce a re-ranking step in order to boost retrieval performance. We show state-of-the-art results for segmentation-free query by string word spotting in single-writer and multi-writer standard datasets.
Suman K. Ghosh, Ernest Valveny
ICDAR2
2015 ICDAR 2015 competition on Robust Reading
abstract
Results of the ICDAR 2015 Robust Reading Competition are presented. A new Challenge 4 on Incidental Scene Text has been added to the Challenges on Born-Digital Images, Focused Scene Images and Video Text. Challenge 4 is run on a newly acquired dataset of 1,670 images evaluating Text Localisation, Word Recognition and End-to-End pipelines. In addition, the dataset for Challenge 3 on Video Text has been substantially updated with more video sequences and more accurate ground truth data. Finally, tasks assessing End-to-End system performance have been introduced to all Challenges. The competition took place in the first quarter of 2015, and received a total of 44 submissions. Only the tasks newly introduced in 2015 are reported on. The datasets, the ground truth specification and the evaluation protocols are presented together with the results and a brief summary of the participating methods.
Dimosthenis Karatzas, Lluís Gómez i Bigorda, Anguelos Nicolaou, Suman K. Ghosh, Andrew D. Bagdanov, Masakazu Iwamura, Jiri Matas, Lukás Neumann, Vijay Chandrasekhar 0001, Shijian Lu, Faisal Shafait, Seiichi Uchida, Ernest Valveny
ICDAR13
2013 Deformable HOG-Based Shape Descriptor
abstract
In this paper we deal with the problem of recognizing handwritten shapes. We present a new deformable feature extraction method that adapts to the shape to be described, dealing in this way with the variability introduced in the handwriting domain. It consists in a selection of the regions that best define the shape to be described, followed by the computation of histograms of oriented gradients-based features over these points. Our results significantly outperform other descriptors in the literature for the task of hand-drawn shape recognition and handwritten word retrieval.
Jon Almazán, Alicia Fornés, Ernest Valveny
ICDAR3
2013 Unsupervised Wall Detector in Architectural Floor Plans
abstract
Wall detection in floor plans is a crucial step in a complete floor plan recognition system. Walls define the main structure of buildings and convey essential information for the detection of other structural elements. Nevertheless, wall segmentation is a difficult task, mainly because of the lack of a standard graphical notation. The existing approaches are restricted to small group of similar notations or require the existence of pre-annotated corpus of input images to learn each new notation. In this paper we present an automatic wall segmentation system, with the ability to handle completely different notations without the need of any annotated dataset. It only takes advantage of the general knowledge that walls are a repetitive element, naturally distributed within the plan and commonly modeled by straight parallel lines. The method has been tested on four datasets of real floor plans with different notations, and compared with the state-of-the-art. The results show its suitability for different graphical notations, achieving higher recall rates than the rest of the methods while keeping a high average precision.
Lluís-Pere de las Heras, David Fernández Mota, Ernest Valveny, Josep Lladós 0001, Gemma Sánchez
ICDAR3
2012 Document Classification Using Multiple Views
abstract
The combination of multiple features or views when representing documents or other kinds of objects usually leads to improved results in classification (and retrieval) tasks. Most systems assume that those views will be available both at training and test time. However, some views may be too `expensive' to be available at test time. In this paper, we consider the use of Canonical Correlation Analysis to leverage `expensive' views that are available only at training time. Experimental results show that this information may significantly improve the results in a classification task.
Albert Gordo, Florent Perronnin, Ernest Valveny
Document Analysis Systems3
2011 A Non-rigid Feature Extraction Method for Shape Recognition
abstract
This paper presents a methodology for shape recognition that focuses on dealing with the difficult problem of large deformations. The proposed methodology consists in a novel feature extraction technique, which uses a non-rigid representation adaptable to the shape. This technique employs a deformable grid based on the computation of geometrical centroids that follows a region partitioning algorithm. Then, a feature vector is extracted by computing pixel density measures around these geometrical centroids. The result is a shape descriptor that adapts its representation to the given shape and encodes the pixel density distribution. The validity of the method when dealing with large deformations has been experimentally shown over datasets composed of handwritten shapes. It has been applied to signature verification and shape recognition tasks demonstrating high accuracy and low computational cost.
Jon Almazán, Alicia Fornés, Ernest Valveny
ICDAR3
2011 Wall Patch-Based Segmentation in Architectural Floorplans
abstract
Segmentation of architectural floor plans is a challenging task, mainly because of the large variability in the notation between different plans. In general, traditional techniques, usually based on analyzing and grouping structural primitives obtained by vectorization, are only able to handle a reduced range of similar notations. In this paper we propose an alternative patch-based segmentation approach working at pixel level, without need of vectorization. The image is divided into a set of patches and a set of features is extracted for every patch. Then, each patch is assigned to a visual word of a previously learned vocabulary and given a probability of belonging to each class of objects. Finally, a post-process assigns the final label for every pixel. This approach has been applied to the detection of walls on two datasets of architectural floor plans with different notations, achieving high accuracy rates.
Lluís-Pere de las Heras, Joan Mas Romeu, Gemma Sánchez, Ernest Valveny
ICDAR4
2010 A bag of notes approach to writer identification in old handwritten musical scores
abstract
Determining the authorship of a document, namely writer identification, can be an important source of information for document categorization. Contrary to text documents, the identification of the writer of graphical documents is still a challenge. In this paper we present a robust approach for writer identification in a particular kind of graphical documents, old music scores. This approach adapts the bag of visual terms method for coping with graphic documents. The identification is performed only using the graphical music notation. For this purpose, we generate a graphic vocabulary without recognizing any music symbols, and consequently, avoiding the difficulties in the recognition of hand-drawn symbols in old and degraded documents. The proposed method has been tested on a database of old music scores from the 17th to 19th centuries, achieving very high identification rates.
Albert Gordo, Alicia Fornés, Ernest Valveny, Josep Lladós 0001
Document Analysis Systems3
2010 A kernel-based approach to document retrieval
abstract
In this paper we tackle the problem of document image retrieval by combining a similarity measure between documents and the probability that a given document belongs to a certain class. The membership probability to a specific class is computed using Support Vector Machines in conjunction with similarity measure based kernel applied to structural document representations. In the presented experiments, we use different document representations, both visual and structural, and we apply them to a database of historical documents. We show how our method based on similarity kernels outperforms the usual distance-based retrieval.
Albert Gordo, Jaume Gibert, Ernest Valveny, Marçal Rusiñol
Document Analysis Systems3
2010 A system to detect rooms in architectural floor plan images
abstract
In this article, a system to detect rooms in architectural floor plan images is described. We first present a primitive extraction algorithm for line detection. It is based on an original coupling of classical Hough transform with image vectorization in order to perform robust and efficient line detection. We show how the lines that satisfy some graphical arrangements are combined into walls. We also present the way we detect some door hypothesis thanks to the extraction of arcs. Walls and door hypothesis are then used by our room segmentation strategy; it consists in recursively decomposing the image until getting nearly convex regions. The notion of convexity is difficult to quantify, and the selection of separation lines between regions can also be rough. We take advantage of knowledge associated to architectural floor plans in order to obtain mostly rectangular rooms. Qualitative and quantitative evaluations performed on a corpus of real documents show promising results.
Sébastien Macé, Hervé Locteau, Ernest Valveny, Salvatore Tabbone
Document Analysis Systems3
2010 A polar-based logo representation based on topological and colour features
abstract
In this paper, we propose a novel rotation and scale invariant method for colour logo retrieval and classification, which involves performing a simple colour segmentation and subsequently describing each of the resultant colour components based on a set of topological and colour features. A polar representation is used to represent the logo and the subsequent logo matching is based on Cyclic Dynamic Time Warping (CDTW). We also show how combining information about the global distribution of the logo components and their local neighbourhood using the Delaunay triangulation allows to improve the results. All experiments are performed on a dataset of 2500 instances of 100 colour logo images in different rotations and scales.
Farshad Nourbakhsh, Dimosthenis Karatzas, Ernest Valveny
Document Analysis Systems3
2009 A Rotation Invariant Page Layout Descriptor for Document Classification and Retrieval
abstract
Document classification usually requires of structural features such as the physical layout to obtain good accuracy rates on complex documents. This paper introduces a descriptor of the layout and a distance measure based on the cyclic dynamic time warping which can be computed in O(n2). This descriptor is translation invariant and can be easily modified to be scale and rotation invariant. Experiments with this descriptor and its rotation invariant modification are performed on the Girona archives database and compared against another common layout distance, the minimum weight edge cover. The experiments show that these methods outperform the MWEC both in accuracy and speed, particularly on rotated documents.
Albert Gordo, Ernest Valveny
ICDAR2
2008 Performance Evaluation of Symbol Recognition and Spotting Systems: An Overview
abstract
This paper deals with the topic of performance evaluation of the symbol recognition & spotting systems. It presents an overview as a result of the work and the discussions undertaken by a working group on this subject. The paper starts by giving a general view of symbol recognition & spotting and performance evaluation. Next, the two main issues of performance evaluation are discussed: groundtruthing and performance characterization. Different problems related to both issues are addressed: groundtruthing of real documents, generation of synthetic documents, degradation models, the use of a priori knowledge, mapping of the groundtruth with the system results, and so on. Open problems arising from this overview are also discussed at the end of the paper.
Mathieu Delalandre, Ernest Valveny, Josep Lladós 0001
Document Analysis Systems2
2007 Combination of OCR Engines for Page Segmentation Based on Performance Evaluation
abstract
In this paper we present a method to improve the performance of individual page segmentation engines based on the combination of the output of several engines. The rules of combination are designed after analyzing the results of each individual method. This analysis is performed using a performance evaluation framework that aims at characterizing each method according to its strengths and weaknesses rather than computing a single performance measure telling which is the "best" segmentation method.
Miquel Ferrer, Ernest Valveny
ICDAR2
2007 A Review of Shape Descriptors for Document Analysis
abstract
Shape descriptors play an important role in many document analysis application. In this paper we review some of the shape descriptors proposed in the last years from a new point of view. We propose the definitions of descriptor and primitive and introduce the notion of feature extraction method. With these definitions, we propose a new classification of shape descriptors that permits to classify according to their properties pointing out their strengths and weaknesses.
Oriol Ramos Terrades, Salvatore Tabbone, Ernest Valveny
ICDAR3
2005 Local Norm Features based on ridgelets Transform
abstract
We propose a set of shape descriptors for image retrieval of graphic documents based on the ridgelets transform, which can be seen as a combination of the Radon transform and the wavelets transform. It is especially well suited to detect linear features, the most relevant features in graphic documents. It also provides a multiscale representation, useful for indexing and retrieval purposes. From the ridgelets representation of an image, we have defined a set of local norm descriptors based on computing a norm over some specific areas of the image. This kind of descriptors are very flexible since we can define different sets of descriptors just by changing such areas of influence in the image. We have also defined a combination of descriptors at several scales of decomposition in order to improve retrieval results.
Oriol Ramos Terrades, Ernest Valveny
ICDAR2
2004 A Platform to Extract Knowledge from Graphic Documents. Application to an Architectural Sketch Understanding Scenario
Gemma Sánchez, Ernest Valveny, Josep Lladós 0001, Joan Mas Romeu, Narcís Lozano
Document Analysis Systems2
2004 Performance Evaluation of Symbol Recognition
Ernest Valveny, Philippe Dosch
Document Analysis Systems1
2003 Radon Transform for Lineal Symbol Representation
abstract
Content-based retrieval and recognition of graphic images requires good models for symbol representation, able to identify those features providing the most relevant information about the shape and the visual appearance of symbols. In this work we have used the Radon transform as the basis to extract the representation of graphic images as it permits to globally detect lineal singularities in an image, which are the most important source of information in these images. The image obtained after applying Radon transform can be used directly to describe the symbol, or can be used to extract new and compact descriptors from it, which will also be based on lineal information about the image. We present some preliminary results showing the usefulness of this representation with a set of architectural symbols.
Oriol Ramos Terrades, Ernest Valveny
ICDAR2
2003 Numeral recognition for quality control of surgical sachets
abstract
In this paper we describe an application of OCR techniques to quality control in industrial production. The purpose of the system is to verify the correct printing of numerical information in sachets with surgical material. Numerals are printed on an aluminium surface covered by a transparent plastic film, which can produce shadows or reflections in the image. The main difficulties for character recognition arise from low acquisition resolution, noise, heavy or light printing, and different printing patterns. The system must perform with an error rate lower than 0.1% and with the minimal computation time. Thus, we have decided to use well-known and simple algorithms, adding to them some refinements which take advantage of specific domain knowledge, making them more robust and reliable. The system is currently working with real production complying with the required specifications.
Ernest Valveny, Antonio M. López 0001
ICDAR1
2001 Learning of Structural Descriptions of Graphic Symbols Using Deformable Template Matching
abstract
Accurate symbol recognition in graphic documents needs an accurate representation of the symbols to be recognized. If structural approaches are used for recognition, symbols have to be described in terms of their shape, using structural relationships among extracted features. Unlike statistical pattern recognition, in structural methods, symbols are usually, manually defined from expertise knowledge, and not automatically, inferred from sample images. In this work we explain one approach to learn from examples a representative structural description of a symbol, thus providing better information about shape variability. The description of a symbol is based on a probabilistic model. It consists of a set of lines described by, the mean and the variance of line parameters, respectively, providing information about the model of the symbol, and its shape variability. The representation of each image in the sample set as a set of lines is achieved using deformable template matching.
Ernest Valveny, Enric Martí
ICDAR1
1999 Application of Deformable Template Matching to Symbol Recognition in Hand-written Architectural Drawings
abstract
We propose using deformable template matching as a new approach for recognising characters and lineal symbols in handwritten line drawings, instead of traditional methods based on vectorization and feature extraction. Bayesian formulation of the deformable template matching allows combining fidelity of the ideal shape of the symbol with maximum flexibility to get the best fit to the input image. The lineal nature of symbols can be exploited to define a suitable representation of models and the set of deformations to be applied to them. Matching, however, is done over the original binary image to avoid losing relevant features during vectorization. We have applied this method to handwritten architectural drawings and experimental results demonstrate that symbols that are highly distorted from ideal shape can be accurately identified.
Ernest Valveny, Enric Martí
ICDAR1