Pau Riba

dblp:156/2959 · DBLP profile ↗
← Back
28ranked-venue papers
10as first author
10since 2021 · last 2022
0000-0002-4710-0864ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 9 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 7 · 4 first-author · 3 since 2021
YearPublicationVenuePosition
2022 Musigraph: Optical Music Recognition Through Object Detection and Graph Neural Network
Arnau Baró, Pau Riba, Alicia Fornés
ICFHR2
2022 Content and Style Aware Generation of Text-Line Images for Handwriting Recognition
abstract
Handwritten Text Recognition has achieved an impressive performance in public benchmarks. However, due to the high inter- and intra-class variability between handwriting styles, such recognizers need to be trained using huge volumes of manually labeled training data. To alleviate this labor-consuming problem, synthetic data produced with TrueType fonts has been often used in the training loop to gain volume and augment the handwriting style variability. However, there is a significant style bias between synthetic and real data which hinders the improvement of recognition performance. To deal with such limitations, we propose a generative method for handwritten text-line images, which is conditioned on both visual appearance and textual content. Our method is able to produce long text-line samples with diverse handwriting styles. Once properly trained, our method can also be adapted to new target data by only accessing unlabeled text-line images to mimic handwritten styles and produce images with any textual content. Extensive experiments have been done on making use of the generated samples to boost Handwritten Text Recognition performance. Both qualitative and quantitative results demonstrate that the proposed approach outperforms the current state of the art.
Lei Kang 0002, Pau Riba, Marçal Rusiñol, Alicia Fornés, Mauricio Villegas
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 Pay attention to what you read: Non-recurrent handwritten text-Line recognition
Lei Kang 0002, Pau Riba, Marçal Rusiñol, Alicia Fornés, Mauricio Villegas
Pattern Recognit.2
2022 Table detection in business document images by message passing networks
Pau Riba, Lutz Goldmann, Oriol Ramos Terrades, Diede Rusticus, Alicia Fornés, Josep Lladós 0001
Pattern Recognit.1
2021 DocSynth: A Layout Guided Approach for Controllable Document Image Synthesis
Sanket Biswas, Pau Riba, Josep Lladós 0001, Umapada Pal 0001
ICDAR (3)2
2021 Date Estimation in the Wild of Scanned Historical Photos: An Image Retrieval Approach
Adrià Molina, Pau Riba, Lluís Gómez i Bigorda, Oriol Ramos Terrades, Josep Lladós 0001
ICDAR (2)2
2021 Learning to Rank Words: Optimizing Ranking Metrics for Word Spotting
Pau Riba, Adrià Molina, Lluís Gómez i Bigorda, Oriol Ramos Terrades, Josep Lladós 0001
ICDAR (2)1
2021 Beyond document object detection: instance-level segmentation of complex layouts
Sanket Biswas, Pau Riba, Josep Lladós 0001, Umapada Pal 0001
Int. J. Document Anal. Recognit.2
2021 Candidate fusion: Integrating language modelling into a sequence-to-sequence handwritten word recognition architecture
Lei Kang 0002, Pau Riba, Mauricio Villegas, Alicia Fornés, Marçal Rusiñol
Pattern Recognit.2
2021 Learning graph edit distance by graph neural networks
Pau Riba, Andreas Fischer 0002, Josep Lladós 0001, Alicia Fornés
Pattern Recognit.1
2020 GANwriting: Content-Conditioned Generation of Styled Handwritten Word Images
Lei Kang 0002, Pau Riba, Yaxing Wang, Marçal Rusiñol, Alicia Fornés, Mauricio Villegas
ECCV (23)2
2020 Distilling Content from Style for Handwritten Word Recognition
abstract
Despite the latest transcription accuracies reached using deep neural network architectures, handwritten text recognition still remains a challenging problem, mainly because of the large inter-writer style variability. Both augmenting the training set with artificial samples using synthetic fonts, and writer adaptation techniques have been proposed to yield more generic approaches aimed at dodging style unevenness. In this work, we take a step closer to learn style independent features from handwritten word images. We propose a novel method that is able to disentangle the content and style aspects of input images by jointly optimizing a generative process and a handwritten word recognizer. The generator is aimed at transferring writing style features from one sample to another in an image-to-image translation approach, thus leading to a learned content-centric features that shall be independent to writing style attributes. Our proposed recognition model is able then to leverage such writer-agnostic features to reach better recognition performances. We advance over prior training strategies and demonstrate with qualitative and quantitative evaluations the performance of both the generative process and the recognition efficiency in the IAM dataset.
Lei Kang 0002, Pau Riba, Marçal Rusiñol, Alicia Fornés, Mauricio Villegas
ICFHR2
2020 Named Entity Recognition and Relation Extraction with Graph Neural Networks in Semi Structured Documents
abstract
The use of administrative documents to communicate and leave record of business information requires of methods able to automatically extract and understand the content from such documents in a robust and efficient way. In addition, the semi-structured nature of these reports is specially suited for the use of graph-based representations which are flexible enough to adapt to the deformations from the different document templates. Moreover, Graph Neural Networks provide the proper methodology to learn relations among the data elements in these documents. In this work we study the use of Graph Neural Network architectures to tackle the problem of entity recognition and relation extraction in semi-structured documents. Our approach achieves state of the art results in the three tasks involved in the process. Additionally, the experimentation with two datasets of different nature demonstrates the good generalization ability of our approach.
Manuel Carbonell, Pau Riba, Mauricio Villegas, Alicia Fornés, Josep Lladós 0001
ICPR2
2020 Unsupervised Adaptation for Synthetic-to-Real Handwritten Word Recognition
abstract
Handwritten Text Recognition (HTR) is still a challenging problem because it must deal with two important difficulties: the variability among writing styles, and the scarcity of labelled data. To alleviate such problems, synthetic data generation and data augmentation are typically used to train HTR systems. However, training with such data produces encouraging but still inaccurate transcriptions in real words. In this paper, we propose an unsupervised writer adaptation approach that is able to automatically adjust a generic handwritten word recognizer, fully trained with synthetic fonts, towards a new incoming writer. We have experimentally validated our proposal using five different datasets, covering several challenges (i) the document source: modern and historic samples, which may involve paper degradation problems; (ii) different handwriting styles: single and multiple writer collections; and (iii) language, which involves different character combinations. Across these challenging collections, we show that our system is able to maintain its performance, thus, it provides a practical and generic approach to deal with new document collections without requiring any expensive and tedious manual annotation step.
Lei Kang 0002, Marçal Rusiñol, Alicia Fornés, Pau Riba, Mauricio Villegas
WACV4
2020 Hierarchical stochastic graphlet embedding for graph-based pattern recognition
abstract
Abstract Despite being very successful within the pattern recognition and machine learning community, graph-based methods are often unusable because of the lack of mathematical operations defined in graph domain. Graph embedding, which maps graphs to a vectorial space, has been proposed as a way to tackle these difficulties enabling the use of standard machine learning techniques. However, it is well known that graph embedding functions usually suffer from the loss of structural information. In this paper, we consider the hierarchical structure of a graph as a way to mitigate this loss of information. The hierarchical structure is constructed by topologically clustering the graph nodes and considering each cluster as a node in the upper hierarchical level. Once this hierarchical structure is constructed, we consider several configurations to define the mapping into a vector space given a classical graph embedding, in particular, we propose to make use of the stochastic graphlet embedding (SGE). Broadly speaking, SGE produces a distribution of uniformly sampled low-to-high-order graphlets as a way to embed graphs into the vector space. In what follows, the coarse-to-fine structure of a graph hierarchy and the statistics fetched by the SGE complements each other and includes important structural information with varied contexts. Altogether, these two techniques substantially cope with the usual information loss involved in graph embedding techniques, obtaining a more robust graph representation. This fact has been corroborated through a detailed experimental evaluation on various benchmark graph datasets, where we outperform the state-of-the-art methods.
Anjan Dutta 0001, Pau Riba, Josep Lladós 0001, Alicia Fornés
Neural Comput. Appl.2
2020 Hierarchical graphs for coarse-to-fine error tolerant matching
Pau Riba, Josep Lladós 0001, Alicia Fornés
Pattern Recognit. Lett.1
2019 Doodle to Search: Practical Zero-Shot Sketch-Based Image Retrieval
abstract
In this paper, we investigate the problem of zero-shot sketch-based image retrieval (ZS-SBIR), where human sketches are used as queries to conduct retrieval of photos from unseen categories. We importantly advance prior arts by proposing a novel ZS-SBIR scenario that represents a firm step forward in its practical application. The new setting uniquely recognizes two important yet often neglected challenges of practical ZS-SBIR, (i) the large domain gap between amateur sketch and photo, and (ii) the necessity for moving towards large-scale retrieval. We first contribute to the community a novel ZS-SBIR dataset, QuickDraw-Extended, that consists of 330,000 sketches and 204,000 photos spanning across 110 categories. Highly abstract amateur human sketches are purposefully sourced to maximize the domain gap, instead of ones included in existing datasets that can often be semi-photorealistic. We then formulate a ZS-SBIR framework to jointly model sketches and photos into a common embedding space. A novel strategy to mine the mutual information among domains is specifically engineered to alleviate the domain gap. External semantic knowledge is further embedded to aid semantic transfer. We show that, rather surprisingly, retrieval performance significantly outperforms that of state-of-the-art on existing datasets that can already be achieved using a reduced version of our model. We further demonstrate the superior performance of our full model by comparing with a number of alternatives on the newly proposed dataset. The new dataset, plus all training and testing code of our model, will be publicly released to facilitate future research.
Sounak Dey, Pau Riba, Anjan Dutta 0001, Josep Lladós 0001, Yi-Zhe Song
CVPR2
2019 Table Detection in Invoice Documents by Graph Neural Networks
abstract
Tabular structures in documents offer a complementary dimension to the raw textual data, representing logical or quantitative relationships among pieces of information. In digital mail room applications, where a large amount of administrative documents must be processed with reasonable accuracy, the detection and interpretation of tables is crucial. Table recognition has gained interest in document image analysis, in particular in unconstrained formats (absence of rule lines, unknown information of rows and columns). In this work, we propose a graph-based approach for detecting tables in document images. Instead of using the raw content (recognized text), we make use of the location, context and content type, thus it is purely a structure perception approach, not dependent on the language and the quality of the text reading. Our framework makes use of Graph Neural Networks (GNNs) in order to describe the local repetitive structural information of tables in invoice documents. Our proposed model has been experimentally validated in two invoice datasets and achieved encouraging results. Additionally, due to the scarcity of benchmark datasets for this task, we have contributed to the community a novel dataset derived from the RVL-CDIP invoice data. It will be publicly released to facilitate future research.
Pau Riba, Anjan Dutta 0001, Lutz Goldmann, Alicia Fornés, Oriol Ramos Terrades, Josep Lladós 0001
ICDAR1
2019 From Optical Music Recognition to Handwritten Music Recognition: A baseline
Arnau Baró, Pau Riba, Jorge Calvo-Zaragoza, Alicia Fornés
Pattern Recognit. Lett.2
2018 Word-Hunter: A Gamesourcing Experience to Validate the Transcription of Historical Manuscripts
abstract
Nowadays, there are still many handwritten historical documents in archives waiting to be transcribed and indexed. Since manual transcription is tedious and time consuming, the automatic transcription seems the path to follow. However, the performance of current handwriting recognition techniques is not perfect, so a manual validation is mandatory. Crowdsourcing is a good strategy for manual validation, however it is a tedious task. In this paper we analyze experiences based in gamification in order to propose and design a gamesourcing framework that increases the interest of users. Then, we describe and analyze our experience when validating the automatic transcription using the gamesourcing application. Moreover, thanks to the combination of clustering and handwriting recognition techniques, we can speed up the validation while maintaining the performance.
Pau Riba, Alicia Fornés, Joan Mas Romeu, Josep Lladós 0001, Joana Maria Pujadas-Mora
ICFHR2
2018 Learning Graph Distances with Message Passing Neural Networks
abstract
Graph representations have been widely used in pattern recognition thanks to their powerful representation formalism and rich theoretical background. A number of error-tolerant graph matching algorithms such as graph edit distance have been proposed for computing a distance between two labelled graphs. However, they typically suffer from a high computational complexity, which makes it difficult to apply these matching algorithms in a real scenario. In this paper, we propose an efficient graph distance based on the emerging field of geometric deep learning. Our method employs a message passing neural network to capture the graph structure and learns a metric with a siamese network approach. The performance of the proposed graph distance is validated in two application cases, graph classification and graph retrieval of handwritten words, and shows a promising performance when compared with (approximate) graph edit distance benchmarks.
Pau Riba, Andreas Fischer 0002, Josep Lladós 0001, Alicia Fornés
ICPR1
2017 Pyramidal Stochastic Graphlet Embedding for Document Pattern Classification
abstract
Document pattern classification methods using graphs have received a lot of attention because of its robust representation paradigm and rich theoretical background. However, the way of preserving and the process for delineating documents with graphs introduce noise in the rendition of underlying data, which creates instability in the graph representation. To deal with such unreliability in representation, in this paper, we propose Pyramidal Stochastic Graphlet Embedding (PSGE). Given a graph representing a document pattern, our method first computes a graph pyramid by successively reducing the base graph. Once the graph pyramid is computed, we apply Stochastic Graphlet Embedding (SGE) for each level of the pyramid and combine their embedded representation to obtain a global delineation of the original graph. The consideration of pyramid of graphs rather than just a base graph extends the representational power of the graph embedding, which reduces the instability caused due to noise and distortion. When plugged with support vector machine, our proposed PSGE has outperformed the state-of-the-art results in recognition of handwritten words as well as graphical symbols.
Anjan Dutta 0001, Pau Riba, Josep Lladós 0001, Alicia Fornés
ICDAR2
2017 Improving Information Retrieval in Multiwriter Scenario by Exploiting the Similarity Graph of Document Terms
abstract
Information Retrieval (IR) is the activity of obtaining information resources relevant to a questioned information. It usually retrieves a set of objects ranked according to the relevancy to the needed fact. In document analysis, information retrieval receives a lot of attention in terms of symbol and word spotting. However, through decades the community mostly focused either on printed or on single writer scenario, where the state-of-the-art results have achieved reasonable performance on the available datasets. Nevertheless, the existing algorithms do not perform accordingly on multiwriter scenario. A graph representing relations between a set of objects is a structure where each node delineates an individual element and the similarity between them is represented as a weight on the connecting edge. In this paper, we explore different analytics of graphs constructed from words or graphical symbols, such as diffusion, shortest path, etc. to improve the performance of information retrieval methods in multiwriter scenario.
Pau Riba, Anjan Dutta 0001, Sounak Dey, Josep Lladós 0001, Alicia Fornés
ICDAR1
2017 Large-scale graph indexing using binary embeddings of node contexts for information spotting in document image databases
Pau Riba, Josep Lladós 0001, Alicia Fornés, Anjan Dutta 0001
Pattern Recognit. Lett.1
2016 Towards the Recognition of Compound Music Notes in Handwritten Music Scores
abstract
The recognition of handwritten music scores still remains an open problem. The existing approaches can only deal with very simple handwritten scores mainly because of the variability in the handwriting style and the variability in the composition of groups of music notes (i.e. compound music notes). In this work we focus on this second problem and propose a method based on perceptual grouping for the recognition of compound music notes. Our method has been tested using several handwritten music scores of the CVC-MUSCIMA database and compared with a commercial Optical Music Recognition (OMR) software. Given that our method is learning-free, the obtained results are promising.
Arnau Baró, Pau Riba, Alicia Fornés
ICFHR2
2015 Handwritten word spotting by inexact matching of grapheme graphs
abstract
This paper presents a graph-based word spotting for handwritten documents. Contrary to most word spotting techniques, which use statistical representations, we propose a structural representation suitable to be robust to the inherent deformations of handwriting. Attributed graphs are constructed using a part-based approach. Graphemes extracted from shape convexities are used as stable units of handwriting, and are associated to graph nodes. Then, spatial relations between them determine graph edges. Spotting is defined in terms of an error-tolerant graph matching using bipartite-graph matching algorithm. To make the method usable in large datasets, a graph indexing approach that makes use of binary embeddings is used as preprocessing. Historical documents are used as experimental framework. The approach is comparable to statistical ones in terms of time and memory requirements, especially when dealing with large document collections.
Pau Riba, Josep Lladós 0001, Alicia Fornés
ICDAR1
2014 On the Influence of Key Point Encoding for Handwritten Word Spotting
abstract
In this paper we evaluate the influence of the selection of key points and the associated features in the performance of word spotting processes. In general, features can be extracted from a number of characteristic points like corners, contours, skeletons, maxima, minima, crossings, etc. A number of descriptors exist in the literature using different interest point detectors. But the intrinsic variability of handwriting vary strongly on the performance if the interest points are not stable enough. In this paper, we analyze the performance of different descriptors for local interest points. As benchmarking dataset we have used the Barcelona Marriage Database that contains handwritten records of marriages over five centuries.
David Fernández Mota, Pau Riba, Alicia Fornés, Josep Lladós 0001
ICFHR2
2014 e-Crowds: A Mobile Platform for Browsing and Searching in Historical Demography-Related Manuscripts
abstract
This paper presents a prototype system running on portable devices for browsing and word searching through historical handwritten document collections. The platform adapts the paradigm of eBook reading, where the narrative is not necessarily sequential, but centered on the user actions. The novelty is to replace digitally born books by digitized historical manuscripts of marriage licenses, so document analysis tasks are required in the browser. With an active reading paradigm, the user can cast queries of people names, so he/she can implicitly follow genealogical links. In addition, the system allows combined searches: the user can refine a search by adding more words to search. As a second contribution, the retrieval functionality involves as a core technology a word spotting module with an unified approach, which allows combined query searches, and also two input modalities: query-by-example, and query-by-string.
Pau Riba, Jon Almazán, Alicia Fornés, David Fernández Mota, Ernest Valveny, Josep Lladós 0001
ICFHR1