EDBT 2026 Demo / reviewers in the wild / expert
Pau Riba
dblp:156/2959
· DBLP profile ↗
7ranked-venue papers in the field
4as first author
3since 2021 · last 2021
0000-0002-4710-0864ORCID · verified
Domains — venue-derived; a paper can count in several
Other / Interdisciplinary · 7 (4 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | DocSynth: A Layout Guided Approach for Controllable Document Image Synthesis
Sanket Biswas, Pau Riba, Josep Lladós 0001, Umapada Pal 0001 |
ICDAR (3) | 2 |
| 2021 | Date Estimation in the Wild of Scanned Historical Photos: An Image Retrieval Approach
Adrià Molina, Pau Riba, Lluís Gómez i Bigorda, Oriol Ramos Terrades, Josep Lladós 0001 |
ICDAR (2) | 2 |
| 2021 | Learning to Rank Words: Optimizing Ranking Metrics for Word Spotting
Pau Riba, Adrià Molina, Lluís Gómez i Bigorda, Oriol Ramos Terrades, Josep Lladós 0001 |
ICDAR (2) | 1 |
| 2019 | Table Detection in Invoice Documents by Graph Neural NetworksabstractTabular structures in documents offer a complementary dimension to the raw textual data, representing logical or quantitative relationships among pieces of information. In digital mail room applications, where a large amount of administrative documents must be processed with reasonable accuracy, the detection and interpretation of tables is crucial. Table recognition has gained interest in document image analysis, in particular in unconstrained formats (absence of rule lines, unknown information of rows and columns). In this work, we propose a graph-based approach for detecting tables in document images. Instead of using the raw content (recognized text), we make use of the location, context and content type, thus it is purely a structure perception approach, not dependent on the language and the quality of the text reading. Our framework makes use of Graph Neural Networks (GNNs) in order to describe the local repetitive structural information of tables in invoice documents. Our proposed model has been experimentally validated in two invoice datasets and achieved encouraging results. Additionally, due to the scarcity of benchmark datasets for this task, we have contributed to the community a novel dataset derived from the RVL-CDIP invoice data. It will be publicly released to facilitate future research. Pau Riba, Anjan Dutta 0001, Lutz Goldmann, Alicia Fornés, Oriol Ramos Terrades, Josep Lladós 0001 |
ICDAR | 1 |
| 2017 | Pyramidal Stochastic Graphlet Embedding for Document Pattern ClassificationabstractDocument pattern classification methods using graphs have received a lot of attention because of its robust representation paradigm and rich theoretical background. However, the way of preserving and the process for delineating documents with graphs introduce noise in the rendition of underlying data, which creates instability in the graph representation. To deal with such unreliability in representation, in this paper, we propose Pyramidal Stochastic Graphlet Embedding (PSGE). Given a graph representing a document pattern, our method first computes a graph pyramid by successively reducing the base graph. Once the graph pyramid is computed, we apply Stochastic Graphlet Embedding (SGE) for each level of the pyramid and combine their embedded representation to obtain a global delineation of the original graph. The consideration of pyramid of graphs rather than just a base graph extends the representational power of the graph embedding, which reduces the instability caused due to noise and distortion. When plugged with support vector machine, our proposed PSGE has outperformed the state-of-the-art results in recognition of handwritten words as well as graphical symbols. Anjan Dutta 0001, Pau Riba, Josep Lladós 0001, Alicia Fornés |
ICDAR | 2 |
| 2017 | Improving Information Retrieval in Multiwriter Scenario by Exploiting the Similarity Graph of Document TermsabstractInformation Retrieval (IR) is the activity of obtaining information resources relevant to a questioned information. It usually retrieves a set of objects ranked according to the relevancy to the needed fact. In document analysis, information retrieval receives a lot of attention in terms of symbol and word spotting. However, through decades the community mostly focused either on printed or on single writer scenario, where the state-of-the-art results have achieved reasonable performance on the available datasets. Nevertheless, the existing algorithms do not perform accordingly on multiwriter scenario. A graph representing relations between a set of objects is a structure where each node delineates an individual element and the similarity between them is represented as a weight on the connecting edge. In this paper, we explore different analytics of graphs constructed from words or graphical symbols, such as diffusion, shortest path, etc. to improve the performance of information retrieval methods in multiwriter scenario. Pau Riba, Anjan Dutta 0001, Sounak Dey, Josep Lladós 0001, Alicia Fornés |
ICDAR | 1 |
| 2015 | Handwritten word spotting by inexact matching of grapheme graphsabstractThis paper presents a graph-based word spotting for handwritten documents. Contrary to most word spotting techniques, which use statistical representations, we propose a structural representation suitable to be robust to the inherent deformations of handwriting. Attributed graphs are constructed using a part-based approach. Graphemes extracted from shape convexities are used as stable units of handwriting, and are associated to graph nodes. Then, spatial relations between them determine graph edges. Spotting is defined in terms of an error-tolerant graph matching using bipartite-graph matching algorithm. To make the method usable in large datasets, a graph indexing approach that makes use of binary embeddings is used as preprocessing. Historical documents are used as experimental framework. The approach is comparable to statistical ones in terms of time and memory requirements, especially when dealing with large document collections. Pau Riba, Josep Lladós 0001, Alicia Fornés |
ICDAR | 1 |