VLDB 2026 Research / reviewers in the wild / expert
Francisco Claude
dblp:94/4051
· DBLP profile ↗
36ranked-venue papers
24as first author
1since 2021 · last 2021
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 22 · 18 first-authorTheory of computation · 10 · 4 first-author · 1 since 2021Artificial intelligence and machine learning · 3 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-authorSoftware engineering, systems software and programming languages · 1Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Grammar-compressed indexes with logarithmic search time
Francisco Claude, Gonzalo Navarro 0001, Alejandro Pacheco |
J. Comput. Syst. Sci. | 1 |
| 2019 | On the reproducibility of experiments of indexing repetitive document collections
Antonio Fariña, Miguel A. Martínez-Prieto, Francisco Claude, Gonzalo Navarro 0001, Juan J. Lastra-Díaz, Nicola Prezza, Diego Seco Naveiras |
Inf. Syst. | 3 |
| 2017 | Competitive Author Profiling Using Compression-Based StrategiesabstractAuthor profiling consists in determining some demographic attributes — such as gender, age, nationality, language, religion, and others — of an author for a given document. This task, which has applications in fields such as forensics, security, or marketing, has been approached from different areas, especially from linguistics and natural language processing, by extracting different types of features from training documents, usually content — and style-based features. In this paper we address the problem by using several compression-inspired strategies that generate different models without analyzing or extracting specific features from the textual content, making them style-oblivious approaches. We analyze the behavior of these techniques, combine them and compare them with other state-of-the-art methods. We show that they can be competitive in terms of accuracy, giving the best predictions for some domains, and they are efficient in time performance. Francisco Claude, Daniil Galaktionov, Roberto Konow, Susana Ladra, Oscar Pedreira |
Int. J. Uncertain. Fuzziness Knowl. Based Syst. | 1 |
| 2016 | Compression-Inspired Author ProfilingabstractAuthor profiling, that is, determining the demographic attributes -such as gender, age, nationality, language, religion, and others- of an author for a given document, has been approached from different areas, especially from linguistics and natural language processing, by extracting different types of features from training documents, usually content- and style-based features.This work addresses the problem of identifying age and gender of the author of a given document with compression-inspired strategies without analysing or extracting specifc features from the textual content, making them style-oblivious approaches. Since they do not require any a priori knowledge of the linguistic properties, they are of special interest for domain where we do not have an a priori intuition of its properties, such as DNA and protein sequences, stock market data, or medical monitoring. Francisco Claude, Roberto Konow, Susana Ladra |
DCC | 1 |
| 2016 | Universal indexes for highly repetitive document collections
Francisco Claude, Antonio Fariña, Miguel A. Martínez-Prieto, Gonzalo Navarro 0001 |
Inf. Syst. | 1 |
| 2016 | Practical compressed string dictionaries
Miguel A. Martínez-Prieto, Nieves R. Brisaboa, Rodrigo Cánovas, Francisco Claude, Gonzalo Navarro 0001 |
Inf. Syst. | 4 |
| 2015 | The wavelet matrix: An efficient wavelet tree for large alphabets
Francisco Claude, Gonzalo Navarro 0001, Alberto Ordóñez Pereira |
Inf. Syst. | 1 |
| 2015 | Fast in-memory XPath search using compressed indexesabstractSummary Extensible Markup Language (XML) documents consist of text data plus structured data (markup). XPath allows to query both text and structure. Evaluating such hybrid queries is challenging. We present a system for in‐memory evaluation ofXPath search queries, that is, queries with text and structure predicates, yet without advanced features such as backward axes, arithmetics, and joins. We show that for this query fragment, which containsForward Core XPath, our system, dubbed Succinct XML Self‐Index (‘SXSI’), outperforms existing systems by 1–3 orders of magnitude. SXSI is based on state‐of‐the‐art indexes for text and structure data. It combines two novelties. On one hand, it represents the XML data in a compact indexed form, which allows it to handle larger collections in main memory while supporting powerful search and navigation operations over the text and the structure. On the other hand, it features an execution engine that uses tree automata and cleverly chooses evaluation orders that leverage the speeds of the respective indexes. SXSI is modular and allows seamless replacement of its indexes. This is demonstrated through experiments with (1) a text index specialized for search of bio sequences, and (2) a word‐based text index specialized for natural language search. Copyright © 2013 John Wiley & Sons, Ltd. Diego Arroyuelo, Francisco Claude, Sebastian Maneth, Veli Mäkinen, Gonzalo Navarro 0001, Kim Nguyen 0001, Jouni Sirén, Niko Välimäki |
Softw. Pract. Exp. | 2 |
| 2014 | Efficient Indexing and Representation of Web Access Logs
Francisco Claude, Roberto Konow, Gonzalo Navarro 0001 |
SPIRE | 1 |
| 2014 | Efficient Fully-Compressed Sequence Representations
Jérémy Barbay, Francisco Claude, Travis Gagie, Gonzalo Navarro 0001, Yakov Nekrich |
Algorithmica | 2 |
| 2014 | On the compression of search trees
Francisco Claude, Patrick K. Nicholson, Diego Seco Naveiras |
Inf. Process. Manag. | 1 |
| 2013 | Adaptive Data Structures for Permutations and Binary Relations
Francisco Claude, J. Ian Munro |
SPIRE | 1 |
| 2013 | Document Listing on Versioned Documents
Francisco Claude, J. Ian Munro |
SPIRE | 1 |
| 2013 | Compact binary relation representations with rich functionality
Jérémy Barbay, Francisco Claude, Gonzalo Navarro 0001 |
Inf. Comput. | 2 |
| 2012 | Differentially Encoded Search TreesabstractLet X = x1, x2,.... xnbe a sequence of non-decreasing integer values. Storing a compressed representation of X that supports access and search is a problem that occurs in many domains. The most common solution to this problem encodes the differences between consecutive elements in the sequence, and includes additional information (samples) to support efficient searching on the encoded values. We introduce a completely different alternative that achieves compression by encoding the differences in a search tree. Our proposal has many applications such as the representation of posting lists, geographic data, sparse bitmaps, and compressed suffix arrays, to name just a few. The structure is practical and we provide an experimental comparison to show that it is also competitive with the existing techniques. Francisco Claude, Patrick K. Nicholson, Diego Seco Naveiras |
DCC | 1 |
| 2012 | The Wavelet Matrix
Francisco Claude, Gonzalo Navarro 0001 |
SPIRE | 1 |
| 2012 | Improved Grammar-Based Compressed Indexes
Francisco Claude, Gonzalo Navarro 0001 |
SPIRE | 1 |
| 2012 | Word-based self-indexes for natural language textabstractThe inverted index supports efficient full-text searches on natural language text collections. It requires some extra space over the compressed text that can be traded for search speed. It is usually fast for single-word searches, yet phrase searches require more expensive intersections. In this article we introduce a different kind of index. It replaces the text using essentially the same space required by the compressed text alone (compression ratio around 35%). Within this space it supports not only decompression of arbitrary passages, but efficient word and phrase searches. Searches are orders of magnitude faster than those over inverted indexes when looking for phrases, and still faster on single-word searches when little space is available. Our new indexes are particularly fast at counting the occurrences of words or phrases. This is useful for computing relevance of words or phrases. We adapt self-indexes that succeeded in indexing arbitrary strings within compressed space to deal with large alphabets. Natural language texts are then regarded as sequences of words, not characters, to achieve word-based self-indexes. We design an architecture that separates the searchable sequence from its presentation aspects. This permits applying case folding, stemming, removing stopwords, etc. as is usual on inverted indexes. Antonio Fariña, Nieves R. Brisaboa, Gonzalo Navarro 0001, Francisco Claude, Ángeles Saavedra Places |
ACM Trans. Inf. Syst. | 4 |
| 2011 | Indexes for highly repetitive document collectionsabstractWe introduce new compressed inverted indexes for highly repetitive document collections. They are based on run-length, Lempel-Ziv, or grammar-based compression of the differential inverted lists, instead of gap-encoding them as is the usual practice. We show that our compression methods significantly reduce the space achieved by classical compression, at the price of moderate slowdowns. Moreover, many of our methods are universal, that is, they do not need to know the versioning structure of the collection. Francisco Claude, Antonio Fariña, Miguel A. Martínez-Prieto, Gonzalo Navarro 0001 |
CIKM | 1 |
| 2011 | Practical representations for web and social graphsabstractIn this paper we focus on representing Web and social graphs. Our work is motivated by the need of mining information out of these graphs, thus our representations do not only aim at compressing the graphs, but also at supporting efficient navigation. This allows us to process bigger graphs in main memory, avoiding the slowdown brought by resorting on external memory. We first show how by just partitioning the graph and combining two existing techniques for Web graph compression, k2-trees [Brisaboa, Ladra and Navarro, SPIRE 2009] and RePair-Graph [Claude and Navarro, TWEB 2010], exploiting the fact that most links are intra-domain, we obtain the best time/space trade-off for direct and reverse navigation when compared to the state of the art. In social networks, splitting the graph to achieve a good decomposition is not easy. For this case, we explore a new proposal for indexing MPK linearizations [Maserrat and Pei, KDD 2010], which have proven to be an effective way of representing social networks in little space by exploiting common dense subgraphs. Our proposal offers better worst case bounds in space and time, and is also a competitive alternative in practice. Francisco Claude, Susana Ladra |
CIKM | 1 |
| 2011 | Space Efficient Wavelet Tree Construction
Francisco Claude, Patrick K. Nicholson, Diego Seco Naveiras |
SPIRE | 1 |
| 2011 | Compressed String Dictionaries
Nieves R. Brisaboa, Rodrigo Cánovas, Francisco Claude, Miguel A. Martínez-Prieto, Gonzalo Navarro 0001 |
SEA | 3 |
| 2011 | Self-Indexed Grammar-Based CompressionabstractSelf-indexes aim at representing text collections in a compressed format that allows extracting arbitrary portions and also offers indexed searching on the collection. Current self-indexes are unable of fully exploiting the redundancy of highly repetitive text collections that arise in several applications. Grammar-based compression is well suited to exploit such repetitiveness. We introduce the first grammar-based self-index. It builds on Straight-Line Programs (SLPs), a rather general kind of context-free grammars. If an SLP of n rules represents a text T[1, u], then an SLP-compressed representation of T requires 2n log 2 n bits. For that same SLP, our self-index takes O(n log n) + n log 2 u bits. It extracts any text substring of length m in time O((m + h) log n), and finds occ occurrences of a pattern string of length m in time O((m(m + h) + h occ) log n), where h is the height of the parse tree of the SLP. No previous grammar representation had achieved o(n) search time. As byproducts we introduce (i) a representation of SLPs that takes 2n log 2 n(1 + o(1)) bits and efficiently supports more operations than a plain array of rules; (ii) a representation for binary relations with labels supporting various extended queries; (iii) a generalization of our self-index to grammar compressors that reduce T to a sequence of terminals and nonterminals, such as Re-Pair and LZ78. Francisco Claude, Gonzalo Navarro 0001 |
Fundam. Informaticae | 1 |
| 2011 | Untangled monotonic chains and adaptive range search
Diego Arroyuelo, Francisco Claude, Reza Dorrigiv, Stephane Durocher, Meng He 0001, Alejandro López-Ortiz, J. Ian Munro, Patrick K. Nicholson, Alejandro Salinger, Matthew Skala |
Theor. Comput. Sci. | 2 |
| 2010 | Compressed q-Gram Indexing for Highly Repetitive Biological SequencesabstractThe study of compressed storage schemes for highly repetitive sequence collections has been recently boosted by the availability of cheaper sequencing technologies and the flood of data they promise to generate. Such a storage scheme may range from the simple goal of retrieving whole individual sequences to the more advanced one of providing fast searches in the collection. In this paper we study alternatives to implement a particularly popular index, namely, the one able of finding all the positions in the collection of substrings of fixed length ($q$-grams). We introduce two novel techniques and show they constitute practical alternatives to handle this scenario. They excel particularly in two cases: when $q$ is small (up to 6), and when the collection is extremely repetitive (less than 0.01% mutations). Francisco Claude, Antonio Fariña, Miguel A. Martínez-Prieto, Gonzalo Navarro 0001 |
BIBE | 1 |
| 2010 | Fast in-memory XPath search using compressed indexesabstractA large fraction of an XML document typically consists of text data. The XPath query language allows text search via the equal, contains, and starts-with predicates. Such predicates can be efficiently implemented using a compressed self-index of the document's text nodes. Most queries, however, contain some parts querying the text of the document, plus some parts querying the tree structure. It is therefore a challenge to choose an appropriate evaluation order for a given query, which optimally leverages the execution speeds of the text and tree indexes. Here the SXSI system is introduced. It stores the tree structure of an XML document using a bit array of opening and closing brackets plus a sequence of labels, and stores the text nodes of the document using a global compressed self-index. On top of these indexes sits an XPath query engine that is based on tree automata. The engine uses fast counting queries of the text index in order to dynamically determine whether to evaluate top-down or bottom-up with respect to the tree structure. The resulting system has several advantages over existing systems: (1) on pure tree queries (without text search) such as the XPathMark queries, the SXSI system performs on par or better than the fastest known systems MonetDB and Qizx, (2) on queries that use text search, SXSI outperforms the existing systems by 1-3 orders of magnitude (depending on the size of the result set), and (3) with respect to memory consumption, SXSI outperforms all other systems for counting-only queries. Diego Arroyuelo, Francisco Claude, Sebastian Maneth, Veli Mäkinen, Gonzalo Navarro 0001, Kim Nguyen 0001, Jouni Sirén, Niko Välimäki |
ICDE | 2 |
| 2010 | Compact Rich-Functional Binary Relation Representations
Jérémy Barbay, Francisco Claude, Gonzalo Navarro 0001 |
LATIN | 2 |
| 2010 | Range Queries over Untangled Chains
Francisco Claude, J. Ian Munro, Patrick K. Nicholson |
SPIRE | 1 |
| 2010 | Fast and Compact Web Graph RepresentationsabstractCompressed graph representations, in particular for Web graphs, have become an attractive research topic because of their applications in the manipulation of huge graphs in main memory. The state of the art is well represented by the WebGraph project, where advantage is taken of several particular properties of Web graphs to offer a trade-off between space and access time. In this paper we show that the same properties can be exploited with a different and elegant technique that builds on grammar-based compression. In particular, we focus on Re-Pair and on Ziv-Lempel compression, which, although cannot reach the best compression ratios of WebGraph, achieve much faster navigation of the graph when both are tuned to use the same space. Moreover, the technique adapts well to run on secondary memory and in distributed scenarios. As a byproduct, we introduce an approximate Re-Pair version that works efficiently with severely limited main memory. Francisco Claude, Gonzalo Navarro 0001 |
ACM Trans. Web | 1 |
| 2009 | E-Breaker: Flexible, distributed environment for collaborative authoringabstractThis paper presents a system called E-Breaker for supporting small and medium size group authoring of any kind of documents following a regular structure. The system supports a decentralized model of development, thus not requiring a central repository. A set of rules for content ownership maintains the synchronization of the work among all members of the developing team which can work on- or off-line. It allows fine-grained locking of documents' content. Nelson Baloian, Francisco Claude, Roberto Konow, Sebastian Kreft |
CSCWD | 2 |
| 2009 | Untangled Monotonic Chains and Adaptive Range Search
Diego Arroyuelo, Francisco Claude, Reza Dorrigiv, Stephane Durocher, Meng He 0001, Alejandro López-Ortiz, J. Ian Munro, Patrick K. Nicholson, Alejandro Salinger, Matthew Skala |
ISAAC | 2 |
| 2009 | Practical Discrete Unit Disk Cover Using an Exact Line-Separable Algorithm
Francisco Claude, Reza Dorrigiv, Stephane Durocher, Robert Fraser, Alejandro López-Ortiz, Alejandro Salinger |
ISAAC | 1 |
| 2009 | Self-indexed Text Compression Using Straight-Line Programs
Francisco Claude, Gonzalo Navarro 0001 |
MFCS | 1 |
| 2008 | Practical Rank/Select Queries over Arbitrary Sequences
Francisco Claude, Gonzalo Navarro 0001 |
SPIRE | 1 |
| 2008 | Speeding Up Pattern Matching by Text Sampling
Francisco Claude, Gonzalo Navarro 0001, Hannu Peltola, Leena Salmela, Jorma Tarhio |
SPIRE | 1 |
| 2007 | A Fast and Compact Web Graph Representation
Francisco Claude, Gonzalo Navarro 0001 |
SPIRE | 1 |