VLDB 2026 Research / reviewers in the wild / expert
Susana Ladra
dblp:04/3672
· DBLP profile ↗
38ranked-venue papers
7as first author
10since 2021 · last 2024
0000-0003-4616-0774ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 24 · 4 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 since 2021Artificial intelligence and machine learning · 5 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021Theory of computation · 3 · 2 since 2021Security and privacy · 1Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Fed-mRMR: A lossless federated feature selection methodabstractFeature selection has become a mandatory task in data mining, due to the overwhelming amount of features in Big Data problems. To handle this high-dimensional data and avoid the well-known curse of dimensionality, we need to pre-select an optimal subset of features to reduce redundant computations. Federated learning is a machine learning technique based on training an algorithm over many decentralized edge devices holding local rather than global data on a centralized server. Application of this technique is extending to fields such as self-driving cars, medicine and health, and Industry 4.0, where data privacy is compulsory. Feature selection through federated learning is a complicated task since suboptimal features calculated by feature selection methods may be different in heterogeneous datasets from different nodes. In this paper, we propose a lossless federated version of the classic minimum redundancy maximum relevance (mRMR) feature selection algorithm, called federated mRMR (fed-mRMR), which, without losing any effectiveness of the original mRMR method, is applicable to federated learning approaches and capable of dealing with data that are not independent and identically distributed (non-IID data). Implementation can be found at: https://github.com/jorgehermo9/fed-mrmr Jorge Hermo, Verónica Bolón-Canedo, Susana Ladra |
Inf. Sci. | 3 |
| 2023 | Augmented Thresholds for MONIabstractMONI (Rossi et al., 2022) can store a pangenomic dataset T in small space and later, given a pattern P, quickly find the maximal exact matches (MEMs) of P with respect to T. In this paper we consider its one-pass version (Boucher et al., 2021), whose query times are dominated in our experiments by longest common extension (LCE) queries. We show how a small modification lets us avoid most of these queries which significantly speeds up MONI in practice while only slightly increasing its size. César Martínez-Guardiola, Nathaniel K. Brown, Fernando Silva-Coira, Dominik Köppl, Travis Gagie, Susana Ladra |
DCC | 6 |
| 2023 | Reproducible experiments with Learned Metric Index Framework
Terézia Slanináková, Matej Antol, Jaroslav Olha, Vlastislav Dohnal, Susana Ladra, Miguel A. Martínez-Prieto |
Inf. Syst. | 5 |
| 2023 | Faster compressed quadtrees
Guillermo de Bernardo, Travis Gagie, Susana Ladra, Gonzalo Navarro 0001, Diego Seco Naveiras |
J. Comput. Syst. Sci. | 3 |
| 2023 | Map algebra on raster datasets represented by compact data structuresabstractAbstract The increase in the size of data repositories has forced the design of new computing paradigms to be able to process large volumes of data in a reasonable amount of time. One of them is in‐memory computing, which advocates storing all the data in main memory to avoid the disk I/O bottleneck. Compression is one of the key technologies for this approach. For raster data, a compact data structure, called ‐raster, have been recently been proposed. It compresses raster maps while still supporting fast retrieval of a given datum or a portion of the data directly from the compressed data. ‐raster's original work introduced several queries in which it was superior to competitors. However, to be used as the basis of an in‐memory system for raster data, it is mandatory to demonstrate its efficiency when performing more complex operations such as the map algebra operators. In this work, we present the algorithms to run a set of these operators directly on ‐raster without a decompression procedure. Fernando Silva-Coira, José R. Paramá, Susana Ladra |
Softw. Pract. Exp. | 3 |
| 2023 | ViQUF: De Novo Viral Quasispecies Reconstruction Using Unitig-Based Flow NetworksabstractDuring viral infection, intrahost mutation and recombination can lead to significant evolution, resulting in a population of viruses that harbor multiple haplotypes. The task of reconstructing these haplotypes from short-read sequencing data is called viral quasispecies assembly, and it can be categorized as a multiassembly problem. We consider the de novo version of the problem, where no reference is available. We present ViQUF, a de novo viral quasispecies assembler that addresses haplotype assembly and quantification. ViQUF obtains a first draft of the assembly graph from a de Bruijn graph. Then, solving a min-cost flow over a flow network built for each pair of adjacent vertices based on their paired-end information creates an approximate paired assembly graph with suggested frequency values as edge labels, which is the first frequency estimation. Then, original haplotypes are obtained through a greedy path reconstruction guided by a min-cost flow solution in the approximate paired assembly graph. ViQUF outputs the contigs with their frequency estimations. Results on real and simulated data show that ViQUF is at least four times faster using at most half of the memory than previous methods, while maintaining, and in some cases outperforming, the high quality of assembly and frequency estimation of overlap graph-based methodologies, which are known to be more accurate but slower than the de Bruijn graph-based approaches. Borja Freire, Susana Ladra, José R. Paramá, Leena Salmela |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2022 | A practical succinct dynamic graph representation
Miguel E. Coimbra, Joana Hrotkó, Alexandre P. Francisco, Luís M. S. Russo, Guillermo de Bernardo, Susana Ladra, Gonzalo Navarro 0001 |
Inf. Comput. | 6 |
| 2022 | Memory-Efficient Assembly Using FlyeabstractIn the past decade, next-generation sequencing (NGS) enabled the generation of genomic data in a cost-effective, high-throughput manner. The most recent third-generation sequencing technologies produce longer reads; however, their error rates are much higher, which complicates the assembly process. This generates time- and space- demanding long-read assemblers. Moreover, the advances in these technologies have allowed portable and real-time DNA sequencing, enabling in-field analysis. In these scenarios, it becomes crucial to have more efficient solutions that can be executed in computers or mobile devices with minimum hardware requirements. We re-implemented an existing assembler devoted for long reads, more concretely Flye, using compressed data structures. We then compare our version with the original software using real datasets, and evaluate their performance in terms of memory requirements, execution speed, and energy consumption. The assembly results are not affected, as the core of the algorithm is maintained, but the usage of advanced compact data structures leads to improvements in memory consumption that range from 22% to 47% less space, and in the processing time, which range from being on a par up to decreases of 25%. These improvements also cause reductions in energy consumption of around 3-8%, with some datasets obtaining decreases up to 26%. Borja Freire, Susana Ladra, José R. Paramá |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2021 | Inference of viral quasispecies with a paired de Bruijn graphabstractMOTIVATION: RNA viruses exhibit a high mutation rate and thus they exist in infected cells as a population of closely related strains called viral quasispecies. The viral quasispecies assembly problem asks to characterize the quasispecies present in a sample from high-throughput sequencing data. We study the de novo version of the problem, where reference sequences of the quasispecies are not available. Current methods for assembling viral quasispecies are either based on overlap graphs or on de Bruijn graphs. Overlap graph-based methods tend to be accurate but slow, whereas de Bruijn graph-based methods are fast but less accurate. RESULTS: We present viaDBG, which is a fast and accurate de Bruijn graph-based tool for de novo assembly of viral quasispecies. We first iteratively correct sequencing errors in the reads, which allows us to use large k-mers in the de Bruijn graph. To incorporate the paired-end information in the graph, we also adapt the paired de Bruijn graph for viral quasispecies assembly. These features enable the use of long-range information in contig construction without compromising the speed of de Bruijn graph-based approaches. Our experimental results show that viaDBG is both accurate and fast, whereas previous methods are either fast or accurate but not both. In particular, viaDBG has comparable or better accuracy than SAVAGE, while being at least nine times faster. Furthermore, the speed of viaDBG is comparable to PEHaplo but viaDBG is able to retrieve also low abundance quasispecies, which are often missed by PEHaplo. AVAILABILITY AND IMPLEMENTATION: viaDBG is implemented in C++ and it is publicly available at https://bitbucket.org/bfreirec1/viadbg. All datasets used in this article are publicly available at https://bitbucket.org/bfreirec1/data-viadbg/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Borja Freire, Susana Ladra, José R. Paramá, Leena Salmela |
Bioinform. | 2 |
| 2021 | Compact structure for sparse undirected graphs based on a clique graph partition
Felipe Glaria, Cecilia Hernández, Susana Ladra, Gonzalo Navarro 0001, Lilian Salinas |
Inf. Sci. | 3 |
| 2020 | On Dynamic Succinct Graph RepresentationsabstractWe address the problem of representing dynamic graphs using k2-trees. The k2-tree data structure is one of the succinct data structures proposed for representing static graphs, and binary relations in general. It relies on compact representations of bit vectors. Hence, by relying on compact representations of dynamic bit vectors, we can also represent dynamic graphs. In this paper we follow instead the ideas by Munro et al., and we present an alternative implementation for representing dynamic graphs using k2-trees. Our experimental results show that this new implementation is competitive in practice. Miguel E. Coimbra, Alexandre P. Francisco, Luís M. S. Russo, Guillermo de Bernardo, Susana Ladra, Gonzalo Navarro 0001 |
DCC | 5 |
| 2019 | Space- and Time-Efficient Storage of LiDAR Point Clouds
Susana Ladra, Miguel Rodríguez Luaces, José R. Paramá, Fernando Silva-Coira |
SPIRE | 1 |
| 2019 | Set operations over compressed binary relations
Carlos Quijada-Fuentes, Miguel R. Penabad, Susana Ladra, Gilberto Gutiérrez 0001 |
Inf. Syst. | 3 |
| 2019 | Compact and efficient representation of general graph databases
Sandra Álvarez-García, Borja Freire, Susana Ladra, Oscar Pedreira |
Knowl. Inf. Syst. | 3 |
| 2018 | Exploiting Computation-Friendly Graph Compression Methods for Adjacency-Matrix MultiplicationabstractComputing the product of the (binary) adjacency matrix of a large graph with a real-valued vector is an important operation that lies at the heart of various graph analysis tasks, such as computing PageRank. In this paper we show that some well-known Web and social graph compression formats are computation-friendly, in the sense that they allow boosting the computation. In particular, we show that the format of Boldi and Vigna allows computing the product in time proportional to the compressed graph size. Our experimental results show speedups of at least 2 on graphs that were compressed at least 5 times with respect to the original. We show that other successful graph compression formats enjoy this property as well. Alexandre P. Francisco, Travis Gagie, Susana Ladra, Gonzalo Navarro 0001 |
DCC | 3 |
| 2018 | Efficient Processing of top-K Vector-Raster Queries Over Compressed DataabstractIn this work, we propose an efficient algorithm for retrieving K polygons of a vector dataset that overlap cells of a raster dataset, such that the K polygons are those overlapping the highest (or lowest) cell values among all polygons. Gilberto Gutiérrez 0001, Susana Ladra, Juan-Ramón López, José R. Paramá, Fernando Silva-Coira |
DCC | 2 |
| 2017 | Competitive Author Profiling Using Compression-Based StrategiesabstractAuthor profiling consists in determining some demographic attributes — such as gender, age, nationality, language, religion, and others — of an author for a given document. This task, which has applications in fields such as forensics, security, or marketing, has been approached from different areas, especially from linguistics and natural language processing, by extracting different types of features from training documents, usually content — and style-based features. In this paper we address the problem by using several compression-inspired strategies that generate different models without analyzing or extracting specific features from the textual content, making them style-oblivious approaches. We analyze the behavior of these techniques, combine them and compare them with other state-of-the-art methods. We show that they can be competitive in terms of accuracy, giving the best predictions for some domains, and they are efficient in time performance. Francisco Claude, Daniil Galaktionov, Roberto Konow, Susana Ladra, Oscar Pedreira |
Int. J. Uncertain. Fuzziness Knowl. Based Syst. | 4 |
| 2017 | Scalable and queryable compressed storage structure for raster data
Susana Ladra, José R. Paramá, Fernando Silva-Coira |
Inf. Syst. | 1 |
| 2016 | Compression-Inspired Author ProfilingabstractAuthor profiling, that is, determining the demographic attributes -such as gender, age, nationality, language, religion, and others- of an author for a given document, has been approached from different areas, especially from linguistics and natural language processing, by extracting different types of features from training documents, usually content- and style-based features.This work addresses the problem of identifying age and gender of the author of a given document with compression-inspired strategies without analysing or extracting specifc features from the textual content, making them style-oblivious approaches. Since they do not require any a priori knowledge of the linguistic properties, they are of special interest for domain where we do not have an a priori intuition of its properties, such as DNA and protein sequences, stock market data, or medical monitoring. Francisco Claude, Roberto Konow, Susana Ladra |
DCC | 3 |
| 2016 | Compact and queryable representation of raster datasetsabstractCompact data structures combine in a unique data structure a compressed representation of the data and the structures to access such data. The target is to be able to manage data directly in compressed form, and in this way, to keep data always compressed, even in main memory. With this, we obtain two benefits: we can manage larger datasets in main memory and we take advantage of a better usage of the memory hierarchy. Susana Ladra, José R. Paramá, Fernando Silva-Coira |
SSDBM | 1 |
| 2015 | Efficient Set Operations over k2-Treesabstractk2-trees have been proved successful to represent in avery compact way different kinds of binary relations, such as web graphs, RDFs or raster data. In order to be a fully functional succinct representation for these domains, the k2-tree must support all the required operations for binary relations. In their original description, the authors include how to answer some of the most relevant queries over the k2-tree. In this paper, we extend this functionality and detail the algorithms to efficiently compute the k2-tree resulting from the union, intersection, difference or complement of binary relations represented using k2-trees. Nieves R. Brisaboa, Guillermo de Bernardo, Gilberto Gutiérrez 0001, Susana Ladra, Miguel R. Penabad, Brunny Troncoso |
DCC | 4 |
| 2015 | Faster Compressed QuadtreesabstractReal-world point sets tend to be clustered, so using a machine word for each point is wasteful. In this paper we first bound the number of nodes in the quad tree for a point set in terms of the points' clustering. We then describe aqua tree data structure that uses O (1) bits per node and supports faster queries than previous structures with this property. Finally, we present experimental evidence that our structure is practical. Travis Gagie, Javier I. González-Nova, Susana Ladra, Gonzalo Navarro 0001, Diego Seco Naveiras |
DCC | 3 |
| 2014 | Compact representation of Web graphs with extended functionality
Nieves R. Brisaboa, Susana Ladra, Gonzalo Navarro 0001 |
Inf. Syst. | 2 |
| 2013 | Context-Based Algorithms for the List-Update Problem under Alternative Cost ModelsabstractThe List-Update Problem is a well studied online problem with direct applications in data compression. Although the model proposed by Sleator & Tarjan has become the standard in the field for the problem, its applicability in some domains, and in particular for compression purposes, has been questioned. In this paper, we focus on two alternative models for the problem that arguably have more practical significance than the standard model. We provide new algorithms for these models, and show that these algorithms outperform all classical algorithms under the discussed models. This is done via an empirical study of the performance of these algorithms on the reference data set for the list-update problem. The presented algorithms make use of the context-based strategies for compression, which have not been considered before in the context of the list-update problem and lead to improved compression algorithms. In addition, we study the adaptability of these algorithms to different measures of locality of reference and compressibility. Shahin Kamali, Susana Ladra, Alejandro López-Ortiz, Diego Seco Naveiras |
DCC | 2 |
| 2013 | DACs: Bringing direct access to variable-length codes
Nieves R. Brisaboa, Susana Ladra, Gonzalo Navarro 0001 |
Inf. Process. Manag. | 2 |
| 2012 | Exploiting SIMD Instructions in Current Processors to Improve Classical String Algorithms
Susana Ladra, Oscar Pedreira, José Duato, Nieves R. Brisaboa |
ADBIS | 1 |
| 2012 | Approximate all-pairs suffix/prefix overlaps
Niko Välimäki, Susana Ladra, Veli Mäkinen |
Inf. Comput. | 2 |
| 2012 | Implicit indexing of natural language text by reorganizing bytecodes
Nieves R. Brisaboa, Antonio Fariña, Susana Ladra, Gonzalo Navarro 0001 |
Inf. Retr. | 3 |
| 2011 | Practical representations for web and social graphsabstractIn this paper we focus on representing Web and social graphs. Our work is motivated by the need of mining information out of these graphs, thus our representations do not only aim at compressing the graphs, but also at supporting efficient navigation. This allows us to process bigger graphs in main memory, avoiding the slowdown brought by resorting on external memory. We first show how by just partitioning the graph and combining two existing techniques for Web graph compression, k2-trees [Brisaboa, Ladra and Navarro, SPIRE 2009] and RePair-Graph [Claude and Navarro, TWEB 2010], exploiting the fact that most links are intra-domain, we obtain the best time/space trade-off for direct and reverse navigation when compared to the state of the art. In social networks, splitting the graph to achieve a good decomposition is not easy. For this case, we explore a new proposal for indexing MPK linearizations [Maserrat and Pei, KDD 2010], which have proven to be an effective way of representing social networks in little space by exploiting common dense subgraphs. Our proposal offers better worst case bounds in space and time, and is also a competitive alternative in practice. Francisco Claude, Susana Ladra |
CIKM | 2 |
| 2010 | Approximate All-Pairs Suffix/Prefix Overlaps
Niko Välimäki, Susana Ladra, Veli Mäkinen |
CPM | 2 |
| 2010 | Evaluation of information loss for privacy preserving data mining through comparison of fuzzy partitionsabstractIn this paper, we focus on the problem of preserving the data confidentiality when sharing the data for clustering. This problem poses new challenges for novel uses of privacy preserving data mining (PPDM) techniques. Specifically, this paper considers the synthetic data generation as a way to preserve the data privacy. One of the state of the art synthetic data generators is the IPSO family of methods. It has been stated that the use of IPSO to generate synthetic data is appropriate when the user plans to apply clustering to the data. Moreover, this paper aims to associate the same property to the FCRM synthetic data generator, and at the same time, to assess the relationship between the information loss produced when generating synthetic data with FCRM and the clustering similarity between the original and synthetic data. Isaac Cano, Susana Ladra, Vicenç Torra |
FUZZ-IEEE | 2 |
| 2010 | Information Loss for Synthetic Data through Fuzzy ClusteringabstractSynthetic data generators are one of the methods used in privacy preserving data mining for ensuring the privacy of the individuals when their data are published. Synthetic data generators construct artificial data from some models obtained from the original data. Such models are mainly based on statistics and, typically, do not take into account other aspects of interest in artificial intelligence. In this paper we study whether one family of such synthetic data generators (the IPSO family) preserves the properties of the data that are of interest when users plan to apply clustering techniques. In particular, we study the effect of such synthetic data generators on fuzzy clustering. That is, we study the information loss data suffer when the original data are replaced by the synthetic ones. Susana Ladra, Vicenç Torra |
Int. J. Uncertain. Fuzziness Knowl. Based Syst. | 1 |
| 2009 | k2-Trees for Compact Web Graph Representation
Nieves R. Brisaboa, Susana Ladra, Gonzalo Navarro 0001 |
SPIRE | 2 |
| 2009 | Directly Addressable Variable-Length Codes
Nieves R. Brisaboa, Susana Ladra, Gonzalo Navarro 0001 |
SPIRE | 2 |
| 2008 | Cluster-Specific Information Loss Measures in Data Privacy: A ReviewabstractData protection mechanisms need to find a trade-off between information loss and disclosure risk. To this end, information loss and disclosure risk measures have been developed. Due to the fact that when data is published it is usual to ignore which kind of analyses a user will pursue with the data, generic information loss measures are used to analyse the impact of the perturbation method onto the data. Such generic information loss measures are defined in terms of a few general-enough statistics. Nevertheless, a more fine-grained analysis is needed for particular data uses. In this paper we provide the reader with a review of a few results on cluster-specific information loss measures. More specifically, we consider the case of using fuzzy clustering to the perturbated data. Vicenç Torra, Susana Ladra |
ARES | 2 |
| 2008 | Reorganizing compressed textabstractRecent research has demonstrated beyond doubts the benefits of compressing natural language texts using word-based statistical semistatic compression. Not only it achieves extremely competitive compression rates, but also direct search on the compressed text can be carried out faster than on the original text; indexing based on inverted lists benefits from compression as well.Such compression methods assign a variable-length codeword to each different text word. Some coding methods (Plain Huffman and Restricted Prefix Byte Codes) do not clearly mark codeword boundaries, and hence cannot be accessed at random positions nor searched with the fastest text search algorithms. Other coding methods (Tagged Huffman, End-Tagged Dense Code, or (s, c)-Dense Code) do mark codeword boundaries, achieving a self-synchronization property that enables fast search and random access, in exchange for some loss in compression effectiveness.In this paper, we show that by just performing a simple reordering of the target symbols in the compressed text (more precisely, reorganizing the bytes into a wavelet-treelike shape) and using little additional space, searching capabilities are greatly improved without a drastic impact in compression and decompression times. With this approach, all the codes achieve synchronism and can be searched fast and accessed at arbitrary points. Moreover, the reordered compressed text becomes an implicitly indexed representation of the text, which can be searched for words in time independent of the text length. That is, we achieve not only fast sequential search time, but indexed search time, for almost no extra space cost.We experiment with three well-known word-based compression techniques with different characteristics (Plain Huffman, End-Tagged Dense Code and Restricted Prefix Byte Codes), and show the searching capabilities achieved by reordering the compressed representation on several corpora. We show that the reordered versions are not only much more efficient than their classical counterparts, but also more efficient than explicit inverted indexes built on the collection, when using the same amount of space. Nieves R. Brisaboa, Antonio Fariña, Susana Ladra, Gonzalo Navarro 0001 |
SIGIR | 3 |
| 2008 | A Toponym Resolution Service Following the OGC WPS Standard
Susana Ladra, Miguel Rodríguez Luaces, Oscar Pedreira, Diego Seco Naveiras |
W2GIS | 1 |
| 2008 | On the Comparison of Generic Information Loss Measures and Cluster-Specific OnesabstractMasking methods are to protect data bases prior to their public release. They mask an original data file so that the new file ensures the privacy of data respondents. Information loss measures have been developed to evaluate in which extent the masked file diverges from the corresponding original file, and in what extent the same analyses on both files lead to the same results. Generic information loss measures ignore the intended data use of the file. These are the standard measures when data has to be released (e.g. published in the web) and there is no control on what kind of analyses users would perform. In this paper we study generic information loss measures, and we compare such measures with respect to cluster-specific ones. That is, measures specifically defined for the case in which the user will do clustering with the original data. To do so, we define such measures and then we do an extensive comparison of the two measures. The paper shows that the generic measures can cope with the information loss related to clustering. Susana Ladra, Vicenç Torra |
Int. J. Uncertain. Fuzziness Knowl. Based Syst. | 1 |