EDBT 2026 Demo / reviewers in the wild / expert
José R. Paramá
dblp:22/5490
· DBLP profile ↗
40ranked-venue papers
3as first author
7since 2021 · last 2023
0000-0002-8727-0980ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 27 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3Theory of computation · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Compacting Massive Public Transport Data
Benjamín Letelier, Nieves R. Brisaboa, Pablo Gutiérrez-Asorey, José R. Paramá, Tirso V. Rodeiro |
SPIRE | 4 |
| 2023 | Map algebra on raster datasets represented by compact data structuresabstractAbstract The increase in the size of data repositories has forced the design of new computing paradigms to be able to process large volumes of data in a reasonable amount of time. One of them is in‐memory computing, which advocates storing all the data in main memory to avoid the disk I/O bottleneck. Compression is one of the key technologies for this approach. For raster data, a compact data structure, called ‐raster, have been recently been proposed. It compresses raster maps while still supporting fast retrieval of a given datum or a portion of the data directly from the compressed data. ‐raster's original work introduced several queries in which it was superior to competitors. However, to be used as the basis of an in‐memory system for raster data, it is mandatory to demonstrate its efficiency when performing more complex operations such as the map algebra operators. In this work, we present the algorithms to run a set of these operators directly on ‐raster without a decompression procedure. Fernando Silva-Coira, José R. Paramá, Susana Ladra |
Softw. Pract. Exp. | 2 |
| 2023 | ViQUF: De Novo Viral Quasispecies Reconstruction Using Unitig-Based Flow NetworksabstractDuring viral infection, intrahost mutation and recombination can lead to significant evolution, resulting in a population of viruses that harbor multiple haplotypes. The task of reconstructing these haplotypes from short-read sequencing data is called viral quasispecies assembly, and it can be categorized as a multiassembly problem. We consider the de novo version of the problem, where no reference is available. We present ViQUF, a de novo viral quasispecies assembler that addresses haplotype assembly and quantification. ViQUF obtains a first draft of the assembly graph from a de Bruijn graph. Then, solving a min-cost flow over a flow network built for each pair of adjacent vertices based on their paired-end information creates an approximate paired assembly graph with suggested frequency values as edge labels, which is the first frequency estimation. Then, original haplotypes are obtained through a greedy path reconstruction guided by a min-cost flow solution in the approximate paired assembly graph. ViQUF outputs the contigs with their frequency estimations. Results on real and simulated data show that ViQUF is at least four times faster using at most half of the memory than previous methods, while maintaining, and in some cases outperforming, the high quality of assembly and frequency estimation of overlap graph-based methodologies, which are known to be more accurate but slower than the de Bruijn graph-based approaches. Borja Freire, Susana Ladra, José R. Paramá, Leena Salmela |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2022 | Memory-Efficient Assembly Using FlyeabstractIn the past decade, next-generation sequencing (NGS) enabled the generation of genomic data in a cost-effective, high-throughput manner. The most recent third-generation sequencing technologies produce longer reads; however, their error rates are much higher, which complicates the assembly process. This generates time- and space- demanding long-read assemblers. Moreover, the advances in these technologies have allowed portable and real-time DNA sequencing, enabling in-field analysis. In these scenarios, it becomes crucial to have more efficient solutions that can be executed in computers or mobile devices with minimum hardware requirements. We re-implemented an existing assembler devoted for long reads, more concretely Flye, using compressed data structures. We then compare our version with the original software using real datasets, and evaluate their performance in terms of memory requirements, execution speed, and energy consumption. The assembly results are not affected, as the core of the algorithm is maintained, but the usage of advanced compact data structures leads to improvements in memory consumption that range from 22% to 47% less space, and in the processing time, which range from being on a par up to decreases of 25%. These improvements also cause reductions in energy consumption of around 3-8%, with some datasets obtaining decreases up to 26%. Borja Freire, Susana Ladra, José R. Paramá |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2021 | Inference of viral quasispecies with a paired de Bruijn graphabstractMOTIVATION: RNA viruses exhibit a high mutation rate and thus they exist in infected cells as a population of closely related strains called viral quasispecies. The viral quasispecies assembly problem asks to characterize the quasispecies present in a sample from high-throughput sequencing data. We study the de novo version of the problem, where reference sequences of the quasispecies are not available. Current methods for assembling viral quasispecies are either based on overlap graphs or on de Bruijn graphs. Overlap graph-based methods tend to be accurate but slow, whereas de Bruijn graph-based methods are fast but less accurate. RESULTS: We present viaDBG, which is a fast and accurate de Bruijn graph-based tool for de novo assembly of viral quasispecies. We first iteratively correct sequencing errors in the reads, which allows us to use large k-mers in the de Bruijn graph. To incorporate the paired-end information in the graph, we also adapt the paired de Bruijn graph for viral quasispecies assembly. These features enable the use of long-range information in contig construction without compromising the speed of de Bruijn graph-based approaches. Our experimental results show that viaDBG is both accurate and fast, whereas previous methods are either fast or accurate but not both. In particular, viaDBG has comparable or better accuracy than SAVAGE, while being at least nine times faster. Furthermore, the speed of viaDBG is comparable to PEHaplo but viaDBG is able to retrieve also low abundance quasispecies, which are often missed by PEHaplo. AVAILABILITY AND IMPLEMENTATION: viaDBG is implemented in C++ and it is publicly available at https://bitbucket.org/bfreirec1/viadbg. All datasets used in this article are publicly available at https://bitbucket.org/bfreirec1/data-viadbg/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Borja Freire, Susana Ladra, José R. Paramá, Leena Salmela |
Bioinform. | 3 |
| 2021 | An index for moving objects with constant-time access to their compressed trajectoriesabstractAs the number of vehicles and devices equipped with GPS technology has grown explosively, an urgent need has arisen for time- and space-efficient data structures to represent their trajectories. The most commonly desired queries are the following: queries about an object’s trajectory, range queries, and nearest neighbor queries. In this paper, we consider that the objects can move freely and we present a new compressed data structure for storing their trajectories, based on a combination of logs and snapshots, with the logs storing sequences of the objects’ relative movements and the snapshots storing their absolute positions sampled at regular time intervals. We call our data structure ContaCT because it provides Constant- time access to Compressed Trajectories. Its logs are based on a compact partial-sums data structure that returns cumulative displacement in constant time, and allows us to compute in constant time any object’s position at any instant, enabling a speedup when processing several other queries. We have compared ContaCT experimentally with another compact data structure for trajectories, called GraCT, and with a classic spatio-temporal index, the MVR-tree. Our results show that ContaCT outperforms the MVR-tree by orders of magnitude in space and also outperforms the compressed representation in time performance. Nieves R. Brisaboa, Travis Gagie, Adrián Gómez-Brandón, Gonzalo Navarro 0001, José R. Paramá |
Int. J. Geogr. Inf. Sci. | 5 |
| 2021 | Space-efficient representations of raster time seriesabstractRaster time series, a.k.a. temporal rasters, are collections of rasters covering the same region at consecutive timestamps. These data have been used in many different applications ranging from weather forecast systems to monitoring of forest degradation or soil contamination. Many different sensors are generating this type of data, which makes such analyses possible, but also challenges the technological capacity to store and retrieve the data. In this work, we propose a space-efficient representation of raster time series that is based on Compact Data Structures (CDS). Our method uses a strategy of snapshots and logs to represent the data, in which both components are represented using CDS. We study two variants of this strategy, one with regular sampling and another one based on a heuristic that determines at which timestamps should the snapshots be created to reduce the space redundancy. We perform a comprehensive experimental evaluation using real datasets. The results show that the proposed strategy is competitive in space with alternatives based on pure data compression, while providing much more efficient query times for different types of queries. Fernando Silva-Coira, José R. Paramá, Guillermo de Bernardo, Diego Seco Naveiras |
Inf. Sci. | 2 |
| 2019 | Space- and Time-Efficient Storage of LiDAR Point Clouds
Susana Ladra, Miguel Rodríguez Luaces, José R. Paramá, Fernando Silva-Coira |
SPIRE | 3 |
| 2019 | GraCT: A Grammar-based Compressed Index for Trajectory Data
Nieves R. Brisaboa, Adrián Gómez-Brandón, Gonzalo Navarro 0001, José R. Paramá |
Inf. Sci. | 4 |
| 2018 | Efficient Processing of top-K Vector-Raster Queries Over Compressed DataabstractIn this work, we propose an efficient algorithm for retrieving K polygons of a vector dataset that overlap cells of a raster dataset, such that the K polygons are those overlapping the highest (or lowest) cell values among all polygons. Gilberto Gutiérrez 0001, Susana Ladra, Juan-Ramón López, José R. Paramá, Fernando Silva-Coira |
DCC | 4 |
| 2018 | 3DGraCT: A Grammar-Based Compressed Representation of 3D Trajectories
Nieves R. Brisaboa, Adrián Gómez-Brandón, Miguel A. Martínez-Prieto, José R. Paramá |
SPIRE | 4 |
| 2018 | Towards a Compact Representation of Temporal Rasters
Ana Cerdeira-Pena, Guillermo de Bernardo, Antonio Fariña, José R. Paramá, Fernando Silva-Coira |
SPIRE | 4 |
| 2018 | The largest empty circle with location constraints in spatial databases
Gilberto Gutiérrez 0001, Juan-Ramón López, José R. Paramá, Miguel R. Penabad |
Knowl. Inf. Syst. | 3 |
| 2018 | Scalable processing and autocovariance computation of big functional dataabstractSummary This paper presents 2 main contributions. The first is a compact representation of huge sets of functional data or trajectories of continuous‐time stochastic processes, which allows keeping the data always compressed even during the processing in main memory. It is oriented to facilitate the efficient computation of the sample autocovariance function without a previous decompression of the data set, by using only partial local decoding. The second contribution is a new memory‐efficient algorithm to compute the sample autocovariance function. The combination of the compact representation and the new memory‐efficient algorithm obtained in our experiments the following benefits. The compressed data occupy in the disk 75% of the space needed by the original data. The computation of the autocovariance function used up to 13 times less main memory, and run 65% faster than the classical method implemented, for example, in the R package. Nieves R. Brisaboa, Ricardo Cao, José R. Paramá, Fernando Silva-Coira |
Softw. Pract. Exp. | 3 |
| 2017 | Efficient Compression and Indexing of Trajectories
Nieves R. Brisaboa, Travis Gagie, Adrián Gómez-Brandón, Gonzalo Navarro 0001, José R. Paramá |
SPIRE | 5 |
| 2017 | Efficiently Querying Vector and Raster DataabstractEven though the field of spatial databases is more than 40 years old, most existing logical data models are highly focused either on spatial objects (vector data models) or spatial fields (raster data models). Furthermore, spatial index structures and query algorithms are still proposed for one of the approaches and little research work has been dedicated to index structures and query algorithms where both types of information are needed. However, due to the current high availability of different types of data, it is much more common nowadays that applications require querying vector and raster data at the same time. This paper presents a method to perform a spatial query between a vector data set represented using an R-tree and a raster data set represented using a compact and space-efficient data structure called k2-tree that saves main memory space. Therefore, the method described in this paper solves two problems: first, it can be used to evaluate queries between vector and raster data without having to convert one of the data sets to the other data model; and second, it saves main memory space, thus obtaining a more scalable system. Nieves R. Brisaboa, Guillermo de Bernardo, Gilberto Gutiérrez 0001, Miguel Rodríguez Luaces, José R. Paramá |
Comput. J. | 5 |
| 2017 | Scalable and queryable compressed storage structure for raster data
Susana Ladra, José R. Paramá, Fernando Silva-Coira |
Inf. Syst. | 2 |
| 2016 | GraCT: A Grammar Based Compressed Representation of Trajectories
Nieves R. Brisaboa, Adrián Gómez-Brandón, Gonzalo Navarro 0001, José R. Paramá |
SPIRE | 4 |
| 2016 | Compact and queryable representation of raster datasetsabstractCompact data structures combine in a unique data structure a compressed representation of the data and the structures to access such data. The target is to be able to manage data directly in compressed form, and in this way, to keep data always compressed, even in main memory. With this, we obtain two benefits: we can manage larger datasets in main memory and we take advantage of a better usage of the memory hierarchy. Susana Ladra, José R. Paramá, Fernando Silva-Coira |
SSDBM | 2 |
| 2014 | The largest empty rectangle containing only a query object in Spatial Databases
Gilberto Gutiérrez 0001, José R. Paramá, Nieves R. Brisaboa, Antonio Corral |
GeoInformatica | 2 |
| 2014 | Indexing and Self-indexing sequences of IEEE 754 double precision numbers
Antonio Fariña, Alberto Ordóñez Pereira, José R. Paramá |
Inf. Process. Manag. | 3 |
| 2012 | Indexing Sequences of IEEE 754 Double Precision NumbersabstractIn the last decades, much attention has been paid to the development of succinct data structures to store and/or index text, biological collections, source code, etc. Their success was in most cases due to handling data with a relatively small alphabet size and to typically exploit a rather skewed distribution (text) or simply the repetitiveness within the source data (source code repositories, biological sequences of similar individuals). In this work, we face the problem of dealing with collections of floating point data that typically have a large alphabet (a real number hardly ever repeats twice) and a less biased distribution. We present two solutions to store and index such collections. The first one is based on the well-known inverted index. It consumes space around the size of the original collection, providing appealing search times. The second one uses a wavelet tree, which at the expense of slower search times, obtains slightly better space consumption. Antonio Fariña, Alberto Ordóñez Pereira, José R. Paramá |
DCC | 3 |
| 2012 | Finding the Largest Empty Rectangle Containing Only a Query Point in Large Multidimensional Databases
Gilberto Gutiérrez 0001, José R. Paramá |
SSDBM | 2 |
| 2012 | Boosting Text Compression with Word-Based Statistical EncodingabstractSemistatic word-based byte-oriented compressors are known to be attractive alternatives to compress natural language texts. With compression ratios around 30–35%, they allow fast direct searching of compressed text. In this article, we reveal that these compressors have even more benefits. We show that most of the state-of-the-art compressors benefit from compressing not the original text, but the compressed representation obtained by a word-based byte-oriented statistical compressor. For example, p7zip with a dense-coding preprocessing achieves even better compression ratios and much faster compression than p7zip alone. We reach compression ratios below 17% in typical large English texts, which was obtained only by the slow prediction by partial matching compressors. Furthermore, searches perform much faster if the final compressor operates over word-based compressed text. We show that typical self-indexes also profit from our preprocessing step. They achieve much better space and time performance when indexing is preceded by a compression step. Apart from using the well-known Tagged Huffman code, we present a new suffix-free Dense-Code-based compressor that compresses slightly better. We also show how some self-indexes can handle non-suffix-free codes. As a result, the compressed/indexed text requires around 35% of the space of the original text and allows indexed searches for both words and phrases. Antonio Fariña, Gonzalo Navarro 0001, José R. Paramá |
Comput. J. | 3 |
| 2011 | Improving semistatic compression via phrase-based modeling
Nieves R. Brisaboa, Antonio Fariña, Gonzalo Navarro 0001, José R. Paramá |
Inf. Process. Manag. | 4 |
| 2010 | Dynamic lightweight text compressionabstractWe address the problem of adaptive compression of natural language text, considering the case where the receiver is much less powerful than the sender, as in mobile applications. Our techniques achieve compression ratios around 32% and require very little effort from the receiver. Furthermore, the receiver is not only lighter, but it can also search the compressed text with less work than that necessary to decompress it. This is a novelty in two senses: it breaks the usual compressor/decompressor symmetry typical of adaptive schemes, and it contradicts the long-standing assumption that only semistatic codes could be searched more efficiently than the uncompressed text. Our novel compression methods are preferable in several aspects over the existing adaptive and semistatic compressors for natural language texts. Nieves R. Brisaboa, Antonio Fariña, Gonzalo Navarro 0001, José R. Paramá |
ACM Trans. Inf. Syst. | 4 |
| 2008 | Word-Based Statistical Compressors as Natural Language Compression BoostersabstractSemistatic word-based byte-oriented compression codes are known to be attractive alternatives to compress natural language texts. With compression ratios around 30%, they allow direct pattern searching on the compressed text up to 8 times faster than on its uncompressed version. In this paper we reveal that these compressors have even more benefits. We show that most of the state-of-the-art compressors such as the block-wise bzip2, those from the Ziv-Lempel family, and the predictive ppm-based ones, can benefit from compressing not the original text, but its compressed representation obtained by a word-based byte-oriented statistical compressor. In particular, our experimental results show that using Dense-Code-based compression as a preprocessing step to classical compressors like bzip2, gzip, or ppmdi, yields several important benefits. For example, the ppm family is known for achieving the best compression ratios. With a Dense coding preprocessing, ppmdi achieves even better compression ratios (the best we know of on natural language) and much faster compression/decompression than ppmdi alone. Text indexing also profits from our preprocessing step. A compressed self-index achieves much better space and time performance when preceded by a semistatic word-based compression step. We show, for example, that the AF-FMindex coupled with Tagged Huffman coding is an attractive alternative index for natural language texts. Antonio Fariña, Gonzalo Navarro 0001, José R. Paramá |
DCC | 3 |
| 2008 | An Ontology-Based Index to Retrieve Documents with Geographic Information
Miguel Rodríguez Luaces, José R. Paramá, Oscar Pedreira, Diego Seco Naveiras |
SSDBM | 2 |
| 2008 | New adaptive compressors for natural language textabstractAbstract Semistatic byte‐oriented word‐based compression codes have been shown to be an attractive alternative to compress natural language text databases, because of the combination of speed, effectiveness, and direct searchability they offer. In particular, our recently proposed family of dense compression codes has been shown to be superior to the more traditional byte‐oriented word‐based Huffman codes in most aspects. In this paper, we focus on the problem of transmitting texts among peers that do not share the vocabulary. This is the typical scenario for adaptive compression methods. We design adaptive variants of our semistatic dense codes, showing that they are much simpler and faster than dynamic Huffman codes and reach almost the same compression effectiveness. We show that our variants have a very compelling trade‐off between compression/decompression speed, compression ratio, and search speed compared with most of the state‐of‐the‐art general compressors. Copyright © 2008 John Wiley & Sons, Ltd. Nieves R. Brisaboa, Antonio Fariña, Gonzalo Navarro 0001, José R. Paramá |
Softw. Pract. Exp. | 4 |
| 2007 | Lightweight natural language text compression
Nieves R. Brisaboa, Antonio Fariña, Gonzalo Navarro 0001, José R. Paramá |
Inf. Retr. | 4 |
| 2007 | Collecting and publishing large multiscale geographic datasetsabstractAbstract In this paper we present our experience in the development of a geographic information system that includes a large database (over 7 GB) with information about the infrastructure and facilities of the municipalities in the province of A Coruña (northwestern Spain). Three interesting aspects of the whole project are described in some detail due to their intrinsic interest for the development of this kind of system. These aspects are: (1) the design of the data model and the system architecture, which is oriented to support some advanced features such as multiscale active maps; (2) the problem of designing appropriate workflows to populate the database; and (3) the design of a Web‐based application to exploit the geographic database through a user‐friendly interface. Copyright © 2007 John Wiley & Sons, Ltd. Nieves R. Brisaboa, José Antonio Cotelo Lema, Antonio Fariña, Miguel Rodríguez Luaces, José R. Paramá, José R. R. Viqueira |
Softw. Pract. Exp. | 5 |
| 2006 | A semantic approach to optimize linear datalog programs
José R. Paramá, Nieves R. Brisaboa, Miguel R. Penabad, Ángeles Saavedra Places |
Acta Informatica | 1 |
| 2006 | The design of a Virtual Library of Emblem BooksabstractAntique documents, which undoubtedly represent our cultural heritage and can be considered a very rich source of information, are kept in many countries only on libraries with historical archives. The antiquity and fragility of such documents makes their access very restricted. Considering that nowadays the Internet is one of the most interesting places to publish any kind of information, it seems logical to use it to both preserve our cultural heritage and provide a broader access to these documents. This work presents a virtual library that stores data, transcribed texts and digitalized pages of historic Spanish documents from the 16th–18th centuries. This virtual library has two main objectives: first, by offering a set of services, including a powerful user interface to search and browse the documents, a bulletin board, a chat, or mail boxes, the virtual library is transformed into a meeting place for researchers that use emblem books as sources of information for their studies. Second, the virtual library contributes to the preservation of emblem books. We shall describe in this work the project that led to the development of the Virtual Library of Emblem Books, showing its evolution from the beginning (simple search forms and answer pages) to its current state as a virtual library, focusing on the techniques used to build an intuitive and powerful user interface. Copyright © 2006 John Wiley & Sons, Ltd. José R. Paramá, Ángeles Saavedra Places, Nieves R. Brisaboa, Miguel R. Penabad |
Softw. Pract. Exp. | 1 |
| 2005 | Efficiently decodable and searchable natural language adaptive compressionabstractWe address the problem of adaptive compression of natural language text, focusing on the case where low bandwidth is available and the receiver has little processing power, as in mobile applications. Our technique achieves compression ratios around 32% and requires very little effort from the receiver. This tradeoff, not previously achieved with alternative techniques, is obtained by breaking the usual symmetry between sender and receiver dominant in statistical adaptive compression. Moreover, we show that our technique can be adapted to avoid decompression at all in cases where the receiver only wants to detect the presence of some keywords in the document. This is useful in scenarios such as selective dissemination of information, news clipping, alert systems, text categorization, and clustering. Thanks to the asymmetry we introduce, the receiver can search the compressed text much faster than the plain text. This was previously achieved only in semistatic compression scenarios. Nieves R. Brisaboa, Antonio Fariña, Gonzalo Navarro 0001, José R. Paramá |
SIGIR | 4 |
| 2004 | Simple, Fast, and Efficient Natural Language Adaptive Compression
Nieves R. Brisaboa, Antonio Fariña, Gonzalo Navarro 0001, José R. Paramá |
SPIRE | 4 |
| 2004 | A Generic Framework for GIS Applications
Miguel Rodríguez Luaces, Nieves R. Brisaboa, José R. Paramá, José R. R. Viqueira |
W2GIS | 3 |
| 2003 | An Efficient Compression Code for Text Databases
Nieves R. Brisaboa, Eva Lorenzo Iglesias, Gonzalo Navarro 0001, José R. Paramá |
ECIR | 4 |
| 2002 | A Semantic Query Optimization Approach to Optimize Linear Datalog Programs
José R. Paramá, Nieves R. Brisaboa, Miguel R. Penabad, Ángeles Saavedra Places |
ADBIS | 1 |
| 2002 | A general procedure to check conjunctive query containment
Miguel R. Penabad, Nieves R. Brisaboa, Héctor J. Hernández, José R. Paramá |
Acta Informatica | 4 |
| 1998 | Containment of Conjunctive Queries with Built-in Predicates with Variables and Constants over any Ordered Domain
Nieves R. Brisaboa, Héctor J. Hernández, José R. Paramá, Miguel R. Penabad |
ADBIS | 3 |