Senén González

dblp:03/8612 · DBLP profile ↗
← Back
14ranked-venue papers
0as first author
2since 2021 · last 2021
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 7Theory of computation · 3 · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Systems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2021 Descriptive complexity of deterministic polylogarithmic time and space
Flavio Ferrarotti, Senén González, José Maria Turull Torres, Jan Van den Bussche, Jonni Virtema
J. Comput. Syst. Sci.2
2021 A logic for reflective ASMs
Klaus-Dieter Schewe, Flavio Ferrarotti, Senén González
Sci. Comput. Program.3
2020 To index or not to index: Time-space trade-offs for positional ranking functions in search engines
Diego Arroyuelo, Senén González, Mauricio Marín, Mauricio Oyarzún, Torsten Suel, Luis Valenzuela
Inf. Syst.2
2019 Descriptive Complexity of Deterministic Polylogarithmic Time
Flavio Ferrarotti, Senén González, José Maria Turull Torres, Jan Van den Bussche, Jonni Virtema
WoLLIC2
2019 Compressed filesystem for managing large genome collections
abstract
MOTIVATION: Genome repositories are growing faster than our storage capacities, challenging our ability to store, transmit, process and analyze them. While genomes are not very compressible individually, those repositories usually contain myriads of genomes or genome reads of the same species, thereby creating opportunities for orders-of-magnitude compression by exploiting inter-genome similarities. A useful compression system, however, cannot be only usable for archival, but it must allow direct access to the sequences, ideally in transparent form so that applications do not need to be rewritten. RESULTS: We present a highly compressed filesystem that specializes in storing large collections of genomes and reads. The system obtains orders-of-magnitude compression by using Relative Lempel-Ziv, which exploits the high similarities between genomes of the same species. The filesystem transparently stores the files in compressed form, intervening the system calls of the applications without the need to modify them. A client/server variant of the system stores the compressed files in a server, while the client's filesystem transparently retrieves and updates the data from the server. The data between client and server are also transferred in compressed form, which saves an order of magnitude network time. AVAILABILITY AND IMPLEMENTATION: The C++ source code of our implementation is available for download in https://github.com/vsepulve/relz_fs.
Gonzalo Navarro 0001, Victor Sepulveda, Mauricio Marín, Senén González
Bioinform.4
2019 BSP abstract state machines capture bulk synchronous parallel computations
Flavio Ferrarotti, Senén González, Klaus-Dieter Schewe
Sci. Comput. Program.2
2018 Efficient SPARQL Evaluation on Stratified RDF Data with Meta-data
Flavio Ferrarotti, Senén González, Klaus-Dieter Schewe
ADBIS2
2018 Hybrid compression of inverted lists for reordered document collections
Diego Arroyuelo, Mauricio Oyarzún, Senén González, Victor Sepulveda
Inf. Process. Manag.3
2017 On Fragments of Higher Order Logics that on Finite Structures Collapse to Second Order
Flavio Ferrarotti, Senén González, José Maria Turull Torres
WoLLIC2
2013 Document identifier reassignment and run-length-compressed inverted indexes for improved search performance
abstract
Text search engines are a fundamental tool nowadays. Their efficiency relies on a popular and simple data structure: the inverted indexes. Currently, inverted indexes can be represented very efficiently using index compression schemes. Recent investigations also study how an optimized document ordering can be used to assign document identifiers (docIDs) to the document database. This yields important improvements in index compression and query processing time. In this paper we follow this line of research, yet from a different perspective. We propose a docID reassignment method that allows one to focus on a given subset of inverted lists to improve their performance. We then use run-length encoding to compress these lists (as many consecutive 1s are generated). We show that by using this approach, not only the performance of the particular subset of inverted lists is improved, but also that of the whole inverted index. Our experimental results indicate a reduction of about 10% in the space usage of the whole index docID reassignment was focused. Also, decompression speed is up to 1.22 times faster if the runs must be explicitly decompressed and up to 4.58 times faster if implicit decompression of runs is allowed. Finally, we also improve the Document-at-a-Time query processing time of AND queries (by up to 12%), WAND queries (by up to 23%) and full (non-ranked) OR queries (by up to 86%).
Diego Arroyuelo, Senén González, Mauricio Oyarzún, Victor Sepulveda
SIGIR2
2012 To index or not to index: time-space trade-offs in search engines with positional ranking functions
abstract
Positional ranking functions, widely used in Web search engines, improve result quality by exploiting the positions of the query terms within documents. However, it is well known that positional indexes demand large amounts of extra space, typically about three times the space of a basic nonpositional index. Textual data, on the other hand, is needed to produce text snippets. In this paper, we study time-space trade-offs for search engines with positional ranking functions and text snippet generation. We consider both index-based and non-index based alternatives for positional data. We aim to answer the question of whether one should index positional data or not. We show that there is a wide range of practical time-space trade-offs. Moreover, we show that both position and textual data can be stored using about 71% of the space used by traditional positional indexes, with a minor increase in query time. This yields considerable space savings and outperforms, both in space and time, recent alternatives from the literature. We also propose several efficient compressed text representations for snippet generation, which are able to use about half of the space of current state-of-the-art alternatives with little impact in query processing time.
Diego Arroyuelo, Senén González, Mauricio Marín, Mauricio Oyarzún, Torsten Suel
SIGIR2
2012 Distributed search based on self-indexed compressed text
Diego Arroyuelo, Veronica Gil-Costa, Senén González, Mauricio Marín, Mauricio Oyarzún
Inf. Process. Manag.3
2010 Compressed Self-indices Supporting Conjunctive Queries on Document Collections
Diego Arroyuelo, Senén González, Mauricio Oyarzún
SPIRE2
2008 Scheduling Intersection Queries in Term Partitioned Inverted Files
Mauricio Marín, Carlos Gómez-Pantoja, Senén González, Veronica Gil-Costa
Euro-Par3