EDBT 2026 Demo / reviewers in the wild / expert
David Dominguez-Sal
dblp:58/2864 · also David Domínguez-Sal
· DBLP profile ↗
15ranked-venue papers
4as first author
0since 2021 · last 2020
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 11 · 1 first-authorArtificial intelligence and machine learning · 3 · 1 first-authorSystems, architecture and hardware · 3 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
5 papers |
Database system architecture and tuning · 40% Data mining · 31% Information retrieval · 8% | |
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Cloud and datacenter computing · 80% Memory systems · 11% Parallel and multicore computing · 10% | |
| Software engineering, system software, and programming languages
1 paper |
Operating systems · 100% |
Topics — the 16 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Database system architecture and tuning › main-memory database
in-memory OLTP |
0.4 | 1 | 2020 | A system design for elastically scaling transaction processing engines in virtualized servers · Proc. VLDB Endow. 2020 |
Cloud and datacenter computing › resource management › cloud resource management
elastic resource management |
0.4 | 1 | 2020 | A system design for elastically scaling transaction processing engines in virtualized servers · Proc. VLDB Endow. 2020 |
Cloud and datacenter computing
virtualization |
0.4 | 1 | 2020 | A system design for elastically scaling transaction processing engines in virtualized servers · Proc. VLDB Endow. 2020 |
Database system architecture and tuning › parallel database system
shared-nothing architecture |
0.3 | 1 | 2017 | Fiber-based architecture for NFV cloud databases · Proc. VLDB Endow. 2017 |
Operating systems › resource management › process management
user-level threads |
0.3 | 1 | 2017 | Fiber-based architecture for NFV cloud databases · Proc. VLDB Endow. 2017 |
Data mining › structured data mining › graph mining
community detection |
0.2 | 1 | 2014 | High quality, scalable and parallel community detection for large real graphs · WWW 2014 |
Data mining › structured data mining
graph mining |
0.2 | 1 | 2014 | High quality, scalable and parallel community detection for large real graphs · WWW 2014 |
Data mining › structured data mining › graph mining › community detection
scalable community detection |
0.2 | 1 | 2014 | High quality, scalable and parallel community detection for large real graphs · WWW 2014 |
Indexing and storage engines › caching
cache replacement |
0.1 | 1 | 2012 | Using Evolutive Summary Counters for Efficient Cooperative Caching in Search Engines · IEEE Trans. Parallel Distributed Syst. 2012 |
Distributed and cloud data management
distributed caching |
0.1 | 1 | 2012 | Using Evolutive Summary Counters for Efficient Cooperative Caching in Search Engines · IEEE Trans. Parallel Distributed Syst. 2012 |
Information retrieval
search engines |
0.1 | 1 | 2012 | Using Evolutive Summary Counters for Efficient Cooperative Caching in Search Engines · IEEE Trans. Parallel Distributed Syst. 2012 |
Memory systems
non-volatile memory |
0.1 | 1 | 2020 | A system design for elastically scaling transaction processing engines in virtualized servers · Proc. VLDB Endow. 2020 |
Query processing and optimization › sorting
external sorting |
0.1 | 1 | 2010 | Two-way Replacement Selection · Proc. VLDB Endow. 2010 |
Cloud and datacenter computing › virtualization › network virtualization
network function virtualization |
0.1 | 1 | 2017 | Fiber-based architecture for NFV cloud databases · Proc. VLDB Endow. 2017 |
Parallel and multicore computing
parallel computing |
0.1 | 1 | 2014 | High quality, scalable and parallel community detection for large real graphs · WWW 2014 |
Parallel and multicore computing
parallel graph algorithms |
0.1 | 1 | 2014 | High quality, scalable and parallel community detection for large real graphs · WWW 2014 |
Methods — techniques the papers use, named apart from their topics
hypervisor-VM communication · 0.9NUMA-aware resource allocation · 0.9shared-nothing partitioning · 0.9fibers · 0.9parallel community detection · 0.4evolutive summary counters · 0.1ESC-summaries · 0.1merge-sort · 0.1heap · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | A system design for elastically scaling transaction processing engines in virtualized serversabstractOnline Transaction Processing (OLTP) deployments are migrating from on-premise to cloud settings in order to exploit the elasticity of cloud infrastructure which allows them to adapt to workload variations. However, cloud adaptation comes at the cost of redesigning the engine, which has led to the introduction of several, new, cloud-based transaction processing systems mainly focusing on: (i) the transaction coordination protocol, (ii) the data partitioning strategy, and, (iii) the resource isolation across multiple tenants. As a result, standalone OLTP engines cannot be easily deployed with an elastic setting in the cloud and they need to migrate to another, specialized deployment. In this paper, we focus on workload variations that can be addressed by modern multi-socket, multi-core servers and we present a system design for providing fine-grained elasticity to multi-tenant, scale-up OLTP deployments. We introduce novel components to the virtualization software stack that enable on-demand addition and removal of computing and memory resources. We provide a bi-directional, low-overhead communication stack between the virtual machine and the hypervisor, which allows the former to adapt to variations coming both from the workload and the resource availability. We show that our system achieves NUMA-aware, millisecond-level, stateful and fine-grained elasticity, while it is not intrusive to the design of state-of-the-art, in-memory OLTP engines. We evaluate our system through novel use cases demonstrating that scale-up elasticity increases resource utilization, while allowing tenants to pay for actual use of resources and not just their reservation. Angelos-Christos G. Anadiotis, Raja Appuswamy, Anastasia Ailamaki, Ilan Bronshtein, Hillel Avni, David Dominguez-Sal, Shay Goikhman, Eliezer Levy |
Proc. VLDB Endow. | 6 |
| 2017 | Fiber-based architecture for NFV cloud databasesabstractThe telco industry is gradually shifting from using monolithic software packages deployed on custom hardware to using modular virtualized software functions deployed on cloudified data centers using commodity hardware. This transformation is referred to as Network Function Virtualization (NFV). The scalability of the databases (DBs) underlying the virtual network functions is the cornerstone for reaping the benefits from the NFV transformation. This paper presents an industrial experience of applying shared-nothing techniques in order to achieve the scalability of a DB in an NFV setup. The special combination of requirements in NFV DBs are not easily met with conventional execution models. Therefore, we designed a special shared-nothing architecture that is based on cooperative multi-tasking using user-level threads (fibers). We further show that the fiber-based approach outperforms the approach built using conventional multi-threading and meets the variable deployment needs of the NFV transformation. Furthermore, fibers yield a simpler-to-maintain software and enable controlling a trade-off between long-duration computations and real-time requests. Vaidas Gasiunas, David Dominguez-Sal, Ralph Acker, Aharon Avitzur, Ilan Bronshtein, Eli Ginot, Norbert Martínez-Bazan, Alexander Nozdrin, Weijie Ou, Nir Pachter, Dima Sivov, Eliezer Levy |
Proc. VLDB Endow. | 2 |
| 2016 | Put Three and Three Together: Triangle-Driven Community DetectionabstractCommunity detection has arisen as one of the most relevant topics in the field of graph data mining due to its applications in many fields such as biology, social networks, or network traffic analysis. Although the existing metrics used to quantify the quality of a community work well in general, under some circumstances, they fail at correctly capturing such notion. The main reason is that these metrics consider the internal community edges as a set, but ignore how these actually connect the vertices of the community. We propose the Weighted Community Clustering ( WCC ), which is a new community metric that takes the triangle instead of the edge as the minimal structural motif indicating the presence of a strong relation in a graph. We theoretically analyse WCC in depth and formally prove, by means of a set of properties, that the maximization of WCC guarantees communities with cohesion and structure. In addition, we propose Scalable Community Detection (SCD) , a community detection algorithm based on WCC , which is designed to be fast and scalable on SMP machines, showing experimentally that WCC correctly captures the concept of community in social networks using real datasets. Finally, using ground-truth data, we show that SCD provides better quality than the best disjoint community detection algorithms of the state of the art while performing faster. Arnau Prat-Pérez, David Dominguez-Sal, Josep M. Brunat, Josep Lluís Larriba-Pey |
ACM Trans. Knowl. Discov. Data | 2 |
| 2014 | Massive Query Expansion by Exploiting Graph Knowledge Bases for Image RetrievalabstractAnnotation-based techniques for image retrieval suffer from sparse and short image textual descriptions. Moreover, users are often not able to describe their needs with the most appropriate keywords. This situation is a breeding ground for a vocabulary mismatch problem resulting in poor results in terms of retrieval precision. In this paper, we propose a query expansion technique for queries expressed as keywords and short natural language descriptions. We present a new massive query expansion strategy that enriches queries using a graph knowledge base by identifying the query concepts, and adding relevant synonyms and semantically related terms. We propose a topological graph enrichment technique that analyzes the network of relations among the concepts, and suggests semantically related terms by path and community detection analysis of the knowledge graph. We perform our expansions by using two versions of Wikipedia as knowledge base achieving improvements of the system's precision up to more than 27%. Joan Guisado-Gámez, David Dominguez-Sal, Josep Lluís Larriba-Pey |
ICMR | 2 |
| 2014 | High quality, scalable and parallel community detection for large real graphsabstractCommunity detection has arisen as one of the most relevant topics in the field of graph mining, principally for its applications in domains such as social or biological networks analysis. Different community detection algorithms have been proposed during the last decade, approaching the problem from different perspectives. However, existing algorithms are, in general, based on complex and expensive computations, making them unsuitable for large graphs with millions of vertices and edges such as those usually found in the real world. Arnau Prat-Pérez, David Dominguez-Sal, Josep Lluís Larriba-Pey |
WWW | 2 |
| 2012 | Shaping communities out of trianglesabstractCommunity detection has arisen as one of the most relevant topics in the field of graph data mining due to its importance in many fields such as biology, social networks or network traffic analysis. The metrics proposed to shape communities are too lax and do not consider the internal layout of the edges in the community, which lead to undesirable results. We define a new community metric called WCC. The proposed metric meets a minimum set of basic properties that guarantees communities with structure and cohesion. We experimentally show that WCC correctly quantifies the quality of communities and community partitions using real and synthetic datasets, and compare some of the most used community detection algorithms in the state of the art. Arnau Prat-Pérez, David Dominguez-Sal, Josep M. Brunat, Josep Lluís Larriba-Pey |
CIKM | 2 |
| 2012 | Efficient graph management based on bitmap indicesabstractThe increasing amount of graph like data from social networks, science and the web has grown an interest in analyzing the relationships between different entities. New specialized solutions in the form of graph databases, which are generic and able to adapt to any schema as an alternative to RDBMS, have appeared to manage attributed multigraphs efficiently. In this paper, we describe the internals of DEX graph database, which is based on a representation of the graph and its attributes as maps and bitmap structures that can be loaded and unloaded efficiently from memory. We also present the internal operations used in DEX to manipulate these structures. We show that by using these structures, DEX scales to graphs with billions of vertices and edges with very limited memory requirements. Finally, we compare our graph-oriented approach to other approaches showing that our system is better suited for out-of-core typical graph-like operations. Norbert Martínez-Bazan, Miquel Angel Aguila-Lorente, Victor Muntés-Mulero, David Dominguez-Sal, Sergio Gómez-Villamor, Josep Lluís Larriba-Pey |
IDEAS | 4 |
| 2012 | Using Evolutive Summary Counters for Efficient Cooperative Caching in Search EnginesabstractWe propose and analyze a distributed cooperative caching strategy based on the Evolutive Summary Counters (ESC), a new data structure that stores an approximated record of the data accesses in each computing node of a search engine. The ESC capture the frequency of accesses to the elements of a data collection, and the evolution of the access patterns for each node in a network of computers. The ESC can be efficiently summarized into what we call ESC-summaries to obtain approximate statistics of the document entries accessed by each computing node. We use the ESC-summaries to introduce two algorithms that manage our distributed caching strategy, one for the distribution of the cache contents, ESC-placement, and another one for the search of documents in the distributed cache, ESC-search. While the former improves the hit rate of the system and keeps a large ratio of data accesses local, the latter reduces the network traffic by restricting the number of nodes queried to find a document. We show that our cooperative caching approach outperforms state-of-the-art models in both hit rate, throughput, and location recall for multiple scenarios, i.e., different query distributions and systems with varying degrees of complexity. David Dominguez-Sal, Josep Aguilar-Saborit, Mihai Surdeanu, Josep Lluís Larriba-Pey |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2011 | Social Based Layouts for the Increase of Locality in Graph Operations
Arnau Prat-Pérez, David Dominguez-Sal, Josep Lluís Larriba-Pey |
DASFAA (1) | 2 |
| 2011 | ParallelGDB: a parallel graph database based on cache specializationabstractThe need for managing massive attributed graphs is becoming common in many areas such as recommendation systems, proteomics analysis, social network analysis or bibliographic analysis. This is making it necessary to move towards parallel systems that allow managing graph databases containing millions of vertices and edges. Previous work on distributed graph databases has focused on finding ways to partition the graph to reduce network traffic and improve execution time. However, partitioning a graph and keeping the information regarding the location of vertices might be unrealistic for massive graphs. In this paper, we propose Parallel-GDB, a new system based on specializing the local caches of any node in this system, providing a better cache hit ratio. ParallelGDB uses a random graph partitioning, avoiding complex partition methods based on the graph topology, that usually require managing extra data structures. This proposed system provides an efficient environment for distributed graph databases. Luis Barguñó, Victor Muntés-Mulero, David Dominguez-Sal, Patrick Valduriez |
IDEAS | 3 |
| 2010 | Two-way Replacement SelectionabstractThe performance of external sorting using merge sort is highly dependent on the length of the runs generated. One of the most commonly used run generation strategies is Replacement Selection (RS) because, on average, it generates runs that are twice the size of the memory available. However, the length of the runs generated by RS is downsized for data with certain characteristics, like inputs sorted inversely with respect to the desired output order. The goal of this paper is to propose and analyze two-way replacement selection (2WRS), which is a generalization of RS obtained by implementing two heaps instead of the single heap implemented by RS. The appropriate management of these two heaps allows generating runs larger than the memory available in a stable way, i.e. independent from the characteristics of the datasets. Depending on the changing characteristics of the input dataset, 2WRS assigns a new data record to one or the other heap, and grows or shrinks each heap, accommodating to the growing or decreasing tendency of the dataset. On average, 2WRS creates runs of at least the length generated by RS, and longer for datasets that combine increasing and decreasing data subsets. We tested both algorithms on large datasets with different characteristics and 2WRS achieves speedups at least similar to RS, and over 2.5 when RS fails to generate large runs. Xavier Martinez-Palau, David Dominguez-Sal, Josep Lluís Larriba-Pey |
Proc. VLDB Endow. | 2 |
| 2009 | Cache-aware load balancing vs. cooperative caching for distributed search enginesabstractIn this paper we study the performance of a distributed search engine from a data caching point of view. We compare and combine two different approaches to achieve better hit rates: (a) send the queries to the node which currently has the related data in its local memory (cache-aware load balancing), and (b) send the cached contents to the node where a query is being currently processed (cooperative caching). Furthermore, we study the best scheduling points in the query computation in which they can be reassigned to another node, and how this reassignation should be performed. Our analysis is guided by statistical tools on a real question answering system for several query distributions, which are typically found in query logs. David Dominguez-Sal, Marta Pérez-Casany, Josep Lluís Larriba-Pey |
HPCC | 1 |
| 2008 | Cache-aware load balancing for question answeringabstractThe need for high performance and throughput Question Answering (QA) systems demands for their migration to distributed environments. However, even in such cases it is necessary to provide the distributed system with cooperative caches and load balancing facilities in order to achieve the desired goals. Until now, the literature on QA has not considered such a complex system as a whole. Currently, the load balancer regulates the assignment of tasks based only on the CPU and I/O loads without considering the status of the system cache. David Dominguez-Sal, Mihai Surdeanu, Josep Aguilar-Saborit, Josep Lluís Larriba-Pey |
CIKM | 1 |
| 2007 | A Multi-layer Collaborative Cache for Question Answering
David Dominguez-Sal, Josep Lluís Larriba-Pey, Mihai Surdeanu |
Euro-Par | 1 |
| 2006 | Design and performance analysis of a factoid question answering system for spontaneous speech transcriptionsabstractThis paper introduces a QA designed from scratch to handle speech transcriptions. The system’s strength is achieved by analyz-ing the speech transcriptions with a mix of IR-oriented methodolo-gies and a small number of robust NLP components. We evaluate the system on transcriptions of spontaneous speech from several 1-hour-long seminars and presentations and show that the system obtains encouraging performance. Index Terms: question answering, natural language processing 1. Mihai Surdeanu, David Dominguez-Sal, Pere Comas |
INTERSPEECH | 2 |