Iraklis Psaroudakis

dblp:133/9910 · DBLP profile ↗
← Back
12ranked-venue papers
7as first author
2since 2021 · last 2025
0009-0002-3500-8451ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 8 · 5 first-authorSystems, architecture and hardware · 3 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
7 papers
Query processing and optimization · 43% Database system architecture and tuning · 37% Graph data management · 16%
Computer architecture, parallel and distributed computing, and storage systems
4 papers
Memory systems · 71% Performance modeling and evaluation · 13% Distributed systems · 9%
Software engineering, system software, and programming languages
1 paper
Runtime systems and virtual machines · 100%

Topics — the 17 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Graph data management › graph query processing
distributed graph queries
0.512021
aDFS: An Almost Depth-First-Search Distributed Graph-Querying System · USENIX ATC 2021
Database system architecture and tuning
task scheduling
0.522016
Adaptive NUMA-aware data placement and task scheduling for analytical workloads in main-memory column-stores · Proc. VLDB Endow. 2016
Scaling Up Concurrent Main-Memory Column-Store Scans: Towards Adaptive NUMA-aware Data and Task Placement · Proc. VLDB Endow. 2015
Query processing and optimization › query execution
concurrent query execution
0.422014
Reactive and proactive sharing across concurrent analytical queries · SIGMOD Conference 2014
Sharing Data and Work Across Concurrent Analytical Queries · Proc. VLDB Endow. 2013
Runtime systems and virtual machines
managed runtime
0.312018
Analytics with smart arrays: adaptive and efficient language-independent data · EuroSys 2018
Memory systems
memory management
0.312018
Analytics with smart arrays: adaptive and efficient language-independent data · EuroSys 2018
Memory systems › non-uniform memory access
NUMA data placement
0.312018
Analytics with smart arrays: adaptive and efficient language-independent data · EuroSys 2018
Memory systems › data locality
cache locality
0.322015
How to stop under-utilization and love multicores · ICDE 2015
How to stop under-utilization and love multicores · SIGMOD Conference 2014
Memory systems › memory hierarchy
memory hierarchy optimization
0.322015
How to stop under-utilization and love multicores · ICDE 2015
How to stop under-utilization and love multicores · SIGMOD Conference 2014
Performance modeling and evaluation › parallel performance evaluation
multicore scalability
0.212015
How to stop under-utilization and love multicores · ICDE 2015
Query processing and optimization › multi-query optimization
query plan sharing
0.212014
Reactive and proactive sharing across concurrent analytical queries · SIGMOD Conference 2014
Query processing and optimization › multi-query optimization
query sharing
0.212014
Reactive and proactive sharing across concurrent analytical queries · SIGMOD Conference 2014
Query processing and optimization
shared computation
0.212013
Sharing Data and Work Across Concurrent Analytical Queries · Proc. VLDB Endow. 2013
Distributed systems › distributed database
distributed query processing
0.112021
aDFS: An Almost Depth-First-Search Distributed Graph-Querying System · USENIX ATC 2021
Query processing and optimization
analytical workloads
0.112016
Adaptive NUMA-aware data placement and task scheduling for analytical workloads in main-memory column-stores · Proc. VLDB Endow. 2016
Processor architecture and microarchitecture
instruction-level parallelism
0.112015
How to stop under-utilization and love multicores · ICDE 2015
High-performance computing
performance optimization at scale
0.112015
How to stop under-utilization and love multicores · ICDE 2015
Data integration and cleaning
data warehouse
0.012013
Sharing Data and Work Across Concurrent Analytical Queries · Proc. VLDB Endow. 2013

Methods — techniques the papers use, named apart from their topics

bit compression · 1.0adaptive data placement · 1.0adaptive algorithms · 0.2work sharing · 0.2sensitivity analysis · 0.2scheduling · 0.2data sharing · 0.2simultaneous pipelining · 0.2global query plans · 0.2
YearPublicationVenuePosition
2025 Serverless Elasticsearch: the Architecture Transformation from Stateful to Stateless
abstract
Elasticsearch (ES) is a distributed search and analytics engine consisting of a cluster of nodes, each hosting a disjoint subset of data. ES has a shared-nothing architecture that relies on local disks to store cluster data such as index files, transaction logs, and cluster state metadata. This stateful architecture couples compute with storage, and leads to different data tiers (e.g., hot, warm, cold, frozen) of hardware and configurations that the administrator chooses from to balance cost, performance, and high availability. In this paper, we show a new serverless architecture that decouples compute from storage. Serverless ES offloads data to an affordable, highly available cloud object store, while supporting the same APIs and read-after-write semantics. We show why and how this stateless architecture simplifies the tiers to just two: indexing and search, allowing indexing and searching practically limitless data while scaling each tier independently. We describe how we wrap index data in a custom batch commit format to the object store to decrease upload costs by up to 100x, how we batch transaction log uploads to decrease upload costs by up to 30x, and how we delete files from the object store. We experimentally show that Serverless ES can get twice the indexing throughput of (stateful) ES on comparable compute hardware by using object storage for durability instead of replication, and can scale linearly to match ingestion workloads.
Iraklis Psaroudakis, Pooya Salehi, Jason Bryan, Francisco Fernández Castaño, Brendan Cully, Ankita Kumar, Henning Andersen, Thomas Repantis
SoCC1
2021 aDFS: An Almost Depth-First-Search Distributed Graph-Querying System
Vasileios Trigonakis, Jean-Pierre Lozi, Tomás Faltín, Nicholas P. Roth, Iraklis Psaroudakis, Arnaud Delamare, Vlad Haprian, Calin Iorgulescu, Petr Koupy, Jinsoo Lee, Sungpack Hong, Hassan Chafi
USENIX ATC5
2020 CSR++: A Fast, Scalable, Update-Friendly Graph Data Structure
abstract
The graph model enables a broad range of analysis, thus graph processing is an invaluable tool in data analytics. At the heart of every graph-processing system lies a concurrent graph data structure storing the graph. Such a data structure needs to be highly efficient for both graph algorithms and queries. Due to the continuous evolution, the sparsity, and the scale-free nature of real-world graphs, graph-processing systems face the challenge of providing an appropriate graph data structure that enables both fast analytical workloads and low-memory graph mutations. Existing graph structures offer a hard trade-off between read-only performance, update friendliness, and memory consumption upon updates. In this paper, we introduce CSR++, a new graph data structure that removes these trade-offs and enables both fast read-only analytics and quick and memory-friendly mutations. CSR++ combines ideas from CSR, the fastest read-only data structure, and adjacency lists to achieve the best of both worlds. We compare CSR++ to CSR, adjacency lists from the Boost Graph Library, and LLAMA, a state-of-the-art update-friendly graph structure. In our evaluation, which is based on popular graph-processing algorithms executed over real-world graphs, we show that CSR++ remains close to CSR in read-only concurrent performance (within 10% on average), while significantly outperforming CSR (by an order of magnitude) and LLAMA (by almost 2×) with frequent updates.
Soukaina Firmli, Vasileios Trigonakis, Jean-Pierre Lozi, Iraklis Psaroudakis, Alexander Weld, Dalila Chiadmi, Sungpack Hong, Hassan Chafi
OPODIS4
2018 Analytics with smart arrays: adaptive and efficient language-independent data
abstract
This paper introduces smart arrays, an abstraction for providing adaptive and efficient language-independent data storage. Their smart functionalities include NUMA-aware data placement across sockets and bit compression. We show how our single C++ implementation can be used efficiently from both native C++ and compiled Java code. We experimentally evaluate smart arrays on a diverse set of C++ and Java analytics workloads. Further, we show how their smart functionalities affect performance and lead to differences in hardware resource demands on multicore machines, motivating the need for adaptivity. We observe that smart arrays can significantly decrease the memory space requirements of analytics workloads, and improve their performance by up to 4x. Smart arrays are the first step towards general smart collections with various smart functionalities that enable the consumption of hardware resources to be traded-off against one another.
Iraklis Psaroudakis, Stefan Kaestle, Matthias Grimmer, Jean-Pierre Lozi, Tim Harris 0001
EuroSys1
2016 Adaptive NUMA-aware data placement and task scheduling for analytical workloads in main-memory column-stores
abstract
Non-uniform memory access (NUMA) architectures pose numerous performance challenges for main-memory column-stores in scaling up analytics on modern multi-socket multi-core servers. A NUMA-aware execution engine needs a strategy for data placement and task scheduling that prefers fast local memory accesses over remote memory accesses, and avoids an imbalance of resource utilization, both CPU and memory bandwidth, across sockets. State-of-the-art systems typically use a static strategy that always partitions data across sockets, and always allows inter-socket task stealing. In this paper, we show that adapting data placement and task stealing to the workload can improve throughput by up to a factor of 4 compared to a static approach. We focus on highly concurrent workloads dominated by operators working on a single table or table group (copartitioned tables). Our adaptive data placement algorithm tracks the resource utilization of tasks, partitions of tables and table groups, and sockets. When a utilization imbalance across sockets is detected, the algorithm corrects it by moving or repartitioning tables. Also, inter-socket task stealing is dynamically disabled for memory-intensive tasks that could otherwise hurt performance.
Iraklis Psaroudakis, Tobias Scheuer, Norman May, Abdelkader Sellami, Anastasia Ailamaki
Proc. VLDB Endow.1
2015 How to stop under-utilization and love multicores
abstract
Hardware trends oblige software to overcome three major challenges against systems scalability: (1) taking advantage of the implicit/vertical parallelism within a core that is enabled through the aggressive micro-architectural features, (2) exploiting the explicit/horizontal parallelism provided by multicores, and (3) achieving predictively efficient execution despite the variability in communication latencies among cores on multisocket multicores. In this three hour tutorial, we shed light on the above three challenges and survey recent proposals to alleviate them. The first part of the tutorial describes the instruction- and data-level parallelism opportunities in a core coming from the hardware and software side. In addition, it examines the sources of under-utilization in a modern processor and presents insights and hardware/software techniques to better exploit the micro-architectural resources of a processor by improving cache locality at the right level of the memory hierarchy. The second part focuses on the scalability bottlenecks of database applications at the level of multicore and multisocket multicore architectures. It first presents a systematic way of eliminating such bottlenecks in online transaction processing workloads, which is based on minimizing unbounded communication, and shows several techniques that minimize bottlenecks in major components of database management systems. Then, it demonstrates the data and work sharing opportunities for analytical workloads, and reviews advanced scheduling mechanisms that are aware of non-uniform memory accesses and alleviate bandwidth saturation.
Anastasia Ailamaki, Erietta Liarou, Pinar Tözün, Danica Porobic, Iraklis Psaroudakis
ICDE5
2015 Extending database task schedulers for multi-threaded application code
abstract
Modern databases can run application logic defined in stored procedures inside the database server to improve application speed. The SQL standard specifies how to call external stored routines implemented in programming languages, such as C, C++, or JAVA, to complement declarative SQL-based application logic. This is beneficial for scientific and analytical algorithms because they are usually too complex to be implemented entirely in SQL. At the same time, database applications like matrix calculations or data mining algorithms benefit from multi-threading to parallelize compute-intensive operations. Multi-threaded application code, however, introduces a resource competition between the threads of applications and the threads of the database task scheduler. In this paper, we show that multi-threaded application code can render the database's workload scheduling ineffective and decrease the core throughput of the database by up to 50%. We present a general approach to address this issue by integrating shared memory programming solutions into the task schedulers of databases. In particular, we describe the integration of OpenMP into databases. We implement and evaluate our approach using SAP HANA. Our experiments show that our integration does not introduce overhead, and can improve the throughput of core database operations by up to 15%.
Florian Wolf 0002, Iraklis Psaroudakis, Norman May, Anastasia Ailamaki, Kai-Uwe Sattler
SSDBM2
2015 Scaling Up Concurrent Main-Memory Column-Store Scans: Towards Adaptive NUMA-aware Data and Task Placement
abstract
Main-memory column-stores are called to efficiently use modern non-uniform memory access (NUMA) architectures to service concurrent clients on big data. The efficient usage of NUMA architectures depends on the data placement and scheduling strategy of the column-store. Most column-stores choose a static strategy that involves partitioning all data across the NUMA architecture, and employing a stealing-based task scheduler. In this paper, we implement different strategies for data placement and task scheduling for the case of concurrent scans. We compare these strategies with an extensive sensitivity analysis. Our most significant findings include that unnecessary partitioning can hurt throughput by up to 70%, and that stealing memory-intensive tasks can hurt throughput by up to 58%. Based on our analysis, we envision a design that adapts the data placement and task scheduling strategy to the workload.
Iraklis Psaroudakis, Tobias Scheuer, Norman May, Abdelkader Sellami, Anastasia Ailamaki
Proc. VLDB Endow.1
2014 Dynamic fine-grained scheduling for energy-efficient main-memory queries
abstract
Power and cooling costs are some of the highest costs in data centers today, which make improvement in energy efficiency crucial. Energy efficiency is also a major design point for chips that power whole ranges of computing devices. One important goal in this area is energy proportionality, arguing that the system's power consumption should be proportional to its performance. Currently, a major trend among server processors, which stems from the design of chips for mobile devices, is the inclusion of advanced power management techniques, such as dynamic voltage-frequency scaling, clock gating, and turbo modes.
Iraklis Psaroudakis, Thomas Kissinger, Danica Porobic, Thomas Ilsche, Erietta Liarou, Pinar Tözün, Anastasia Ailamaki, Wolfgang Lehner
DaMoN1
2014 How to stop under-utilization and love multicores
abstract
Designing scalable database management systems on modern hardware has been a challenge for almost a decade. Hardware trends oblige software to overcome three major challenges against systems scalability: (1) Exploiting the abundant thread-level parallelism provided by multicores, (2) Achieving predictively efficient execution despite the variability in communication latencies among cores on multisocket multicores, and (3) Taking advantage of the aggressive micro-architectural features. In this tutorial, we shed light on the above three challenges and survey recent proposals to alleviate them. First, we present a systematic way of eliminating scalability bottlenecks based on minimizing unbounded communication and show several techniques that minimize bottlenecks in major components of database management systems. In addition, we demonstrate methods to parallelize major database operations. Then, we analyze the problems that arise from the non-uniform nature of communication latencies on modern multisockets and ways to address them for systems that already scale well on multicores. Finally, we examine the sources of under-utilization within a modern processor and present insights and techniques to better exploit the micro-architectural resources of a processor by improving cache locality at the right level of the memory hierarchy.
Anastasia Ailamaki, Erietta Liarou, Pinar Tözün, Danica Porobic, Iraklis Psaroudakis
SIGMOD Conference5
2014 Reactive and proactive sharing across concurrent analytical queries
abstract
Today an ever increasing amount of data is collected and analyzed by researchers, businesses, and scientists in data warehouses (DW). In addition to the data size, the number of users and applications querying data grows exponentially. The increasing concurrency is itself a challenge in query execution, but also introduces an opportunity favoring synergy between concurrent queries. Traditional execution engines of DW follows a query-centric approach, where each query is optimized and executed independently. On the other hand, workloads with increased concurrency have several queries with common parts of data and work, creating the opportunity for sharing among concurrent queries. Sharing can be reactive to the inherently existing sharing opportunities, or proactive by redesigning query operators to maximize the sharing opportunities. This demonstration showcases the impact of proactive and reactive sharing by comparing and integrating representative state-of-the-art techniques: Simultaneous Pipelining (SP), for reactive sharing, which shares intermediate results of common sub-plans, and Global Query Plans (GQP) for proactive sharing, which build and evaluate a single query plan with shared operators. We visually demonstrate, in an interactive interface, the behavior of both sharing approaches on top of a state-of-the-art storage engine using the original prototypes. We show that pull-based sharing for SP eliminates the serialization point imposed by the original push-based approach. Then, we compare, through a sensitivity analysis, the performance of SP and GQP. Finally, we show that SP can improve the performance of GQP for a query mix with common sub-plans.
Iraklis Psaroudakis, Manos Athanassoulis, Matthaios Olma, Anastasia Ailamaki
SIGMOD Conference1
2013 Sharing Data and Work Across Concurrent Analytical Queries
abstract
Today's data deluge enables organizations to collect massive data, and analyze it with an ever-increasing number of concurrent queries. Traditional data warehouses (DW) face a challenging problem in executing this task, due to their query-centric model: each query is optimized and executed independently. This model results in high contention for resources. Thus, modern DW depart from the query-centric model to execution models involving sharing of common data and work. Our goal is to show when and how a DW should employ sharing. We evaluate experimentally two sharing methodologies, based on their original prototype systems, that exploit work sharing opportunities among concurrent queries at run-time: Simultaneous Pipelining (SP), which shares intermediate results of common sub-plans, and Global Query Plans (GQP), which build and evaluate a single query plan with shared operators. First, after a short review of sharing methodologies, we show that SP and GQP are orthogonal techniques. SP can be applied to shared operators of a GQP, reducing response times by 20%-48% in workloads with numerous common sub-plans. Second, we corroborate previous results on the negative impact of SP on performance for cases of low concurrency. We attribute this behavior to a bottleneck caused by the push-based communication model of SP. We show that pull-based communication for SP eliminates the overhead of sharing altogether for low concurrency, and scales better on multi-core machines than push-based SP, further reducing response times by 82%-86% for high concurrency. Third, we perform an experimental analysis of SP, GQP and their combination, and show when each one is beneficial. We identify a trade-off between low and high concurrency. In the former case, traditional query-centric operators with SP perform better, while in the latter case, GQP with shared operators enhanced by SP give the best results.
Iraklis Psaroudakis, Manos Athanassoulis, Anastasia Ailamaki
Proc. VLDB Endow.1