Manos Karpathiotakis

dblp:41/9728 · DBLP profile ↗
← Back
20ranked-venue papers
5as first author
1since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 19 · 4 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
11 papers
Query processing and optimization · 44% Data integration and cleaning · 27% Indexing and storage engines · 8%
Computer architecture, parallel and distributed computing, and storage systems
3 papers
Distributed systems · 32% Cloud and datacenter computing · 32% GPUs and heterogeneous computing · 28%
Software engineering, system software, and programming languages
1 paper
Runtime systems and virtual machines · 100%

Topics — the 22 heaviest of 28, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Query processing and optimization
query execution
0.822020
Cleaning Denial Constraint Violations through Relaxation · SIGMOD Conference 2020
HetExchange: Encapsulating heterogeneous CPU-GPU parallelism in JIT compiled engines · Proc. VLDB Endow. 2019
Data integration and cleaning › data preprocessing
data cleaning
0.722020
Cleaning Denial Constraint Violations through Relaxation · SIGMOD Conference 2020
CleanM: An Optimizable Query Language for Unified Scale-Out Data Cleaning · Proc. VLDB Endow. 2017
Query processing and optimization › query execution
in-situ query processing
0.722020
Adaptive partitioning and indexing for in situ query processing · VLDB J. 2020
Slalom: Coasting Through Raw Data via Adaptive Partitioning and Indexing · Proc. VLDB Endow. 2017
Data integration and cleaning › data preprocessing › data cleaning
query-driven data cleaning
0.412020
Query-driven Repair of Functional Dependency Violations · ICDE 2020
Distributed and cloud data management
data partitioning
0.422020
Slalom: Coasting Through Raw Data via Adaptive Partitioning and Indexing · Proc. VLDB Endow. 2017
Adaptive partitioning and indexing for in situ query processing · VLDB J. 2020
Query processing and optimization › join processing › join algorithms
hash join
0.412019
Hardware-Conscious Hash-Joins on GPUs · ICDE 2019
Query processing and optimization
parallel query processing
0.412019
HetExchange: Encapsulating heterogeneous CPU-GPU parallelism in JIT compiled engines · Proc. VLDB Endow. 2019
Runtime systems and virtual machines › dynamic compilation
just-in-time compilation
0.412019
HetExchange: Encapsulating heterogeneous CPU-GPU parallelism in JIT compiled engines · Proc. VLDB Endow. 2019
GPUs and heterogeneous computing
GPU query processing
0.412019
Hardware-Conscious Hash-Joins on GPUs · ICDE 2019
GPUs and heterogeneous computing
heterogeneous query processing
0.412019
HetExchange: Encapsulating heterogeneous CPU-GPU parallelism in JIT compiled engines · Proc. VLDB Endow. 2019
Indexing and storage engines
adaptive indexing
0.312017
Slalom: Coasting Through Raw Data via Adaptive Partitioning and Indexing · Proc. VLDB Endow. 2017
Indexing and storage engines
caching
0.312017
ReCache: Reactive Caching for Fast Analytics over Heterogeneous Data · Proc. VLDB Endow. 2017
Data models and query languages
query language design
0.312017
CleanM: An Optimizable Query Language for Unified Scale-Out Data Cleaning · Proc. VLDB Endow. 2017
Storage systems
data placement
0.312025
Scribe: How Meta transports zettabytes per day in real time · Proc. VLDB Endow. 2025
Query processing and optimization › query compilation
code generation for query execution
0.212016
Fast Queries Over Heterogeneous Data Through Engine Customization · Proc. VLDB Endow. 2016
Query processing and optimization › query execution
raw data querying
0.212014
Adaptive Query Processing on RAW Data · Proc. VLDB Endow. 2014
Data models and query languages › multidimensional database
array DBMS
0.112012
TELEIOS: A Database-Powered Virtual Earth Observatory · Proc. VLDB Endow. 2012
Knowledge graphs
semantic web
0.112012
TELEIOS: A Database-Powered Virtual Earth Observatory · Proc. VLDB Endow. 2012
Data mining
exploratory data analysis
0.112020
Cleaning Denial Constraint Violations through Relaxation · SIGMOD Conference 2020
Distributed and cloud data management
scale-out
0.112017
CleanM: An Optimizable Query Language for Unified Scale-Out Data Cleaning · Proc. VLDB Endow. 2017
Database system architecture and tuning › database system implementation
query engine design
0.112016
Fast Queries Over Heterogeneous Data Through Engine Customization · Proc. VLDB Endow. 2016
Data integration and cleaning
heterogeneous data sources
0.112014
Adaptive Query Processing on RAW Data · Proc. VLDB Endow. 2014

Methods — techniques the papers use, named apart from their topics

exchange operator · 1.1JIT compilation · 1.1probabilistic repair · 0.9opportunistic data placement · 0.9multi-hop write path · 0.9adaptive replica placement · 0.9radix partitioning · 0.8hardware-conscious join design · 0.8relaxation · 0.4query plan weaving · 0.4adaptive indexing · 0.4multi-level query optimization · 0.3lightweight monitoring · 0.3stSPARQL · 0.1stRDF · 0.1column store · 0.1SciQL · 0.1
YearPublicationVenuePosition
2025 Scribe: How Meta transports zettabytes per day in real time
abstract
Millions of web servers and a multitude of applications are producing ever-increasing amounts of data in real time at Meta. Regardless of how data is generated and how it is processed, there is a need for infrastructure that can accommodate the transport of arbitrarily large data streams from their generation location to their processing location with low latency. This paper presents Scribe, a multi-tenant message queue service that natively supports the requirements of Meta's data-intensive applications, ingesting > 15 TB/s and serving > 110 TB/s to its consumers. Scribe relies on a multi-hop write path and opportunistic data placement to maximise write availability, whereas its read path adapts replica placement and representation based on the incoming workload as a means to minimise resource consumption for both Scribe and its downstreams. The wide range of Scribe use cases can pick from a range of offered guarantees, based on the trade-offs favourable for each one.
Manos Karpathiotakis, Vlassios Rizopoulos, Artem Gelun, Tiziano Carotti, Hazem Nada, Basri Kahveci, Yuri Dolgov
Proc. VLDB Endow.1
2020 Query-driven Repair of Functional Dependency Violations
abstract
Data cleaning is a time-consuming process that depends on the data analysis that users perform. Existing solutions treat data cleaning as a separate offline process that takes place before analysis begins. Applying data cleaning before analysis assumes a priori knowledge of the inconsistencies and the query workload, thereby requiring effort on understanding and cleaning the data that is unnecessary for the analysis. We propose an approach that performs probabilistic repair of functional dependency violations on-demand, driven by the exploratory analysis that users perform. We introduce Daisy, a system that seamlessly integrates data cleaning into the analysis by relaxing query results. Daisy executes analytical query-workloads over dirty data by weaving cleaning operators into the query plan. Our evaluation shows that Daisy adapts to the workload and outperforms traditional offline cleaning on both synthetic and real-world workloads.
Stella Giannakopoulou, Manos Karpathiotakis, Anastasia Ailamaki
ICDE2
2020 Cleaning Denial Constraint Violations through Relaxation
abstract
Data cleaning is a time-consuming process that depends on the data analysis that users perform. Existing solutions treat data cleaning as a separate offline process that takes place before analysis begins. Applying data cleaning before analysis assumes a priori knowledge of the inconsistencies and the query workload, thereby requiring effort on understanding and cleaning the data that is unnecessary for the analysis. We propose an approach that performs probabilistic repair of denial constraint violations on-demand, driven by the exploratory analysis that users perform. We introduce Daisy, a system that seamlessly integrates data cleaning into the analysis by relaxing query results. Daisy executes analytical query-workloads over dirty data by weaving cleaning operators into the query plan. Our evaluation shows that Daisy adapts to the workload and outperforms traditional offline cleaning on both synthetic and real-world workloads.
Stella Giannakopoulou, Manos Karpathiotakis, Anastasia Ailamaki
SIGMOD Conference2
2020 Adaptive partitioning and indexing for in situ query processing
Matthaios Olma, Manos Karpathiotakis, Ioannis Alagiannis, Manos Athanassoulis, Anastasia Ailamaki
VLDB J.2
2019 Hardware-Conscious Hash-Joins on GPUs
abstract
Traditionally, analytical database engines have used task parallelism provided by modern multisocket multicore CPUs for scaling query execution. Over the past few years, GPUs have started gaining traction as accelerators for processing analytical queries due to their massively data-parallel nature and high memory bandwidth. Recent work on designing join algorithms for CPUs has shown that carefully tuned join implementations that exploit underlying hardware can outperform naive, hardware-oblivious counterparts and provide excellent performance on modern multicore servers. However, there has been no such systematic analysis of hardware-conscious join algorithms for GPUs that systematically explores the dimensions of partitioning (partitioned versus non-partitioned joins), data location (data fitting and not fitting in GPU device memory), and access pattern (skewed versus uniform). In this paper, we present the design and implementation of a family of novel, partitioning-based GPU-join algorithms that are tuned to exploit various GPU hardware characteristics for working around the two main limitations of GPUs–limited memory capacity and slow PCIe interface. Using a thorough evaluation, we show that: i) hardware-consciousness plays a key role in GPU joins similar to CPU joins and our join algorithms can process 1 Billion tuples/second even if no data is GPU resident, ii) radix partitioning-based GPU joins that are tuned to exploit GPU hardware can substantially outperform non-partitioned hash joins, iii) hardware-conscious GPU joins can effectively overcome GPU limitations and match, or even outperform, state-of-the-art CPU joins.
Panagiotis Sioulas, Periklis Chrysogelos, Manos Karpathiotakis, Raja Appuswamy, Anastasia Ailamaki
ICDE3
2019 HetExchange: Encapsulating heterogeneous CPU-GPU parallelism in JIT compiled engines
abstract
Modern server hardware is increasingly heterogeneous as hardware accelerators, such as GPUs, are used together with multicore CPUs to meet the computational demands of modern data analytics work-loads. Unfortunately, query parallelization techniques used by analytical database engines are designed for homogeneous multicore servers, where query plans are parallelized across CPUs to process data stored in cache coherent shared memory. Thus, these techniques are unable to fully exploit available heterogeneous hardware, where one needs to exploit task-parallelism of CPUs and data-parallelism of GPUs for processing data stored in a deep, non-cache-coherent memory hierarchy with widely varying access latencies and bandwidth. In this paper, we introduce HetExchange-a parallel query execution framework that encapsulates the heterogeneous parallelism of modern multi-CPU-multi-GPU servers and enables the parallelization of (pre-)existing sequential relational operators. In contrast to the interpreted nature of traditional Exchange, HetExchange is designed to be used in conjunction with JIT compiled engines in order to allow a tight integration with the proposed operators and generation of efficient code for heterogeneous hardware. We validate the applicability and efficiency of our design by building a prototype that can operate over both CPUs and GPUs, and enables its operators to be parallelism- and data-location-agnostic. In doing so, we show that efficiently exploiting CPU-GPU parallelism can provide 2.8x and 6.4x improvement in performance compared to state-of-the-art CPU-based and GPU-based DBMS.
Periklis Chrysogelos, Manos Karpathiotakis, Raja Appuswamy, Anastasia Ailamaki
Proc. VLDB Endow.2
2017 The Case For Heterogeneous HTAP
Raja Appuswamy, Manos Karpathiotakis, Danica Porobic, Anastasia Ailamaki
CIDR2
2017 No data left behind: real-time insights from a complex data ecosystem
abstract
The typical enterprise data architecture consists of several actively updated data sources (e.g., NoSQL systems, data warehouses), and a central data lake such as HDFS, in which all the data is periodically loaded through ETL processes. To simplify query processing, state-of-the-art data analysis approaches solely operate on top of the local, historical data in the data lake, and ignore the fresh tail end of data that resides in the original remote sources. However, as many business operations depend on real-time analytics, this approach is no longer viable. The alternative is hand-crafting the analysis task to explicitly consider the characteristics of the various data sources and identify optimization opportunities, rendering the overall analysis non-declarative and convoluted.
Manos Karpathiotakis, Avrilia Floratou, Fatma Özcan 0001, Anastasia Ailamaki
SoCC1
2017 ReCache: Reactive Caching for Fast Analytics over Heterogeneous Data
abstract
As data continues to be generated at exponentially growing rates in heterogeneous formats, fast analytics to extract meaningful information is becoming increasingly important. Systems widely use in-memory caching as one of their primary techniques to speed up data analytics. However, caches in data analytics systems cannot rely on simple caching policies and a fixed data layout to achieve good performance. Different datasets and workloads require different layouts and policies to achieve optimal performance. This paper presents ReCache, a cache-based performance accelerator that is reactive to the cost and heterogeneity of diverse raw data formats. Using timing measurements of caching operations and selection operators in a query plan, ReCache accounts for the widely varying costs of reading, parsing, and caching data in nested and tabular formats. Combining these measurements with information about frequently accessed data fields in the workload, ReCache automatically decides whether a nested or relational column-oriented layout would lead to better query performance. Furthermore, ReCache keeps track of commonly utilized operators to make informed cache admission and eviction decisions. Experiments on synthetic and real-world datasets show that our caching techniques decrease caching overhead for individual queries by an average of 59%. Furthermore, over the entire workload, ReCache reduces execution time by 19-75% compared to existing techniques.
Tahir Azim, Manos Karpathiotakis, Anastasia Ailamaki
Proc. VLDB Endow.2
2017 CleanM: An Optimizable Query Language for Unified Scale-Out Data Cleaning
abstract
Data cleaning has become an indispensable part of data analysis due to the increasing amount of dirty data. Data scientists spend most of their time preparing dirty data before it can be used for data analysis. At the same time, the existing tools that attempt to automate the data cleaning procedure typically focus on a specific use case and operation. Still, even such specialized tools exhibit long running times or fail to process large datasets. Therefore, from a user's perspective, one is forced to use a different, potentially inefficient tool for each category of errors. This paper addresses the coverage and efficiency problems of data cleaning. It introduces CleanM ( pronounced clean'em ), a language which can express multiple types of cleaning operations. CleanM goes through a three-level translation process for optimization purposes; a different family of optimizations is applied in each abstraction level. Thus, CleanM can express complex data cleaning tasks, optimize them in a unified way, and deploy them in a scaleout fashion. We validate the applicability of CleanM by using it on top of CleanDB, a newly designed and implemented framework which can query heterogeneous data. When compared to existing data cleaning solutions, CleanDB a) covers more data corruption cases, b) scales better, and can handle cases for which its competitors are unable to terminate, and c) uses a single interface for querying and for data cleaning.
Stella Giannakopoulou, Manos Karpathiotakis, Benjamin Gaidioz, Anastasia Ailamaki
Proc. VLDB Endow.2
2017 Slalom: Coasting Through Raw Data via Adaptive Partitioning and Indexing
abstract
The constant flux of data and queries alike has been pushing the boundaries of data analysis systems. The increasing size of raw data files has made data loading an expensive operation that delays the data-to-insight time. Hence, recent in-situ query processing systems operate directly over raw data, alleviating the loading cost. At the same time, analytical workloads have increasing number of queries. Typically, each query focuses on a constantly shifting -- yet small -- range. Minimizing the workload latency, now, requires the benefits of indexing in in-situ query processing. In this paper, we present Slalom, an in-situ query engine that accommodates workload shifts by monitoring user access patterns. Slalom makes on-the-fly partitioning and indexing decisions, based on information collected by lightweight monitoring. Slalom has two key components: (i) an online partitioning and indexing scheme, and (ii) a partitioning and indexing tuner tailored for in-situ query engines. When compared to the state of the art, Slalom offers performance benefits by taking into account user query patterns to (a) logically partition raw data files and (b) build for each partition lightweight partition-specific indexes. Due to its lightweight and adaptive nature, Slalom achieves efficient accesses to raw data with minimal memory consumption. Our experimentation with both micro-benchmarks and real-life workloads shows that Slalom outperforms state-of-the-art in-situ engines (3 -- 10×), and achieves comparable query response times with fully indexed DBMS, offering much lower (∼ 3×) cumulative query execution times for query workloads with increasing size and unpredictable access patterns.
Matthaios Olma, Manos Karpathiotakis, Ioannis Alagiannis, Manos Athanassoulis, Anastasia Ailamaki
Proc. VLDB Endow.2
2016 Fast Queries Over Heterogeneous Data Through Engine Customization
abstract
Industry and academia are continuously becoming more data-driven and data-intensive, relying on the analysis of a wide variety of heterogeneous datasets to gain insights. The different data models and formats pose a significant challenge on performing analysis over a combination of diverse datasets. Serving all queries using a single, general-purpose query engine is slow. On the other hand, using a specialized engine for each heterogeneous dataset increases complexity: queries touching a combination of datasets require an integration layer over the different engines. This paper presents a system design that natively supports heterogeneous data formats and also minimizes query execution times. For multi-format support, the design uses an expressive query algebra which enables operations over various data models. For minimal execution times, it uses a code generation mechanism to mimic the system and storage most appropriate to answer a query fast. We validate our design by building Proteus, a query engine which natively supports queries over CSV, JSON, and relational binary data, and which specializes itself to each query, dataset, and workload via code generation. Proteus outperforms state-of-the-art open-source and commercial systems on both synthetic and real-world workloads without being tied to a single data model or format, all while exposing users to a single query interface.
Manos Karpathiotakis, Ioannis Alagiannis, Anastasia Ailamaki
Proc. VLDB Endow.1
2015 Just-In-Time Data Virtualization: Lightweight Data Management with ViDa
Manos Karpathiotakis, Ioannis Alagiannis, Thomas Heinis, Miguel Branco, Anastasia Ailamaki
CIDR1
2015 Sextant: Visualizing time-evolving linked geospatial data
Charalampos Nikolaou, Kallirroi Dogani, Konstantina Bereta, George Garbis, Manos Karpathiotakis, Kostis Kyzirakos, Manolis Koubarakis
J. Web Semant.5
2014 Adaptive Query Processing on RAW Data
abstract
Database systems deliver impressive performance for large classes of workloads as the result of decades of research into optimizing database engines. High performance, however, is achieved at the cost of versatility. In particular, database systems only operate efficiently over loaded data, i.e., data converted from its original raw format into the system's internal data format. At the same time, data volume continues to increase exponentially and data varies increasingly, with an escalating number of new formats. The consequence is a growing impedance mismatch between the original structures holding the data in the raw files and the structures used by query engines for efficient processing. In an ideal scenario, the query engine would seamlessly adapt itself to the data and ensure efficient query processing regardless of the input data formats, optimizing itself to each instance of a file and of a query by leveraging information available at query time. Today's systems, however, force data to adapt to the query engine during data loading. This paper proposes adapting the query engine to the formats of raw data. It presents RAW, a prototype query engine which enables querying heterogeneous data sources transparently. RAW employs Just-In-Time access paths, which efficiently couple heterogeneous raw files to the query engine and reduce the overheads of traditional general-purpose scan operators. There are, however, inherent overheads with accessing raw data directly that cannot be eliminated, such as converting the raw values. Therefore, RAW also uses column shreds, ensuring that we pay these costs only for the subsets of raw data strictly needed by a query. We use RAW in a real-world scenario and achieve a two-order of magnitude speedup against the existing hand-written solution.
Manos Karpathiotakis, Miguel Branco, Ioannis Alagiannis, Anastasia Ailamaki
Proc. VLDB Endow.1
2014 Wildfire monitoring using satellite images, ontologies and linked geospatial data
Kostis Kyzirakos, Manos Karpathiotakis, George Garbis, Charalampos Nikolaou, Konstantina Bereta, Ioannis Papoutsis, Themos Herekakis, Dimitrios Michail 0001, Manolis Koubarakis, Charalambos Kontoes
J. Web Semant.2
2013 The Spatiotemporal RDF Store Strabon
Kostis Kyzirakos, Manos Karpathiotakis, Konstantina Bereta, George Garbis, Charalampos Nikolaou, Panayiotis Smeros, Stella Giannakopoulou, Kallirroi Dogani, Manolis Koubarakis
SSTD2
2012 Strabon: A Semantic Geospatial DBMS
Kostis Kyzirakos, Manos Karpathiotakis, Manolis Koubarakis
ISWC (1)2
2012 TELEIOS: A Database-Powered Virtual Earth Observatory
abstract
TELEIOS is a recent European project that addresses the need for scalable access to petabytes of Earth Observation data and the discovery and exploitation of knowledge that is hidden in them. TELEIOS builds on scientific database technologies (array databases, SciQL, data vaults) and Semantic Web technologies (stRDF and stSPARQL) implemented on top of a state of the art column store database system (MonetDB). We demonstrate a first prototype of the TELEIOS Virtual Earth Observatory (VEO) architecture, using a forest fire monitoring application as example.
Manolis Koubarakis, Kostis Kyzirakos, Manos Karpathiotakis, Charalampos Nikolaou, Stavros Vassos, George Garbis, Michael Sioutis, Konstantina Bereta, Dimitrios Michail 0001, Charalambos Kontoes, Ioannis Papoutsis, Themos Herekakis, Stefan Manegold, Martin L. Kersten, Milena Ivanova, Holger Pirk, Ying Zhang 0027, Mihai Datcu, Gottfried Schwarz, Corneliu Octavian Dumitru, Daniela Espinoza-Molina, Katrin Molch, Ugo Di Giammatteo, Manuela Sagona, Sergio Perelli, Thorsten Reitz, Eva Klien, Robert Gregor
Proc. VLDB Endow.3
2011 A Semantically Enabled Service Architecture for Mashups over Streaming and Stored Data
Alasdair J. G. Gray, Raúl García-Castro, Kostis Kyzirakos, Manos Karpathiotakis, Jean-Paul Calbimonte, Kevin R. Page, Jason Sadler, Alex Frazer, Ixent Galpin, Alvaro A. A. Fernandes, Norman W. Paton, Óscar Corcho, Manolis Koubarakis, David De Roure, Kirk Martinez, Asunción Gómez-Pérez
ESWC (2)4