EDBT 2026 Demo / reviewers in the wild / expert
Nikolaos Papailiou
dblp:70/11205
· DBLP profile ↗
10ranked-venue papers
3as first author
0since 2021 · last 2019
0000-0001-6030-2762ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 9 · 3 first-authorArtificial intelligence and machine learning · 4 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
4 papers |
Graph data management · 61% Query processing and optimization · 19% Data mining · 13% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Cloud and datacenter computing · 87% GPUs and heterogeneous computing · 13% | |
| Theoretical computer science
1 paper |
Graph algorithms and graph theory · 100% |
Topics — the 14 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Graph data management › RDF data management
RDF query processing |
0.4 | 2 | 2015 | Graph-Aware, Workload-Adaptive SPARQL Query Caching · SIGMOD Conference 2015 H2RDF+: an efficient data management system for big RDF graphs · SIGMOD Conference 2014 |
Graph data management › graph data model
uncertain graph |
0.4 | 1 | 2019 | Uncertain Graph Sparsification (Extended Abstract) · ICDE 2019 |
Graph algorithms and graph theory
graph sparsification |
0.4 | 1 | 2019 | Uncertain Graph Sparsification (Extended Abstract) · ICDE 2019 |
Data mining › structured data mining
graph mining |
0.3 | 1 | 2018 | Uncertain Graph Sparsification · IEEE Trans. Knowl. Data Eng. 2018 |
Graph data management › graph transformation
graph sparsification |
0.3 | 1 | 2018 | Uncertain Graph Sparsification · IEEE Trans. Knowl. Data Eng. 2018 |
Graph data management › graph processing
uncertain graph processing |
0.3 | 1 | 2018 | Uncertain Graph Sparsification · IEEE Trans. Knowl. Data Eng. 2018 |
Query processing and optimization
query result caching |
0.2 | 1 | 2015 | Graph-Aware, Workload-Adaptive SPARQL Query Caching · SIGMOD Conference 2015 |
Cloud and datacenter computing
big data analytics |
0.2 | 1 | 2015 | IReS: Intelligent, Multi-Engine Resource Scheduler for Big Data Analytics Workflows · SIGMOD Conference 2015 |
Cloud and datacenter computing
cluster resource management and scheduling |
0.2 | 1 | 2015 | IReS: Intelligent, Multi-Engine Resource Scheduler for Big Data Analytics Workflows · SIGMOD Conference 2015 |
Query processing and optimization › join processing
distributed join |
0.2 | 1 | 2014 | H2RDF+: an efficient data management system for big RDF graphs · SIGMOD Conference 2014 |
Distributed and cloud data management
distributed query processing |
0.2 | 1 | 2014 | H2RDF+: an efficient data management system for big RDF graphs · SIGMOD Conference 2014 |
Graph data management › RDF data management
RDF triple store |
0.1 | 1 | 2015 | Graph-Aware, Workload-Adaptive SPARQL Query Caching · SIGMOD Conference 2015 |
GPUs and heterogeneous computing › heterogeneous architecture
heterogeneous processors |
0.1 | 1 | 2015 | IReS: Intelligent, Multi-Engine Resource Scheduler for Big Data Analytics Workflows · SIGMOD Conference 2015 |
Query processing and optimization › query optimization
cost-based optimization |
0.1 | 1 | 2014 | H2RDF+: an efficient data management system for big RDF graphs · SIGMOD Conference 2014 |
Methods — techniques the papers use, named apart from their topics
edge probability redistribution · 0.3dynamic programming planner · 0.2cost modeling · 0.2canonical labelling · 0.2sort-merge join · 0.2merge join · 0.2greedy planner · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | Uncertain Graph Sparsification (Extended Abstract)abstractUncertain graphs are prevalent in several applications including communications systems, biological databases and social networks. The ever increasing size of the underlying data renders both graph storage and query processing extremely expensive. Sparsification has often been used to reduce the size of deterministic graphs by maintaining only the important edges. However, adaptation of deterministic sparsification methods fails in the uncertain setting. To overcome this problem, we introduce the first sparsification techniques aimed explicitly at uncertain graphs. The proposed methods reduce the number of edges and redistribute their probabilities in order to decrease the graph size, while preserving its underlying structure. The resulting graph can be used to efficiently and accurately approximate any query and mining tasks on the original graph, including clustering coefficient, page rank, reliability and shortest path distance. Panos Parchas, Nikolaos Papailiou, Dimitris Papadias, Francesco Bonchi |
ICDE | 2 |
| 2018 | Uncertain Graph SparsificationabstractUncertain graphs are prevalent in several applications including communications systems, biological databases, and social networks. The ever increasing size of the underlying data renders both graph storage and query processing extremely expensive. Sparsification has often been used to reduce the size of deterministic graphs by maintaining only the important edges. However, adaptation of deterministic sparsification methods fails in the uncertain setting. To overcome this problem, we introduce the first sparsification techniques aimed explicitly at uncertain graphs. The proposed methods reduce the number of edges and redistribute their probabilities in order to decrease the graph size, while preserving its underlying structure. The resulting graph can be used to efficiently and accurately approximate any query and mining tasks on the original graph. An extensive experimental evaluation with real and synthetic datasets illustrates the effectiveness of our techniques on several common graph tasks, including clustering coefficient, page rank, reliability, and shortest path distance. Panos Parchas, Nikolaos Papailiou, Dimitris Papadias, Francesco Bonchi |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2016 | Mix 'n' match multi-engine analyticsabstractCurrent platforms fail to efficiently cope with the data and task heterogeneity of modern analytics workflows due to their adhesion to a single data and/or compute model. As a remedy, we present IReS, the Intelligent Resource Scheduler for complex analytics workflows executed over multi-engine environments. IReS is able to optimize a workflow with respect to a user-defined policy relying on cost and performance models of the required tasks over the available platforms. This optimization consists in allocating distinct workflow parts to the most advantageous execution and/or storage engine among the available ones and deciding on the exact amount of resources provisioned. Our current prototype supports 5 compute and 3 data engines, yet new ones can effortlessly be added to IReS by virtue of its engine-agnostic mechanisms. Our extensive experimental evaluation confirms that IReS speeds up diverse and realistic workflows by up to 30% compared to their optimal single-engine plan by automatically scattering parts of them to different execution engines and datastores. Its optimizer incurs only marginal overhead to the workflow execution performance, managing to discover the optimal execution plan within a few seconds, even for large-scale workflow instances. Katerina Doka, Nikolaos Papailiou, Victor Giannakouris, Dimitrios Tsoumakos, Nectarios Koziris |
IEEE BigData | 2 |
| 2016 | MuSQLE: Distributed SQL query execution over multiple engine environmentsabstractMulti-engine analytics has been gaining an increasing amount of attention from both the academic and the industrial community as it can successfully cope with the heterogeneity and complexity that the plethora of frameworks, technologies and requirements have brought forth. It is now common for a data analyst to combine data that resides on multiple and totally independent engines and perform complex analytics queries. Multi-engine solutions based on SQL can facilitate such efforts, as SQL is a popular standard that the majority of data-scientists understands. Existing solutions propose a middleware that centrally optimizes query execution for multiple engines. Yet, this approach requires manual integration of every primitive engine operator along with its cost model, rendering the process of adding new operators or engines highly inextensible. To address this issue we present MuSQLE, a system for SQL-based analytics over multi-engine environments. MuSQLE can efficiently utilize external SQL engines allowing for both intra and inter engine optimizations. Our framework adopts a novel API-based strategy. Instead of manual integration, MuSQLE specifies a generic API, used for the cost estimation and query execution, that needs to be implemented for each SQL engine endpoint. Our engine API is integrated with a state-of-the-art query optimizer, adding support for location-based, multi-engine query optimization and letting individual runtimes perform sub-query physical optimization. The derived multi-engine plans are executed using the Spark distributed execution framework. Our detailed experimental evaluation, integrating PostgreSQL, MemSQL and SparkSQL under MuSQLE, demonstrates its ability to accurately decide on the most suitable execution engine. MuSQLE can provide speedups of up to 1 order of magnitude for TPCH queries, leveraging different engines for the execution of individual query parts. Victor Giannakouris, Nikolaos Papailiou, Dimitrios Tsoumakos, Nectarios Koziris |
IEEE BigData | 2 |
| 2015 | PANIC: Modeling Application Performance over Virtualized ResourcesabstractIn this work we address the problem of predicting the performance of a complex application deployed over virtualized resources. Cloud computing has enabled numerous companies to develop and deploy their applications over cloud infrastructures for a wealth of reasons including (but not limited to) decrease costs, avoid administrative effort, rapidly allocate new resources, etc. Virtualization however, adds an extra layer in the software stack, hardening the prediction of the relation between the resources and the application performance, which is a key factor for every industry. To address this challenge we propose PANIC, a system which obtains knowledge for the application by actually deploying it over a cloud infrastructure and then, approximating the performance of the application for the all possible deployment configurations. The user of PANIC defines a set of resources along with their respective ranges and then the system samples the deployment space formed by all the combinations of the resources, deploys the application in some representative points and utilizes a wealth of approximation techniques to predict the behavior of the application in the remainder space. The experimental evaluation has indicated that a small portion of the possible deployment configurations is enough to create profiles with high accuracy for three real world applications. Ioannis Giannakopoulos, Dimitrios Tsoumakos, Nikolaos Papailiou, Nectarios Koziris |
IC2E | 3 |
| 2015 | IReS: Intelligent, Multi-Engine Resource Scheduler for Big Data Analytics WorkflowsabstractBig data analytics tools are steadily gaining ground at becoming indispensable to businesses worldwide. The complexity of the tasks they execute is ever increasing due to the surge in data and task heterogeneity. Current analytics platforms, while successful in harnessing multiple aspects of this ``data deluge", bind their efficacy to a single data and compute model and often depend on proprietary systems. However, no single execution engine is suitable for all types of computation and no single data store is suitable for all types of data. To this end, we demonstrate IReS, the Intelligent Resource Scheduler for complex analytics workflows executed over multi-engine environments. Our system models the cost and performance of the required tasks over the available platforms. IReS is then able to match distinct workflow parts to the execution and/or storage engine among the available ones in order to optimize with respect to a user-defined policy. During the demo, the attendees will be able to execute workflows that match real use cases and parametrize the input datasets and optimization policy. The underlying platform supports multiple compute and data engines, allowing the user to choose any subset of them. Through the inspection of the produced plan, its execution and the collection and presentation of numerous cost and performance metrics, the audience will experience first-hand how IReS takes advantage of heterogeneous runtimes and data stores and effectively models operator cost and performance for actual and diverse workflows. Katerina Doka, Nikolaos Papailiou, Dimitrios Tsoumakos, Christos Mantas, Nectarios Koziris |
SIGMOD Conference | 2 |
| 2015 | Graph-Aware, Workload-Adaptive SPARQL Query CachingabstractThe pace at which data is described, queried and exchanged using the RDF specification has been ever increasing with the proliferation of Semantic Web. Minimizing SPARQL query response times has been an open issue for the plethora of RDF stores, yet SPARQL result caching techniques have not been extensively utilized. In this work we present a novel system that addresses graph-based, workload-adaptive indexing of large RDF graphs by caching SPARQL query results. At the heart of the system lies a SPARQL query canonical labelling algorithm that is used to uniquely index and reference SPARQL query graphs as well as their isomorphic forms. We integrate our canonical labelling algorithm with a dynamic programming planner in order to generate the optimal join execution plan, examining the utilization of both primitive triple indexes and cached query results. By monitoring cache requests, our system is able to identify and cache SPARQL queries that, even if not explicitly issued, greatly reduce the average response time of a workload. The proposed cache is modular in design, allowing integration with different RDF stores. Incorporating it to an open-source, distributed RDF engine that handles large scale RDF datasets, we prove that workload-adaptive caching can reduce average response times by up to two orders of magnitude and offer interactive response times for complex workloads and huge RDF datasets. Nikolaos Papailiou, Dimitrios Tsoumakos, Panagiotis Karras, Nectarios Koziris |
SIGMOD Conference | 1 |
| 2014 | CELAR: Automated application elasticity platformabstractOne of the main promises of the cloud computing paradigm is the ability to scale resources on-demand. This feature characterizes the cloud era, where the overhead of early expenditure for infrastructure is eliminated. Innovative services are thus able to enter the market quicker and adopt faster to new challenges and user demand. One of the main aspects of this on-demand nature is the concept of elasticity, i.e., the ability of autonomously provision and de-provision resources by reacting to changes in the incoming load. An elastic service is able to operate with an optimal cost by expanding and contracting its used resources at runtime and according to demand. This does not only minimizes running cost, but also avoids disruptive outages due to spikes in service usage. While the various layers comprising a cloud service can be scaled, this does not happen in a unified manner. The vision of CELAR is to provide a fully integrated software stack that manages resource allocation for cloud applications in an autonomous, efficient and generic manner. In order to achieve that, CELAR incorporates novel methodologies for describing cloud applications, monitoring the use of various resources, evaluating cost, taking informed decisions and interacting with the underlying cloud infrastructure. Our goal is two-fold. On the one hand is developing the methodologies for achieving multi-grained, automatic elasticity control on both application and infrastructure level. On the other hand is developing the open-source tools that implement those methods in an integrated manner. Hereby we present an overview of the CELAR platform, explaining its architectural components and some basic workflows that show how they interact in order to achieve the core functionalities. Ioannis Giannakopoulos, Nikolaos Papailiou, Christos Mantas, Ioannis Konstantinou, Dimitrios Tsoumakos, Nectarios Koziris |
IEEE BigData | 2 |
| 2014 | H2RDF+: an efficient data management system for big RDF graphsabstractThe proliferation of data in RDF format has resulted in the emergence of a plethora of specialized management systems. While the ability to adapt to the complexity of a SPARQL query -- given their inherent diversity -- is crucial, current approaches do not scale well when faced with substantially complex, non-selective joins, resulting in exponential growth of execution times. In this demonstration we present H2 RDF+, an RDF store that efficiently performs distributed Merge and Sort-Merge joins using a multiple-index scheme over HBase indexes. Through a greedy planner that incorporates our cost-model, it adaptively commands for either single or multi-machine query execution based on join complexity. In this paper, we present its key scientific contributions and allow participants to interact with an H2RDF+ deployment over a Cloud infrastructure. Using a web-based GUI we allow users to load different datasets (both real and synthetic), apply any query (custom or predefined) and monitor its execution. By allowing real-time inspection of cluster status, response times and committed resources the audience will evaluate the validity of H2RDF+'s claims and perform direct comparisons to two other state-of-the-art RDF stores. Nikolaos Papailiou, Dimitrios Tsoumakos, Ioannis Konstantinou, Panagiotis Karras, Nectarios Koziris |
SIGMOD Conference | 1 |
| 2013 | H2RDF+: High-performance distributed joins over large-scale RDF graphsabstractThe proliferation of data in RDF format calls for efficient and scalable solutions for their management. While scalability in the era of big data is a hard requirement, modern systems fail to adapt based on the complexity of the query. Current approaches do not scale well when faced with substantially complex, non-selective joins, resulting in exponential growth of execution times. In this work we present H2RDF+, an RDF store that efficiently performs distributed Merge and Sort-Merge joins over a multiple index scheme. H2RDF+ is highly scalable, utilizing distributed MapReduce processing and HBase indexes. Utilizing aggressive byte-level compression and result grouping over fast scans, it can process both complex and selective join queries in a highly efficient manner. Furthermore, it adaptively chooses for either single- or multi-machine execution based on join complexity estimated through index statistics. Our extensive evaluation demonstrates that H2RDF+ efficiently answers non-selective joins an order of magnitude faster than both current state-of-the-art distributed and centralized stores, while being only tenths of a second slower in simple queries, scaling linearly to the amount of available resources. Nikolaos Papailiou, Ioannis Konstantinou, Dimitrios Tsoumakos, Panagiotis Karras, Nectarios Koziris |
IEEE BigData | 1 |