VLDB 2026 Research / reviewers in the wild / expert
Stathis Maroulis
dblp:194/7743
· DBLP profile ↗
8ranked-venue papers
4as first author
0since 2021 · last 2019
0000-0002-2872-7821ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 3 · 1 first-authorArtificial intelligence and machine learning · 2 · 1 first-authorSystems, architecture and hardware · 2 · 2 first-authorSecurity and privacy · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Cloud and datacenter computing · 48% Energy-efficient computing · 48% Performance modeling and evaluation · 4% |
Topics — the 7 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Cloud and datacenter computing
cluster resource management and scheduling |
0.7 | 2 | 2019 | A Holistic Energy-Efficient Real-Time Scheduler for Mixed Stream and Batch Processing Workloads · IEEE Trans. Parallel Distributed Syst. 2019 Dione: A Framework for Automatic Profiling and Tuning Big Data Applications · ICDE 2018 |
Energy-efficient computing › power management
dynamic voltage and frequency scaling |
0.4 | 1 | 2019 | A Holistic Energy-Efficient Real-Time Scheduler for Mixed Stream and Batch Processing Workloads · IEEE Trans. Parallel Distributed Syst. 2019 |
Energy-efficient computing
energy-aware scheduling |
0.4 | 1 | 2019 | A Holistic Energy-Efficient Real-Time Scheduler for Mixed Stream and Batch Processing Workloads · IEEE Trans. Parallel Distributed Syst. 2019 |
Energy-efficient computing
power management |
0.4 | 1 | 2019 | A Holistic Energy-Efficient Real-Time Scheduler for Mixed Stream and Batch Processing Workloads · IEEE Trans. Parallel Distributed Syst. 2019 |
Cloud and datacenter computing
configuration tuning |
0.3 | 1 | 2018 | Dione: A Framework for Automatic Profiling and Tuning Big Data Applications · ICDE 2018 |
Cloud and datacenter computing › big data analytics
batch and stream processing |
0.1 | 1 | 2019 | A Holistic Energy-Efficient Real-Time Scheduler for Mixed Stream and Batch Processing Workloads · IEEE Trans. Parallel Distributed Syst. 2019 |
Performance modeling and evaluation
workload characterization |
0.1 | 1 | 2018 | Dione: A Framework for Automatic Profiling and Tuning Big Data Applications · ICDE 2018 |
Methods — techniques the papers use, named apart from their topics
time-series segmentation · 0.4regression · 0.4DVFS · 0.4prediction model · 0.3execution plan similarity · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | Fast. Efficient Performance Predictions for Big Data ApplicationsabstractIn recent years we observe a rapid growth in the deployment of machine learning workloads on big data analytics frameworks like Apache Spark and Apache Flink. These workloads are typically represented as graphs, run on shared infrastructures and often have much more demanding resource requirements than those traditionally found in typical enterprise settings. However, predicting the execution times of the workloads is important as they often run on shared public or private infrastructures and, thus, their execution is greatly affected by the resource sharing, the hardware infrastructure utilized as well as the choice of the configuration parameters provided by the frameworks. In this work, we propose a fast and efficient performance prediction system to address the challenge of predicting the execution times of big data workloads, exploiting the fact that workloads are represented as processing graphs and often share similar structures and parameters. Thus, we can use the performance models we have built for already deployed workloads, to estimate the end-to-end execution time for a new workload. Previous works assume that a large number of profiling runs can be utilized for building the prediction models. However, this assumption is not always valid and more elaborate mechanisms need to be applied. Our detailed experimental evaluation on our local Spark cluster illustrates that our approach can predict accurately the execution time of a wide range of Spark workloads. Stathis Maroulis, Nikos Zacheilas, Thanasis Theocharis, Vana Kalogeraki |
ISORC | 1 |
| 2019 | ESCAPe: Elastic Caching For Big Data SystemsabstractIn recent years, in-memory cache systems have been commonly utilized to help maintain low application response times, compared to traditional relational databases, where the data items are stored in disk drives. Although cache memory systems offer improved performance, running everything in memory might not be cost-effective. In this paper, we present ESCAPe, an elastic high throughput and low latency key-value in-memory cache system. Unlike existing schemes, ESCAPe offers an elastic mechanism that proactively adds or removes nodes to scale-down or scale-up to meet fluctuating application demands, and incorporates a dynamic redistribution scheme that prioritizes the distribution of the keys at the nodes, while keeping the overhead cost as low as possible. We have evaluated our approach in a real cluster, using ESCAPe as the memcache system for Web Applications using different workload traces and comparing our approach with state of the art schemes. Our results illustrate that ESCAPe is able to select the most useful items to keep in memory, significantly reducing the end-to-end latency experienced by the applications. Thanasis Priovolos, Stathis Maroulis, Vana Kalogeraki |
SRDS | 2 |
| 2019 | A Framework for Managing an Elastic Redis CacheabstractIn this demonstration we present ESCAPe, a framework for automatically monitoring and managing an elastic Redis cache. System administrators can configure ESCAPe in concert with their Redis cluster, to meet application real-time objectives while minimizing the cost associated with scaling the applications. ESCAPe comprises Agents that are deployed at each node in the Redis cluster in order to implement the scaling decisions and apply the proposed eviction policy, and a ESCAPe Manager that communicates with all other nodes and makes the scaling decisions in response to the monitored response times and costs. Finally ESCAPe's Web User Interface visualizes realtime statistics related to response time and the distribution of object types in the cluster nodes as well as historical statistics. The system administrators can further fine-tune the cluster parameters and monitor the impact of their decisions on the application performance and cost. Thanasis Priovolos, Stathis Maroulis, Vana Kalogeraki |
SRDS | 2 |
| 2019 | A Holistic Energy-Efficient Real-Time Scheduler for Mixed Stream and Batch Processing WorkloadsabstractIn recent years we have experienced a wide adoption of novel distributed processing frameworks such as Apache Spark for handling batch and stream processing big data applications. An important aspect that has not been examined in these systems yet, is the energy consumption during the applications' execution. Reducing the energy consumption of modern datacenters is a necessity, as datacenters contribute over 2 percent of the total US electric usage. However, efficiently scheduling applications in distributed processing systems can be challenging as there is a trade-off between minimizing the datacenter's energy usage and satisfying the application performance requirements. In this work we propose, ExpREsS, a scheduler for orchestrating the execution of Spark applications in a way that enables us to minimize the energy consumption while ensuring that the applications' performance requirements are met. Our approach exploits time-series segmentation for capturing the applications' energy usage and execution times, and then applies a novel DVFS technique to minimize the energy consumption. In order to tackle the limited number of application's profiling runs, we exploit regression techniques to predict the applications' execution times and power consumption. Our detailed experimental evaluation using realistic workloads on our local cluster illustrates the working and benefits of our approach. Stathis Maroulis, Nikos Zacheilas, Vana Kalogeraki |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2018 | Dione: A Framework for Automatic Profiling and Tuning Big Data ApplicationsabstractIn this demonstration we presentDionea novel framework for automatic profiling and tuning big data applications. Our system allows a non-expert user to submit Spark or Flink applications to his/her cluster and Dione automatically determines the impact of different configuration parameters on the application's execution time and monetary cost. Dione is the first framework that exploits similarities in the execution plans of different applications to narrow down the amount of profiling runs that are required for building prediction models that capture the impact of the configuration parameters on the metrics of interest. Dione exploits these prediction models to tune the configuration parameters in a way that minimizes the application's execution time or the user's budget. Finally, Dione's Web-UI visualizes the impact of the configuration parameters on the execution time and the monetary cost, and enables the user to submit the application with the recommended parameters' values. Nikos Zacheilas, Stathis Maroulis, Thanasis Priovolos, Vana Kalogeraki, Dimitrios Gunopulos |
ICDE | 2 |
| 2017 | Dione: Profiling spark applications exploiting graph similarityabstractIn recent years distributed processing frameworks such as Apache Spark have been utilized for running big data applications. Predicting the application's execution time has been an important goal since it can help the end user to determine the necessary processing resources to be reserved. While there have been some previous works that examine the problem of profiling Spark applications, they mainly focus on specific application types (e.g., Machine learning applications) and rely on the existence of a large number of previous execution runs. In this work we aim at overcoming these limitations by minimizing the number of past execution runs needed for the profiling phase. Furthermore, we identify patterns of continuous identical dataset transformations between different applications to cope with the limited historical data availability. We propose an on-line profiling framework, called Dione, that estimates the running times of new applications, even if no historical data is available. Finally, in our detailed experimental evaluation, using practical workloads on our local cluster, we illustrate that our approach accurately predicts the execution times of Spark applications and requires 30% less training time and monetary cost compared to the current state-of-the-art techniques. Nikos Zacheilas, Stathis Maroulis, Vana Kalogeraki |
IEEE BigData | 2 |
| 2017 | A Framework for Efficient Energy Scheduling of Spark WorkloadsabstractNowadays distributed processing frameworks like Apache Spark have been successfully used for the execution of big data applications. Despite their wide adoption little work has been done in terms of controlling the applications' energy consumption. Datacenters contribute over 2 % of the total US electric usage therefore minimizing the energy utilization of Spark application can be extremely helpful. Solving this energy consumption problem requires the scheduling of Spark applications in an energy-efficient way. However, the problem is challenging as we also have to consider application performance requirements. In this work, we provide the overview of a novel framework that orchestrates the execution order of Spark applications, exploiting DVFS to tune the computing nodes CPU frequencies in order to minimize the energy consumption and satisfy application's performance requirements. Our early experimental results illustrate the working and benefits of our framework. Stathis Maroulis, Nikos Zacheilas, Vana Kalogeraki |
ICDCS | 1 |
| 2016 | Context-aware point of interest recommendation using tensor factorizationabstractThe wide adoption of Location Based Social Networks along with advances in mobile technology, has brought forth as a core service the analysis of large volumes of location-based data for personalized Point of Interest (POIs) recommendations. The majority of the existing recommendation systems take advantage of Collaborative Filtering, but they fail to exploit the contextual information involved with POI checkins (i.e., POI category, location, or the checkin timestamp). In this paper we propose CoTF, a Context-Aware Point of Interest Recommendation system using Tensor Factorization, that aims at enhancing the user experience by providing personalized context aware POI recommendations. Our approach exploits Category-based context related to checkins without the need of any pre-or post-filtering techniques. Our detailed experimental evaluation using real data from the Foursquare location-based social network illustrates that our approach can efficiently produce personalized recommendations to users, while significantly reducing the training time compared to current state-of-the-art methods. Stathis Maroulis, Ioannis Boutsis, Vana Kalogeraki |
IEEE BigData | 1 |