EDBT 2026 Demo / reviewers in the wild / expert
Herodotos Herodotou
dblp:26/5170
· DBLP profile ↗
27ranked-venue papers in the field
13as first author
12since 2021 · last 2024
0000-0002-8717-1691ORCID · corroborated
Domains — venue-derived; a paper can count in several
Database Systems & Data Management · 25 (11 first)Data Mining & Knowledge Discovery · 1 (1 first)Big Data, Cloud & Distributed Data Systems · 1 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Optimizing Distributed Tiered Data Storage Systems with DITISabstractModern data storage systems are characterized by a distributed architecture as well as the presence of multiple storage tiers and caches. Both system developers and operators are challenged with the complexity of such systems as it is hard to evaluate how a configuration change will impact the workload or system performance and identify the best configuration to satisfy some performance objective. DITIS is a new simulator that models the end-to-end execution of file requests on distributed tiered storage systems that addresses the aforementioned challenges efficiently without any costly system redeployments. The demonstration will showcase the key functionalities and benefits offered by DITIS, including (i) analyzing workload traces to understand their characteristics and the behavior of the underlying storage system; (ii) running simulations with different configurations to evaluate their impact on performance; and (iii) running optimizations over custom search spaces to find the best configuration that satisfies a given objective. Sotiris Vasileiadis, Matthew Paraskeva, George Savva, Andreas Efstathiou, Edson Ramiro Lucas Filho, Jianqiang Shen, Lun Yang, Ke-Bo Fu, Herodotos Herodotou |
Proc. VLDB Endow. | 9 |
| 2023 | On combining system and machine learning performance tuning for distributed data stream applications
Lambros Odysseos, Herodotos Herodotou |
Distributed Parallel Databases | 2 |
| 2023 | Cost-based Data Prefetching and Scheduling in Big Data Platforms over Tiered Storage SystemsabstractThe use of storage tiering is becoming popular in data-intensive compute clusters due to the recent advancements in storage technologies. The Hadoop Distributed File System, for example, now supports storing data in memory, SSDs, and HDDs, while OctopusFS and hatS offer fine-grained storage tiering solutions. However, current big data platforms (such as Hadoop and Spark) are not exploiting the presence of storage tiers and the opportunities they present for performance optimizations. Specifically, schedulers and prefetchers will make decisions only based on data locality information and completely ignore the fact that local data are now stored on a variety of storage media with different performance characteristics. This article presents Trident, a scheduling and prefetching framework that is designed to make task assignment, resource scheduling, and prefetching decisions based on both locality and storage tier information. Trident formulates task scheduling as a minimum cost maximum matching problem in a bipartite graph and utilizes two novel pruning algorithms for bounding the size of the graph, while still guaranteeing optimality. In addition, Trident extends YARN’s resource request model and proposes a new storage-tier-aware resource scheduling algorithm. Finally, Trident includes a cost-based data prefetching approach that coordinates with the schedulers for optimizing prefetching operations. Trident is implemented in both Spark and Hadoop and evaluated extensively using a realistic workload derived from Facebook traces as well as an industry-validated benchmark, demonstrating significant benefits in terms of application performance and cluster efficiency. Herodotos Herodotou, Elena Kakoulli |
ACM Trans. Database Syst. | 1 |
| 2022 | Automatic Performance Tuning for Distributed Data Stream Processing SystemsabstractDistributed data stream processing systems (DSPSs) such as Storm, Flink, and Spark Streaming are now routinely used to process continuous data streams in (near) real-time. However, achieving the low latency and high throughput demanded by today's streaming applications can be a daunting task, especially since the performance of DSPSs highly depends on a large number of system parameters that control load balancing, degree of parallelism, buffer sizes, and various other aspects of system execution. This tutorial offers a comprehensive review of the state-of-the-art automatic performance tuning approaches that have been proposed in recent years. The approaches are organized into five main categories based on their methodologies and features: cost modeling, simulation-based, experiment-driven, machine learning, and adaptive tuning. The categories of approaches will be analyzed in depth and compared to each other, exposing their various strengths and weaknesses. Finally, we will identify several open research problems and challenges related to automatic performance tuning for DSPSs. Herodotos Herodotou, Lambros Odysseos, Yuxing Chen 0003, Jiaheng Lu |
ICDE | 1 |
| 2022 | Estimation of Sea Surface Current Velocities using AIS DataabstractThe Automatic Identification System (AIS) provides information for tracking and monitoring vessel activity in real time. The vessel traffic data from AIS includes position coordinates in latitude and longitude, speed and course over ground, the vessel's unique identification number, and many more. In this work, we investigate the use of AIS data for estimating sea surface current velocities in the Eastern Mediterranean sea. Specifically, we apply the dead reckoning technique to compute the difference between the projected position and the true position of a vessel over time, which is mainly attributed on the force of sea surface currents. The estimated sea surface current velocities and directions are compared with the ones provided by the Copernicus Marine ocean product system. The analysis reveals that the dead reckoning technique can be used reliably for estimating sea currents at a very fine granularity, especially in high-traffic and coastal areas, where there is an increased complexity of obtaining accurate results from other sources. Konstantinos Christodoulou, Herodotos Herodotou, Michalis P. Michaelides |
MDM | 2 |
| 2022 | An Intelligent Framework for Vessel Traffic Monitoring Using AIS DataabstractAutomatic identification system (AIS) data provides a wealth of information regarding vessel traffic and is used for a variety of applications such as collision detection and avoidance, route prediction and optimization, search and rescue operations, etc. However, several challenges exist when working with AIS data including huge volume and velocity (as AIS signals are sent by vessels every few seconds), message duplication, various types of data irregularities, as well as the need for real-time processing and analysis. This paper presents a new framework for collecting, processing, storing, and analyzing AIS data in real time plus a set of algorithms for doing so in an efficient and scalable way. At the same time, a set of intelligent services are provided as building blocks for improving and creating new AIS data driven applications. This framework has been operational for the past few years in Cyprus, and has collected and processed around one billion AIS messages from the Eastern Mediterranean Sea. Nicos Evmides, Lambros Odysseos, Michalis P. Michaelides, Herodotos Herodotou |
MDM | 4 |
| 2022 | Online Analytical Processing of Port Calls for Decision SupportabstractThe port call process encapsulates a visitation cycle of a ship to a port and can generate a wealth of data. The real time analysis of port call data can be used to find bottlenecks in the port call process, establish targets based on key performance indicators (KPIs), and to understand how shipping traffic impacts a port's efficiency. This demonstration will showcase a new Power BI interactive report powered by a multidimensional OLAP cube for very fast performance, which is built on top of a data warehouse collecting information from various sources in real time. The report currently visualizes several KPIs and other types of information that can be filtered per port, time-period, vessel type, origin or destination ports, and various other categories to help manage arrivals, departures, and port operations. Aidan Worth, Aris Televantos, Nicos Evmides, Michalis P. Michaelides, Herodotos Herodotou |
MDM | 5 |
| 2022 | Introduction to the special issue on self‑managing and hardware‑optimized database systems 2020
Herodotos Herodotou, Panos K. Chrysanthis, Shimin Chen, Meichun Hsu, Khuzaima Daudjee, Yingjun Wu, Constantinos Costa |
Distributed Parallel Databases | 1 |
| 2022 | Hihooi: A Database Replication Middleware for Scaling Transactional Databases ConsistentlyabstractWith the advent of the Internet and Internet-connected devices, modern business applications can experience rapid increases as well as variability in transactional workloads. Database replication has been employed to scale performance and improve availability of relational databases but past approaches have suffered from various issues including limited scalability, performance versus consistency tradeoffs, and requirements for database or application modifications. This paper presents Hihooi, a replication-based middleware system that is able to achieve workload scalability, strong consistency guarantees, and elasticity for existing transactional databases at a low cost. A novel replication algorithm enables Hihooi to propagate database modifications asynchronously to all replicas at high speeds, while ensuring that all replicas are consistent. At the same time, a fine-grained routing algorithm is used to load balance incoming transactions to available replicas in a consistent way. Our thorough experimental evaluation with several well-established benchmarks shows how Hihooi is able to achieve almost linear workload scalability for transactional databases. Michael A. Georgiou, Aristodemos Paphitis, Michael Sirivianos, Herodotos Herodotou |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2021 | Catching them red-handed: Real-time Aggression Detection on Social MediaabstractAggression on social media has evolved into a major point of concern. However, recently proposed machine learning (ML) approaches to detect various types of aggressive behavior fall short, due to the fast and increasing pace of content generation as well as evolution of such behavior over time. This work introduces the first, practical, real-time framework for detecting aggression on Twitter via embracing the streaming ML paradigm. This method adapts its ML binary classifiers in an incremental fashion, while receiving new annotated examples, and achieves similar performance as batch-based ML models, with 82-93% accuracy, precision, and recall. Experimental analysis on real Twitter data reveals how this framework, implemented in Spark Streaming, easily scales to process millions of tweets in minutes. Herodotos Herodotou, Despoina Chatzakou, Nicolas Kourtellis |
ICDE | 1 |
| 2021 | Attaining Workload Scalability and Strong Consistency for Replicated Databases with HihooiabstractDatabase replication can be employed for scaling transactional workloads while maintaining strong consistency semantics. However, past approaches suffer from various issues such as limited scalability, performance versus consistency tradeoffs, and requirements for database or application modifications. Hihooi is a new replication-based master-slave middleware system that is able to overcome the aforementioned limitations. The novelty of Hihooi lies in its modern architecture as well as its replication and transaction routing algorithms. In particular, Hihooi replicates all write statements asynchronously and applies them in parallel at the replica nodes, while ensuring replica consistency. At the same time, a fine-grained transaction routing algorithm ensures that all read transactions are load balanced to the replicas consistently. This demonstration will showcase the key functionalities of Hihooi, including (i) practical management of system components and databases (e.g., add a new replica node), (ii) increased scalability compared to state-of-the-art approaches, and (iii) support for elasticity by suspending and resuming database replicas online without service interruption. Michael A. Georgiou, Michael Panayiotou, Lambros Odysseos, Aristodemos Paphitis, Michael Sirivianos, Herodotos Herodotou |
SIGMOD Conference | 6 |
| 2021 | Trident: Task Scheduling over Tiered Storage Systems in Big Data PlatformsabstractThe recent advancements in storage technologies have popularized the use of tiered storage systems in data-intensive compute clusters. The Hadoop Distributed File System (HDFS), for example, now supports storing data in memory, SSDs, and HDDs, while OctopusFS and hatS offer fine-grained storage tiering solutions. However, the task schedulers of big data platforms (such as Hadoop and Spark) will assign tasks to available resources only based on data locality information, and completely ignore the fact that local data is now stored on a variety of storage media with different performance characteristics. This paper presents Trident, a principled task scheduling approach that is designed to make optimal task assignment decisions based on both locality and storage tier information. Trident formulates task scheduling as a minimum cost maximum matching problem in a bipartite graph and uses a standard solver for finding the optimal solution. In addition, Trident utilizes two novel pruning algorithms for bounding the size of the graph, while still guaranteeing optimality. Trident is implemented in both Spark and Hadoop, and evaluated extensively using a realistic workload derived from Facebook traces as well as an industry-validated benchmark, demonstrating significant benefits in terms of application performance and cluster efficiency. Herodotos Herodotou, Elena Kakoulli |
Proc. VLDB Endow. | 1 |
| 2020 | A Streaming Machine Learning Framework for Online Aggression Detection on TwitterabstractThe rise of online aggression on social media is evolving into a major point of concern. Several machine and deep learning approaches have been proposed recently for detecting various types of aggressive behavior. However, social media are fast paced, generating an increasing amount of content, while aggressive behavior evolves over time. In this work, we introduce the first, practical, real-time framework for detecting aggression on Twitter via embracing the streaming machine learning paradigm. Our method adapts its ML classifiers in an incremental fashion as it receives new annotated examples and is able to achieve the same (or even higher) performance as batch-based ML models, with over 90% accuracy, precision, and recall. At the same time, our experimental analysis on real Twitter data reveals how our framework can easily scale to accommodate the entire Twitter Firehose (of 778 million tweets per day) with only 3 commodity machines. Finally, we show that our framework is general enough to detect other related behaviors such as sarcasm, racism, and sexism in real time. Herodotos Herodotou, Despoina Chatzakou, Nicolas Kourtellis |
IEEE BigData | 1 |
| 2019 | Automating Distributed Tiered Storage Management in Cluster ComputingabstractData-intensive platforms such as Hadoop and Spark are routinely used to process massive amounts of data residing on distributed file systems like HDFS. Increasing memory sizes and new hardware technologies (e.g., NVRAM, SSDs) have recently led to the introduction of storage tiering in such settings. However, users are now burdened with the additional complexity of managing the multiple storage tiers and the data residing on them while trying to optimize their workloads. In this paper, we develop a general framework for automatically moving data across the available storage tiers in distributed file systems. Moreover, we employ machine learning for tracking and predicting file access patterns, which we use to decide when and which data to move up or down the storage tiers for increasing system performance. Our approach uses incremental learning to dynamically refine the models with new file accesses, allowing them to naturally adjust and adapt to workload changes over time. Our extensive evaluation using realistic workloads derived from Facebook and CMU traces compares our approach with several other policies and showcases significant benefits in terms of both workload performance and cluster efficiency. Herodotos Herodotou, Elena Kakoulli |
Proc. VLDB Endow. | 1 |
| 2019 | Speedup Your Analytics: Automatic Parameter Tuning for Databases and Big Data SystemsabstractDatabase and big data analytics systems such as Hadoop and Spark have a large number of configuration parameters that control memory distribution, I/O optimization, parallelism, and compression. Improper parameter settings can cause significant performance degradation and stability issues. However, regular users and even expert administrators struggle to understand and tune them to achieve good performance. In this tutorial, we review existing approaches on automatic parameter tuning for databases, Hadoop, and Spark, which we classify into six categories: rule-based, cost modeling, simulation-based, experiment-driven, machine learning, and adaptive tuning. We describe the foundations of different automatic parameter tuning algorithms and present pros and cons of each approach. We also highlight real-world applications and systems, and identify research challenges for handling cloud services, resource heterogeneity, and real-time analytics. Jiaheng Lu, Yuxing Chen 0003, Herodotos Herodotou, Shivnath Babu |
Proc. VLDB Endow. | 3 |
| 2018 | OctopusFS in Action: Tiered Storage Management for Data Intensive ComputingabstractThe continuous improvements in memory, storage devices, and network technologies of commodity hardware introduce new challenges and opportunities in tiered storage management. Whereas past work is exploiting storage tiers in pairs or for specific applications, OctopusFS---a novel distributed file system that is aware of the underlying storage media---offers a comprehensive solution to managing multiple storage tiers in a distributed setting. OctopusFS contains auto-mated data-driven policies for managing the placement and retrieval of data across the nodes and storage tiers of the cluster. It also exposes the network locations and storage tiers of the data in order to allow higher-level systems to make locality-aware and tier-aware decisions. This demonstration will showcase the web interface of OctopusFS, which enables users to (i) view detailed utilization information for the various storage tiers and nodes, (ii) browse the directory namespace and perform file-related actions, and (iii) execute caching-related operations while observing their performance impact on MapReduce and Spark workloads. Elena Kakoulli, Nikolaos Karmiris, Herodotos Herodotou |
Proc. VLDB Endow. | 3 |
| 2017 | OctopusFS: A Distributed File System with Tiered Storage ManagementabstractThe ever-growing data storage and I/O demands of modern large-scale data analytics are challenging the current distributed storage systems. A promising trend is to exploit the recent improvements in memory, storage media, and networks for sustaining high performance and low cost. While past work explores using memory or SSDs as local storage or combine local with network-attached storage in cluster computing, this work focuses on managing multiple storage tiers in a distributed setting. We present OctopusFS, a novel distributed file system that is aware of heterogeneous storage media (e.g., memory, SSDs, HDDs, NAS) with different capacities and performance characteristics. The system offers a variety of pluggable policies for automating data management across the storage tiers and cluster nodes. The policies employ multi-objective optimization techniques for making intelligent data management decisions based on the requirements of fault tolerance, data and load balancing, and throughput maximization. At the same time, the storage media are explicitly exposed to users and applications, allowing them to choose the distribution and placement of replicas in the cluster based on their own performance and fault tolerance requirements. Our extensive evaluation shows the immediate benefits of using OctopusFS with data-intensive processing systems, such as Hadoop and Spark, in terms of both increased performance and better cluster utilization. Elena Kakoulli, Herodotos Herodotou |
SIGMOD Conference | 2 |
| 2016 | Enhancing Virtual Reality Systems with Smart Wearable DevicesabstractThe proliferation of wearable and smartphone devices with embedded sensors has enabled researchers and engineers to study and understand user behavior at an extremely high fidelity, particularly for use in industries such as entertainment, health, and retail. However, identified user patterns are yet to be integrated into modern systems with immersive capabilities, such as VR systems, which still remain constrained by limited application interaction models exposed to developers. In this paper, we present Smart VR, a platform that allows developers to seamlessly incorporate user behavior into VR apps. We present the high-level architecture of Smart VR, and show how it facilitates communication, data acquisition, and context recognition between smart wearable devices and mediator systems (e.g., smartphones, tablets, PCs). We demonstrate Smart VR in the context of a VR app for retail stores to show how it can be used to substitute the requirement of cumbersome input devices (e.g., mouse, keyboard) with more natural means of user-app interaction (e.g., user gestures such as swiping and tapping) to improve user experience. Salah Eddin Alshaal, Stylianos Michael, Andreas Pamboris, Herodotos Herodotou, George Samaras, Panayiotis Andreou |
MDM | 4 |
| 2014 | PStorM: Profile Storage and Matching for Feedback-Based Tuning of MapReduce JobsabstractThe MapReduce programming model has become widely adopted for large scale analytics on big data. MapReduce systems such as Hadoop have many tuning parameters, many of which have a significant impact on performance. The map and reduce functions that make up a MapReduce job are developed using arbitrary programming constructs, which makes them black-box in nature and prevents users from making good parameter tuning decisions for a submitted MapReduce job. Some research projects, such as the Starfish system, aim to provide automatic tuning decisions for input MapReduce jobs. Starfish and similar systems rely on an execution profile of a MapReduce job being tuned, and this profile is assumed to come from a previous execution of the same job. Managing these execution profiles has not been previously studied. This thesis presents PStorM, a profile store that organizes the collected profiling information in a scalable and extensible data model, and a profile matcher that accurately picks the relevant profiling information even for previously unseen MapReduce jobs. PStorM is currently integrated with the Starfish system, providing the necessary profiles that Starfish needs to tune a job. The thesis presents results that demonstrate the accuracy and efficiency of profile matching. The results also show that the profiles returned by PStorM lead to Starfish tuning decisions that are as good as the decisions made by profiles collected from a previous run of the job. Mostafa Ead, Herodotos Herodotou, Ashraf Aboulnaga, Shivnath Babu |
EDBT | 2 |
| 2014 | Scalable near real-time failure localization of data center networksabstractLarge-scale data center networks are complex---comprising several thousand network devices and several hundred thousand links---and form the critical infrastructure upon which all higher-level services depend on. Despite the built-in redundancy in data center networks, performance issues and device or link failures in the network can lead to user-perceived service interruptions. Therefore, determining and localizing user-impacting availability and performance issues in the network in near real time is crucial. Traditionally, both passive and active monitoring approaches have been used for failure localization. However, data from passive monitoring is often too noisy and does not effectively capture silent or gray failures, whereas active monitoring is potent in detecting faults but limited in its ability to isolate the exact fault location depending on its scale and granularity. Herodotos Herodotou, Bolin Ding, Shobana Balakrishnan, Geoff Outhred, Percy Fitter |
KDD | 1 |
| 2012 | Stubby: A Transformation-based Optimizer for MapReduce WorkflowsabstractThere is a growing trend of performing analysis on large datasets using workflows composed of MapReduce jobs connected through producer-consumer relationships based on data. This trend has spurred the development of a number of interfaces---ranging from program-based to query-based interfaces---for generating MapReduce workflows. Studies have shown that the gap in performance can be quite large between optimized and unoptimized workflows. However, automatic cost-based optimization of MapReduce workflows remains a challenge due to the multitude of interfaces, large size of the execution plan space, and the frequent unavailability of all types of information needed for optimization. We introduce a comprehensive plan space for MapReduce workflows generated by popular workflow generators. We then propose Stubby , a cost-based optimizer that searches selectively through the subspace of the full plan space that can be enumerated correctly and costed based on the information available in any given setting. Stubby enumerates the plan space based on plan-to-plan transformations and an efficient search algorithm. Stubby is designed to be extensible to new interfaces and new types of optimizations, which is a desirable feature given how rapidly MapReduce systems are evolving. Stubby's efficiency and effectiveness have been evaluated using representative workflows from many domains. Harold Lim, Herodotos Herodotou, Shivnath Babu |
Proc. VLDB Endow. | 2 |
| 2011 | Starfish: A Self-tuning System for Big Data Analytics
Herodotos Herodotou, Harold Lim, Nedyalko Borisov, Fatma Bilgen Cetin, Shivnath Babu |
CIDR | 1 |
| 2011 | Query optimization techniques for partitioned tablesabstractTable partitioning splits a table into smaller parts that can be accessed, stored, and maintained independent of one another. From their traditional use in improving query performance, partitioning strategies have evolved into a powerful mechanism to improve the overall manageability of database systems. Table partitioning simplifies administrative tasks like data loading, removal, backup, statistics maintenance, and storage provisioning. Query language extensions now enable applications and user queries to specify how their results should be partitioned for further use. However, query optimization techniques have not kept pace with the rapid advances in usage and user control of table partitioning. We address this gap by developing new techniques to generate efficient plans for SQL queries involving multiway joins over partitioned tables. Our techniques are designed for easy incorporation into bottom-up query optimizers that are in wide use today. We have prototyped these techniques in the PostgreSQL optimizer. An extensive evaluation shows that our partition-aware optimization techniques, with low optimization overhead, generate plans that can be an order of magnitude better than plans produced by current optimizers. Herodotos Herodotou, Nedyalko Borisov, Shivnath Babu |
SIGMOD Conference | 1 |
| 2011 | Profiling, What-if Analysis, and Cost-based Optimization of MapReduce Programs
Herodotos Herodotou, Shivnath Babu |
Proc. VLDB Endow. | 1 |
| 2011 | MapReduce Programming and Cost-based Optimization? Crossing this Chasm with Starfish
Herodotos Herodotou, Shivnath Babu |
Proc. VLDB Endow. | 1 |
| 2010 | Xplus: A SQL-Tuning-Aware Query OptimizerabstractThe need to improve a suboptimal execution plan picked by the query optimizer for a repeatedly run SQL query arises routinely. Complex expressions, skewed or correlated data, and changing conditions can cause the optimizer to make mistakes. For example, the optimizer may pick a poor join order, overlook an important index, use a nested-loop join when a hash join would have done better, or cause an expensive, but avoidable, sort to happen. SQL tuning is also needed while tuning multi-tier services to meet service-level objectives. The difficulty of SQL tuning can be lessened considerably if users and higher-level tuning tools can tell the optimizer: "I am not satisfied with the performance of the plan p being used for the query Q that runs repeatedly. Can you generate a (δ%) better plan?" This paper designs, implements, and evaluates Xplus which, to our knowledge, is the first query optimizer to provide this feature. Xplus goes beyond the traditional plan-first-execute-next approach: Xplus runs some (sub)plans proactively, collects monitoring data from the runs, and iterates. A nontrivial challenge is in choosing a small set of plans to run. Xplus guides this process efficiently using an extensible architecture comprising SQL-tuning experts with different goals, and a policy to arbitrate among the experts. We show the effectiveness of Xplus on real-life tuning scenarios created using TPC-H queries on a PostgreSQL database. Herodotos Herodotou, Shivnath Babu |
Proc. VLDB Endow. | 1 |
| 2009 | RIOT: I/O-Efficient Numerical Computing without SQL
Yi Zhang 0011, Herodotos Herodotou, Jun Yang 0001 |
CIDR | 2 |