EDBT 2026 Demo / reviewers in the wild / expert
Vana Kalogeraki
dblp:k/VanaKalogeraki
· DBLP profile ↗
54ranked-venue papers in the field
2as first author
11since 2021 · last 2026
0000-0002-6421-9947ORCID · verified
Domains — venue-derived; a paper can count in several
Database Systems & Data Management · 35 (1 first)Big Data, Cloud & Distributed Data Systems · 12Data Mining & Knowledge Discovery · 5Information Retrieval & Web Search · 2 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | The CoDiet Messaging App: A Mobile App for Transmitting Personalized Dietary Event Alerts
Michalis Tsenos, Christos Gunopulos, Vana Kalogeraki |
MDM | 3 |
| 2025 | Trustworthy Scheduling for Big Data Applications
Dimitrios Tomaras, Vana Kalogeraki, Dimitrios Gunopulos |
IEEE Big Data | 2 |
| 2025 | CONCERTO: Constrained Linear Multi-Objective Routing Path OptimizationabstractRouting through large urban areas is a daily problem for the vast majority of the commuters. In this paper, we consider the extension of the routing problem where we want to find the optimal path for different objectives. Specifically, we consider the problem for finding paths that are quick, safe or efficient in the use of resources. Multi-objective routing path optimization is an active research field in the area of path optimization. We present an efficient technique for finding the optimal path which satisfies user defined constraints. Experimental evaluation leads us to the conclusion that multi-objective routing path optimization can be deployed efficiently in terms of computational cost and solution accuracy. We perform experiments comparing our work with a state of the art technique which concerns the multi-objective shortest path optimization problem. Athanasios Makropoulos, Dimitrios Gunopulos, Vana Kalogeraki, Nikolaos Zygouras |
MDM | 3 |
| 2024 | TIMBER: On supporting data pipelines in Mobile Cloud EnvironmentsabstractThe radical advances in mobile computing, the IoT technological evolution along with cyberphysical components (e.g., sensors, actuators, control centers) have led to the development of smart city applications that generate raw or preprocessed data, enabling workflows involving the city to better sense the urban environment and support citizens’ everyday lives. Recently, a new era of Mobile Edge Cloud (MEC) infrastructures has emerged to support smart city applications that aim to address the challenges raised due to the spatio-temporal dynamics of the urban crowd as well as bring scalability and on-demand computing capacity to urban system applications for timely response. In these, resource capabilities are distributed at the edge of the network and in close proximity to end-users, making it possible to perform computation and data processing at the network edge. However, there are important challenges related to real-time execution, not only due to the highly dynamic and transient crowd, the bursty and highly unpredictable amount of requests but also due to the resource constraints imposed by the Mobile Edge Cloud environment. In this paper, we present TIMBER, our framework for efficiently supporting mobile daTa processing pIpelines in MoBile cloud EnviRonments that effectively addresses the aforementioned challenges. Our detailed experimental results illustrate that our approach can reduce the operating costs by 66.245% on average and achieve up to 96.4% similar throughput performance for agnostic workloads. Dimitrios Tomaras, Michalis Tsenos, Vana Kalogeraki, Dimitrios Gunopulos |
MDM | 3 |
| 2023 | A holistic approach for modeling and predicting bike demand
Dimitrios Tomaras, Ioannis Boutsis, Vana Kalogeraki |
Inf. Syst. | 3 |
| 2022 | Practical Privacy Preservation in a Mobile Cloud EnvironmentabstractThe proliferation of smartphone devices has led to the emergence of powerful user services from enabling interactions with friends and business associates to mapping, finding nearby businesses and alerting users in real-time. Moreover, users do not realize that continuously sharing their trajectory data with online systems may end up revealing a great amount of information in terms of their behavior, mobility patterns and social relationships. Thus, addressing these privacy risks is a fundamental challenge. In this work, we present$TP^{3}$, a Privacy Protection system for Trajectory analytics. Our contributions are the following: (1) we model a new type of attack, namely “social link exploitation attack”, (2) we utilize the coresets theory, a fast and accurate technique which approximates well the original data using a small data set, and running queries on the coreset produces similar results to the original data, and (3) we employ the Serverless computing paradigm to accommodate a set of privacy operations for achieving high system performance with minimized provisioning costs, while preserving the users' privacy. We have developed these techniques in our$TP^{3}$system that works with state-of-the-art trajectory analytics apps and applies different types of privacy operations. Our detailed experimental evaluation illustrates that our approach is both efficient and practical. Dimitrios Tomaras, Michalis Tsenos, Vana Kalogeraki |
MDM | 3 |
| 2022 | A Framework for Supporting Privacy Preservation Functions in a Mobile Cloud EnvironmentabstractThe problem of privacy protection of trajectory data has received increasing attention in recent years with the significant grow in the volume of users that contribute trajectory data with rich user information. This creates serious privacy concerns as exposing an individual's privacy information may result in attacks threatening the user's safety. In this demonstration we present$TP^{3}$a novel practical framework for supporting trajectory privacy preservation in Mobile Cloud Environments (MCEs). In$TP^{3}$, non-expert users submit their trajectories and the system is responsible to determine their privacy exposure before sharing them to data analysts in return for various benefits, e.g. better recommendations.$TP^{3}$makes a number of contributions: (a) It evaluates the privacy exposure of the users utilizing various privacy operations, (b) it is latency-efficient as it implements the privacy operations as serverless functions which can scale automatically to serve an increasing number of users with low latency, and (c) it is practical and cost-efficient as it exploits the serverless model to adapt to the demands of the users with low operational costs for the service provider. Finally,$TP^{3}$'s Web-UI provides insights to the service provider regarding the performance and the respective revenue from the service usage, while enabling the user to submit the trajectories with recommended preferences of privacy. Dimitrios Tomaras, Michalis Tsenos, Vana Kalogeraki |
MDM | 3 |
| 2021 | Evaluating Actions in Sports Analytics with Deep LearningabstractIn recent years, Sports Analytics has attracted a lot of interest in the research community. This presents an exciting opportunity as there is a plethora of complex, real-time events that can be analyzed to gain better insights or assist in decision-making. In this paper, we address the problem of evaluating actions of players in a football game using Deep Learning techniques. We propose and evaluate four different Deep Learning models, namely, Fully Convolutional Neural Networks, Long Short Term Models and combinations of them, to make more robust predictions for unseen data. Our detailed experimental evaluation illustrates that our approach can be successfully applied to structural data to compute the probabilities of scoring and conceding a goal in future actions. Dimitrios Klagkos, Vana Kalogeraki |
IEEE BigData | 2 |
| 2021 | Cherry: A Distributed Task-Aware Shuffle Service for Serverless AnalyticsabstractWhile there has been a lot of effort in recent years in optimising Big Data systems like Apache Spark and Hadoop, the all-to-all transfer of data between a MapReduce computation step, i.e., the shuffle data mechanism between cluster nodes remains always a serious bottleneck. In this work, we present Cherry, an open-source distributed task-aware Caching sHuffle sErvice for seRveRless analYtics. Our thorough experiments on a cloud testbed using realistic and synthetic workloads showcase that Cherry can achieve an almost 23% to 39% reduction in completion of the reduce stage with small shuffle block sizes, a 10% reduction in execution time on real workloads, while it can efficiently handle Spark execution failures with a constant task time re-computation overhead compared to existing approaches. Nikolaos Nikitas, Ioannis Konstantinou, Vana Kalogeraki, Nectarios Koziris |
IEEE BigData | 3 |
| 2021 | Dynamic Rate Control for Topic-based Pub/Sub SystemsabstractPub/sub systems have been widely utilized in the industry as the connecting piece between high rate producers or mobile clients and end-services, due to their ability to handle messages of high volume and velocity and achieve high throughput. Apache Kafka is one of the most popular Big Data messaging systems. Although Kafka's modular Consumer API allows businesses to take advantage of the rich set of features it provides, it often suffers from improper configured topics/partitions or load imbalances due to spikes or high volume of messages injected in the same topic. In our research our goal is to develop a novel framework which investigates smart queuing mechanisms to overcome these issues and deal effectively with sudden bursts and overloads which are frequently experienced in these systems. Our approach provides rate control capabilities to the Kafka's consumer API which allows us to effectively meet the requirements of different end-services without interference among them. Michalis Tsenos, Vana Kalogeraki |
MDM | 2 |
| 2021 | A General Framework for First Story Detection Utilizing Entities and Their RelationsabstractNews portals, such as Yahoo News or Google News, collect large amounts of news articles from a variety of sources on a daily basis. Only a small portion of these documents can be selected and displayed on the homepage. Thus, there is a strong preference for major, recent events. In this work, we propose a scalable First Story Detection (FSD) pipeline that identifies fresh news. This pipeline is used in order to instantiate a variety of FSD approaches. In addition we suggest a novel FSD technique that in comparison to existing systems, relies on relation extraction algorithms and exploits the named entities and their relations in order to decide about the freshness of an article. We evaluate our technique by instantiating existing state of art FSD techniques within our generic pipeline. As ground truth we use multiple datasets that cover different categories. Experimental results demonstrate that our FSD method in many cases provides an improvement over state-of-the-art techniques. In addition, we show using a large synthetic dataset that our general FSD pipeline has constant space and time requirements and is suitable for very high volume streams. Nikolaos Panagiotou, Cem Akkaya, Kostas Tsioutsiouliklis, Vana Kalogeraki, Dimitrios Gunopulos |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2020 | Cost-Aware Influence Maximization in Multi-Attribute NetworksabstractThe popularity of Online Social Networks (OSNs) led to numerous applications that harness the benefits of immediate information exchange among numerous users of the network. This property is particularly utilized by OSNs campaigns that exploit of the "word-of-mouth" effect exhibited in the network. The problem Influence Maximization (IM), i.e., identifying the appropriate subset of users to initiate the propagation of a specific campaign, is widely studied in the literature. Various models have been proposed to capture the way information propagates in the network, yet a unified model that considers the important parameters of (i) correlation among campaigns propagating in the network and (ii) the different attributes of the propagating entities coupled with the users distinct preferences in certain attributes, is lacking. Additionally, the majority of the works assume uniform costs and revenues among users. Finally, the IM problem is addressed solely offline, i.e., after the seed selection process no further action is defined. In this work we propose the Multi-Attribute Correlated Independent Cascade (MAC-IC) propagation model to tackle the aforementioned limitations of existing propagation models. Given the MAC-IC model we design a two-phase Greedy Offer Selection (GOES) algorithm to address the IM problem under variable costs and revenues generated by the users. During the offline phase, the seeds to initiate the propagation of a specific item are identified. In the online phase, the propagation is monitored, the blockers are detected and real-time incentives may be offered to convince them to participate in the campaign. We prove that the GOES offline phase achieves an approximation ratio of 1 - 1/e. Through an extensive experimental evaluation we demonstrate the efficiency of our approach compared to state-of-the-art schemes. Juliana Litou, Vana Kalogeraki |
IEEE BigData | 2 |
| 2019 | Attendance Maximization for Successful Social Event Planning
Nikos Bikakis, Vana Kalogeraki, Dimitrios Gunopulos |
EDBT | 2 |
| 2018 | Influence Maximization in Evolving Multi-Campaign EnvironmentsabstractOnline Social Networks (OSNs) are being extensively used in a variety of campaigns, with the objective to raise the awareness of the audience regarding a specific piece of information, e.g, product awareness, political positions, etc. An important characteristic is the correlations of different strength and types, i.e., either positive or negative, that appear among the plethora of information diffused in the network. Therefore, the breadth of the diffusion of a specific campaign is significantly affected by the spread of the correlated information. In addition to the correlations of the diffusing information, another important factor affecting the spread of a specific information, is the network connectivity. Social networks are constantly evolving, with links among users being continuously formed or discontinued and communication patterns being highly fluctuating. The aim of this work is to address the problem of Influence Maximization in Online Social Networks by bridging the gap of the existing literature regarding the plethora of correlated information simultaneously propagating in the network and the dynamic communication patterns observed among users in the network. Towards this, we propose a mechanism that estimates users' interaction probabilities and formulate a propagation model that captures how the exposure to many correlated campaigns impacts on the users' behavior to support any of them. We finally design a greedy seed selection approach that approximates the optimal solution at a ratio of 1 - 1/e and prove through extensive experimental evaluation the superiority of our approach compared to state-of-the-art approaches. Juliana Litou, Vana Kalogeraki |
IEEE BigData | 2 |
| 2018 | Scalable Distributed Top-k Join Queries in Topic-Based Pub/Sub SystemsabstractIn this paper, we provide a novel approach that enables the execution of top-k join queries over sliding windows in a way that reduces the amount of data that need to be analyzed by the stream processing operators. The main idea is that brokers individually invoke the query on their received messages and forward the top-k results to a stream processing operator that performs the merging of the results and provides to the end-user the final top-k results. Moreover, our system exploits the Bayesian Optimization technique to determine automatically the number of top-k results that should be provided by each broker. Our approach has been developed in the Kappa architecture that exploits topic-based scalable publish/subscribe (pub/sub) systems like Apache Kafka to efficiently forward the high volume of incoming messages to distributed processing systems (i.e., Apache Spark or Apache Flink) that perform the batch and stream analytics operations. Our detailed experimental evaluation on our local cluster illustrates that we can efficiently execute top-k join queries on our system with high accuracy and low latency. Nikos Zacheilas, Dimitris Dedousis, Vana Kalogeraki |
IEEE BigData | 3 |
| 2018 | Social Event SchedulingabstractA major challenge for social event organizers (e.g., event planning and marketing companies, venues) is attracting the maximum number of participants, since it has great impact on the success of the event, and, consequently, the expected gains (e.g., revenue, artist/brand publicity). In this paper, we introduce the Social Event Scheduling (SES) problem, which schedules a set of social events considering user preferences and behavior, events' spatiotemporal conflicts, and competing events, in order to maximize the overall number of attendees. We show that SES is strongly NP-hard, even in highly restricted instances. To cope with the hardness of the SES problem we design a greedy approximation algorithm. Finally, we evaluate our method experimentally using a dataset from the Meetup event-based social network. Nikos Bikakis, Vana Kalogeraki, Dimitrios Gunopulos |
ICDE | 2 |
| 2018 | Dione: A Framework for Automatic Profiling and Tuning Big Data ApplicationsabstractIn this demonstration we presentDionea novel framework for automatic profiling and tuning big data applications. Our system allows a non-expert user to submit Spark or Flink applications to his/her cluster and Dione automatically determines the impact of different configuration parameters on the application's execution time and monetary cost. Dione is the first framework that exploits similarities in the execution plans of different applications to narrow down the amount of profiling runs that are required for building prediction models that capture the impact of the configuration parameters on the metrics of interest. Dione exploits these prediction models to tune the configuration parameters in a way that minimizes the application's execution time or the user's budget. Finally, Dione's Web-UI visualizes the impact of the configuration parameters on the execution time and the monetary cost, and enables the user to submit the application with the recommended parameters' values. Nikos Zacheilas, Stathis Maroulis, Thanasis Priovolos, Vana Kalogeraki, Dimitrios Gunopulos |
ICDE | 4 |
| 2018 | A Cost-Aware Incentive Mechanism in Mobile Crowdsourcing SystemsabstractThe rapid growth of ubiquitous mobile smart devices has led to the creation of a new era of mobile crowdsourcing applications, where human workers participate and perform tasks in exchange of a monetary reward. Such crowdsourcing systems can play a vital role during emergency events, where fast and accurate responses are needed. However, a commonly ignored aspect is how the price (i.e. the reward paid to workers) must be set in order for the system to meet two important requirements: (i) to timely receive an adequate number of responses which is crucial during emergencies, and (ii) to meet budget constraints. In the majority of the existing systems, the price per task is set up-front and remains unchanged for all upcoming tasks, leading to either higher monetary cost than necessary or to significantly larger latency than expected. In this work, we provide a formulation based on Kalman Filters that enables the system to estimate the user/worker behavior, i.e., the likelihood over time for a user to provide answers for a specific reward. Specifically, we focus on the problem of developing an adaptive pricing policy to incentivize the users to rapidly provide their responses. Our mechanism can be adjusted dynamically to bridge the gap among the users' behavior and the system's needs so as to maximize the overall utility of the system. We simulate our model and through extensive experimental evaluation we show how our system performs and provides benefits to both the users and the system operator. Ellen Mitsopoulou, Ioannis Boutsis, Vana Kalogeraki, Jia Yuan Yu |
MDM | 3 |
| 2018 | Crowd-Based Ecofriendly Trip PlanningabstractIn recent years we have witnessed a growing interest in trip planning systems aiming at organizing daily travel schedules in smart cities. Such systems use specialized engines to find optimal means of transport between two geospatial endpoints to provide recommendations to citizens for short routes across the city. At the same time, alternative means of transportation, such as bike sharing systems, have enjoyed tremendous success since they offer a green and facile solution for daily commuters and tourists. However, one major challenge of the bike sharing systems is that the distribution of bikes among the stations can be quite uneven during rush hours or due to topography. This often results in shortage of bikes and increasing numbers of disappointed users. Existing works in the literature are limited since they only focus on predicting the demand or apply a-posteriori methods for balancing the load of stations. Furthermore, none of these works consider the benefit of these systems in concert. In this work, we present "MOToR" (MultimOdal Trip Rebalancing), a system that builds upon the OpenTripPlanner framework to incorporate dynamic transit schedule data while balancing the availability of bikes among the bike stations. Our experimental evaluation shows that our approach is practical, efficient and outperforms state-of-the-art methods for route planning. Dimitrios Tomaras, Vana Kalogeraki, Thomas Liebig, Dimitrios Gunopulos |
MDM | 2 |
| 2017 | Dione: Profiling spark applications exploiting graph similarityabstractIn recent years distributed processing frameworks such as Apache Spark have been utilized for running big data applications. Predicting the application's execution time has been an important goal since it can help the end user to determine the necessary processing resources to be reserved. While there have been some previous works that examine the problem of profiling Spark applications, they mainly focus on specific application types (e.g., Machine learning applications) and rely on the existence of a large number of previous execution runs. In this work we aim at overcoming these limitations by minimizing the number of past execution runs needed for the profiling phase. Furthermore, we identify patterns of continuous identical dataset transformations between different applications to cope with the limited historical data availability. We propose an on-line profiling framework, called Dione, that estimates the running times of new applications, even if no historical data is available. Finally, in our detailed experimental evaluation, using practical workloads on our local cluster, we illustrate that our approach accurately predicts the execution times of Spark applications and requires 30% less training time and monetary cost compared to the current state-of-the-art techniques. Nikos Zacheilas, Stathis Maroulis, Vana Kalogeraki |
IEEE BigData | 3 |
| 2017 | Revealing the Hidden Links in Content Networks: An Application to Event DiscoveryabstractSocial networks have become the de facto online resource for people to share, comment on and be informed about events pertinent to their interests and livelihood, ranging from road traffic or an illness to concerts and earthquakes, to economics and politics. This has been the driving force behind research endeavors that analyse such data. In this paper, we focus on how Content Networks can help us identify events effectively. Content Networks incorporate both structural and content-related information of a social network in a unified way, at the same time, bringing together two disparate lines of research: graph-based and content-based event discovery in social media. We model interactions of two types of nodes, users and content, and introduce an algorithm that builds heterogeneous, dynamic graphs, in addition to revealing content links in the network's structure. By linking similar content nodes and tracking connected components over time, we can effectively identify different types of events. Our evaluation on social media streaming data suggests that our approach outperforms state-of-the-art techniques, while showcasing the significance of hidden links to the quality of the results. Antonia Saravanou, Ioannis Katakis 0001, George Valkanas, Vana Kalogeraki, Dimitrios Gunopulos |
CIKM | 4 |
| 2017 | Mining Urban Data (Part C)
Gennady L. Andrienko, Dimitrios Gunopulos, Yannis E. Ioannidis, Vana Kalogeraki, Ioannis Katakis 0001, Katharina Morik, Olivier Verscheure |
Inf. Syst. | 4 |
| 2017 | Efficient techniques for time-constrained information dissemination using location-based social networks
Juliana Litou, Ioannis Boutsis, Vana Kalogeraki |
Inf. Syst. | 3 |
| 2016 | Mining hidden constrained streams in practice: Informed search in dynamic filter spacesabstractIn this paper we tackle the recently proposed problem of hidden streams. In many situations, the data stream that we are interested in, is not directly accessible. Instead, part of the data can be accessed only through applying filters (e.g. keyword filtering). In fact this is the case of the most discussed social stream today, Twitter. The problem in this case is how to retrieve as many relevant documents as possible by applying the most appropriate set of filters to the original stream and, at the same time, respect a number of constrains (e.g. maximum number of filters that can be applied). In this work we introduce a search approach on a dynamic filter space. We utilize heterogeneous filters (not only keywords) making no assumptions about the attributes of the individual filters. We advance current research by considering realistically hard constraints based on real-world scenarios that require tracking of multiple dynamic topics. We demonstrate the effectiveness of our approaches on a set of topics of static and dynamic nature. The development of the approach was motivated by a real application. Our system is deployed in Dublin City's Traffic Management Center and allows the city officers to analyze large sources of heterogeneous data and identify events related to traffic as well as emergencies. Nikolaos Panagiotou, Ioannis Katakis 0001, Dimitrios Gunopulos, Vana Kalogeraki, Elizabeth Daly, Jia Yuan Yu, Brendan O'Brien |
ASONAM | 4 |
| 2016 | Context-aware point of interest recommendation using tensor factorizationabstractThe wide adoption of Location Based Social Networks along with advances in mobile technology, has brought forth as a core service the analysis of large volumes of location-based data for personalized Point of Interest (POIs) recommendations. The majority of the existing recommendation systems take advantage of Collaborative Filtering, but they fail to exploit the contextual information involved with POI checkins (i.e., POI category, location, or the checkin timestamp). In this paper we propose CoTF, a Context-Aware Point of Interest Recommendation system using Tensor Factorization, that aims at enhancing the user experience by providing personalized context aware POI recommendations. Our approach exploits Category-based context related to checkins without the need of any pre-or post-filtering techniques. Our detailed experimental evaluation using real data from the Foursquare location-based social network illustrates that our approach can efficiently produce personalized recommendations to users, while significantly reducing the training time compared to current state-of-the-art methods. Stathis Maroulis, Ioannis Boutsis, Vana Kalogeraki |
IEEE BigData | 3 |
| 2016 | LOCAl: a personalized cache mechanism for location-based social networksabstractRecommending nearby Points of Interest (POI) has received growing interest in mobile location-based networks today, where users share content embedded with location information. In this work, we propose a novel caching framework to support personalised proactive caching for mobile location-based social networks. We propose "LOCAI", which uses a probabilistic approach in order to predict the POIs that users will access and retrieve the appropriate data objects that will fulfill user preferences. Our detailed experimental evaluation, using data from the Foursquare location-based social network, illustrates that LOCAI minimizes the user latency to retrieve the data objects they are interested in, is efficient and practical. Dimitrios Tomaras, Ioannis Boutsis, Vana Kalogeraki, Dimitrios Gunopulos |
SIGSPATIAL/GIS | 3 |
| 2016 | State Detection Using Adaptive Human Sensor SamplingabstractWith the massive prevalence of smartphones, mobile social sensing systems in which humans acting as social sensors respond to geo-located crowdsourcing tasks, became extremely popular. Such systems can provide significant benefits particularly during crisis management and emergency situations. However, not only querying users can be extremely costly but also human sensors are mobile, subjective and their response delays can highly vary. In this paper we develop a social sensing system that performs sampling on mobile social sensors to achieve accurate and real-time detection of the state of emergency events. Our contributions are two-fold: (i) our approach can capture well emergencies even in large geographical regions, and (ii) our sampling approach considers the individual characteristics of the social sensors to maximize the probability of receiving accurate responses in a timely manner. We provide comprehensive experiments that indicate that our approach accurately identifies critical real-world events, has low overhead and reduces the classification error up to 90% compared to traditional approaches. Ioannis Boutsis, Vana Kalogeraki, Dimitrios Gunopulos |
HCOMP | 2 |
| 2016 | Real-Time and Cost-Effective Limitation of Misinformation PropagationabstractOnline Social Networks (OSNs) constitute one of the most important communication channels and are widely utilized as news sources. Information spreads widely and rapidly in OSNs through the word-of-mouth effect. However, it is not uncommon for misinformation to propagate in the network. Misinformation dissemination may lead to undesirable effects, especially in cases where the non-credible information concerns emergency events. Therefore, it is essential to timely limit the propagation of misinformation. Towards this goal, we suggest a novel propagation model, namely the Dynamic Linear Threshold (DLT) model, that effectively captures the way contradictory information, i.e., misinformation and credible information, propagates in the network. The DLT model considers the probability of a user alternating between competing beliefs, assisting in either the propagation of misinformation or credible news. Based on the DLT model, we formulate an optimization problem that aims in identifying the most appropriate subset of users to limit the spread of misinformation by initiating the propagation of credible information. Through extensive experimental evaluation we demonstrate that our approach outperforms its competitors. Juliana Litou, Vana Kalogeraki, Ioannis Katakis 0001, Dimitrios Gunopulos |
MDM | 2 |
| 2016 | INSIGHT: Dynamic Traffic Management Using Heterogeneous Urban Data
Nikolaos Panagiotou, Nikolaos Zygouras, Ioannis Katakis 0001, Dimitrios Gunopulos, Nikos Zacheilas, Ioannis Boutsis, Vana Kalogeraki, Stephen Lynch, Brendan O'Brien, Dermot Kinane, Jakub Marecek, Jia Yuan Yu, Rudi Verago, Elizabeth Daly, Nico Piatkowski, Thomas Liebig, Christian Bockermann, Katharina Morik, François Schnitzler, Matthias Weidlich 0001, Avigdor Gal, Shie Mannor, Hendrik Stange, Werner Halft, Gennady L. Andrienko |
ECML/PKDD (3) | 7 |
| 2016 | Intelligent Urban Data Monitoring for Smart Cities
Nikolaos Panagiotou, Nikolaos Zygouras, Ioannis Katakis 0001, Dimitrios Gunopulos, Nikos Zacheilas, Ioannis Boutsis, Vana Kalogeraki, Stephen Lynch, Brendan O'Brien |
ECML/PKDD (3) | 7 |
| 2016 | Mining Urban Data (Part B)
Gennady L. Andrienko, Dimitrios Gunopulos, Yannis E. Ioannidis, Vana Kalogeraki, Ioannis Katakis 0001, Katharina Morik, Olivier Verscheure |
Inf. Syst. | 4 |
| 2015 | Elastic complex event processing exploiting predictionabstractSupporting real-time, cost-effective execution of Complex Event processing applications in the cloud has been an important goal for many scientists in recent years. Distributed Stream Processing Systems (DSPS) have been widely adopted by major computing companies as a powerful approach for large-scale Complex Event processing (CEP). However, determining the appropriate degree of parallelism of the DSPS' components can be particularly challenging as the volume of data streams is becoming increasingly large, the rule set is becoming continuously complex, and the system must be able to handle such large data stream volumes in real-time, taking into consideration changes in the burstiness levels and data characteristics. In this paper we describe our solution to building elastic complex event processing systems on top of our distributed CEP system which combines two commonly used frameworks, Storm and Esper, in order to provide both ease of usage and scalability. Our approach makes the following contributions: (i) we provide a mechanism for predicting the load and latency of the Esper engines in upcoming time windows, and (ii) we propose a novel algorithm for automatically adjusting the number of engines to use in the upcoming windows, taking into account the cost and the performance gains of possible changes. Our detailed experimental evaluation with a real traffic monitoring application that analyzes bus traces from the city of Dublin indicates the benefits in the working of our approach. Our proposal outperforms the current state of the art technique in regards to the amount of tuples that it can process by four orders of magnitude. Nikos Zacheilas, Vana Kalogeraki, Nikolaos Zygouras, Nikolaos Panagiotou, Dimitrios Gunopulos |
IEEE BigData | 2 |
| 2015 | Insights on a Scalable and Dynamic Traffic Management SystemabstractComplex Event Processing (CEP) systems process large streams of data trying to detect events of interest. Traditional CEP systems, such as Esper, lack the required scalability and processing capability to cope with the constantly increasing amount of data that needs to be processed. Furthermore, user defined rules are static so changes in the monitored environment cannot be easily detected. In this paper we investigate the development of a scalable and dynamic traffic management system. Our work makes several contributions: We propose a novel system that combines Esper with a stream processing framework, Storm, in order to parallelize the processing of larger amounts of data. We propose a novel rules’ assignment algorithm for distributing Esper rules to the available CEP engines, in a way that maximizes the overall system’s throughput. Finally, our system adapts to changes of the environment by processing historical data via Hadoop and dynamically updating the Esper rules based on the generated results. Our work has been evaluated using real data, in several traffic monitoring scenarios for the city of Dublin. Our detailed experimental results indicate the benefits in the working of our approach and the significant increase in the system’s throughput when a large number of Esper rules were examined concurrently. Nikolaos Zygouras, Nikos Zacheilas, Vana Kalogeraki, Dermot Kinane, Dimitrios Gunopulos |
EDBT | 3 |
| 2015 | Personalized Event Recommendations Using Social NetworksabstractIn recent years we have observed a significant increase in the popularity of location-based social networks for exchanging news and experiences, sharing location information, or publishing real world events. One important challenge in such networks is to understand human crowd mobility behavior based on user social activities and interactions. In this paper we introduce PRESENT, our middleware that utilizes a Mixed Markov Model to extract the behavioral patterns of the users in social groups, to make personalized event recommendations. Our detailed experimental evaluation, using data from the Meet up location-based social network, illustrates that our approach is efficient, practical and achieves an average prediction for the user attendance of over 73%. Ioannis Boutsis, Stavroula Karanikolaou, Vana Kalogeraki |
MDM (1) | 3 |
| 2014 | Heterogeneous Stream Processing and Crowdsourcing for Urban Traffic ManagementabstractUrban traffic gathers increasing interest as cities become bigger, crowded and “smart”. We present a system for het-erogeneous stream processing and crowdsourcing supporting intelligent urban traffic management. Complex events related to traffic congestion (trends) are detected from heterogeneous sources involving fixed sensors mounted on intersections and mobile sensors mounted on public transport vehicles. To deal with data veracity, a crowdsourcing component handles and resolves sensor disagreement. Furthermore, to deal with data sparsity, a traffic modelling component offers information in areas with low sensor coverage. We demonstrate the system with a real-world use-case from Dublin city, Ireland. Alexander Artikis, Matthias Weidlich 0001, François Schnitzler, Ioannis Boutsis, Thomas Liebig, Nico Piatkowski, Christian Bockermann, Katharina Morik, Vana Kalogeraki, Jakub Marecek, Avigdor Gal, Shie Mannor, Dimitrios Gunopulos, Dermot Kinane |
EDBT | 9 |
| 2014 | Using Location-Based Social Networks for Time-Constrained Information DisseminationabstractLocation-based social networks have evolved into powerful tools in recent years. The ability to embed location information in Social Networks such as Facebook, Foursquare and Twitter creates exciting opportunities for users to disseminate and exchange geolocation information in a variety of domains. The problem of exploiting the social ties between the users for maximizing information reach has become a topic of great interest, and many challenges have to be met. In this work we study the problem of efficient information dissemination in location-based social networks under time constraints. The objective is to identify a subset of individuals to propagate the information and make intelligent route selection that can result in maximizing the reach within a time window. Our detailed experimental results illustrate the feasibility and performance of our approach. Juliana Litou, Ioannis Boutsis, Vana Kalogeraki |
MDM (1) | 3 |
| 2014 | Heterogeneous Stream Processing and Crowdsourcing for Traffic Monitoring: Highlights
François Schnitzler, Alexander Artikis, Matthias Weidlich 0001, Ioannis Boutsis, Thomas Liebig, Nico Piatkowski, Christian Bockermann, Katharina Morik, Vana Kalogeraki, Jakub Marecek, Avigdor Gal, Shie Mannor, Dermot Kinane, Dimitrios Gunopulos |
ECML/PKDD (3) | 9 |
| 2014 | Supporting historic queries in sensor networks with flash storage
Adam Ji Dou, Vana Kalogeraki, Dimitrios Gunopulos |
Inf. Syst. | 3 |
| 2013 | Self-adaptive event recognition for intelligent transport managementabstractIntelligent transport management involves the use of voluminous amounts of uncertain sensor data to identify and effectively manage issues of congestion and quality of service. In particular, urban traffic has been in the eye of the storm for many years now and gathers increasing interest as cities become bigger, crowded, and “smart”. In this work we tackle the issue of uncertainty in transportation systems stream reporting. The variety of existing data sources opens new opportunities for testing the validity of sensor reports and self-adapting the recognition of complex events as a result. We report on the use of a logic-based event reasoning tool to identify regions of uncertainty within a stream and demonstrate our method with a real-world use-case from the city of Dublin. Our empirical analysis shows the feasibility of the approach when dealing with voluminous and highly uncertain streams. Alexander Artikis, Matthias Weidlich 0001, Avigdor Gal, Vana Kalogeraki, Dimitrios Gunopulos |
IEEE BigData | 4 |
| 2013 | Mobile Stream Sampling under Time ConstraintsabstractThe proliferation of mobile networking and the increasing capabilities of smartphone devices in the recent years have resulted in transforming mobile smartphone devices into ubiquitous sensing platforms. In this new class of “Community-based Participatory Sensing” systems, users actively participate in the data collection and sharing for the benefit of the community, in a wide range of application areas from entertainment, to transportation, to environmental monitoring. These approaches, however, generate large amounts of transient data streams, leading to real-time computational challenges. In this paper we propose sampling algorithms on streams of mobile data generated by ubiquitous sensing devices that need to be processed under time constraints. In our approach users participate in the system by sensing and sharing streams of data. The system then uses a sampling mechanism to select a subset of data streams that preserves the characteristics of the stream data and provides the highest “information gain” to the system, given the real-time, budget and resource constraints. Detailed experimental results illustrate that our approach is practical, efficient and depicts good performance. Ioannis Boutsis, Vana Kalogeraki |
MDM (1) | 2 |
| 2013 | Guest editorial: special issue on mobile data management
Dipanjan Chakraborty 0001, Vana Kalogeraki, Mohamed F. Mokbel |
Distributed Parallel Databases | 2 |
| 2013 | SmartMonitor: Using Smart Devices to Perform Structural Health MonitoringabstractIn this demonstration, we are presenting SmartMonitor, a distributed Structural Health Monitoring (SHM) system consisting of smart devices. Over the last few years, the vast majority of smart devices is equipped with accelerometers that can be utilized towards building SHM systems with hundreds of nodes. We describe a scalable, fault-tolerant communication protocol, that performs best-effort time synchronization of the nodes and is used to implement a decentralized version of the popular peak-picking SHM method. The implemented interactive system can be easily installed in any accelerometer-equipped Android device and the user has a number of options for configuring the system or analyzing the collected data and computed outcomes. Dimitrios Kotsakos, Panos Sakkos, Vana Kalogeraki, Dimitrios Gunopulos |
Proc. VLDB Endow. | 3 |
| 2012 | Misco: A System for Data Analysis Applications on Networks of Smartphones Using MapReduceabstractThe recent years have seen a proliferation of community sensing or participatory sensing paradigms, where individuals rely on the use of smart and powerful mobile devices to collect, store and analyze data from everyday life. Due to this massive collection of the data, a key challenge to all such developments, is to provide a simple but efficient way to facilitate the programming of distributed applications on the embedded devices. We will demonstrate a novel system that provides a principled approach to developing distributed data clustering applications on networks of smartphones and other mobile devices. The system comprises three components: (a) a distributed framework, implemented on mobile phones that eases the programmability and deployment of applications on the devices using simple programming primitives, (b) a data gathering component that tracks the movement of wireless device users and collects sensor data (i.e., GPS and accelerometer sensor data), and (c) a distributed data clustering algorithm that allows users to combine their individual data, that is distributed and energy efficient. Using a road traffic monitoring application we demonstrate how MISCO can efficiently identify anomalies in the road surface conditions and illustrate that our system is practical and has low energy and resource overhead. Theofilos Kakantousis, Ioannis Boutsis, Vana Kalogeraki, Dimitrios Gunopulos, Giorgos Gasparis, Adam Ji Dou |
MDM | 3 |
| 2007 | Improving process models by discovering decision points
Sharmila Subramaniam, Vana Kalogeraki, Dimitrios Gunopulos, Fabio Casati, Malú Castellanos, Umeshwar Dayal, Mehmet Sayal |
Inf. Syst. | 2 |
| 2007 | Efficient Approximate Query Processing in Peer-to-Peer NetworksabstractPeer-to-peer (P2P) databases are becoming prevalent on the Internet for distribution and sharing of documents, applications, and other digital media. The problem of answering large-scale ad hoc analysis queries, for example, aggregation queries, on these databases poses unique challenges. Exact solutions can be time consuming and difficult to implement, given the distributed and dynamic nature of P2P databases. In this paper, we present novel sampling-based techniques for approximate answering of ad hoc aggregation queries in such databases. Computing a high-quality random sample of the database efficiently in the P2P environment is complicated due to several factors: the data is distributed (usually in uneven quantities) across many peers, within each peer, the data is often highly correlated, and, moreover, even collecting a random sample of the peers is difficult to accomplish. To counter these problems, we have developed an adaptive two-phase sampling approach based on random walks of the P2P graph, as well as block-level sampling techniques. We present extensive experimental evaluations to demonstrate the feasibility of our proposed solution. Benjamin Arai, Gautam Das 0001, Dimitrios Gunopulos, Vana Kalogeraki |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2006 | Approximating Aggregation Queries in Peer-to-Peer NetworksabstractPeer-to-peer databases are becoming prevalent on the Internet for distribution and sharing of documents, applications, and other digital media. The problem of answering large scale, ad-hoc analysis queries ― e.g., aggregation queries ― on these databases poses unique challenges. Exact solutions can be time consuming and difficult to implement given the distributed and dynamic nature of peer-to-peer databases. In this paper we present novel sampling-based techniques for approximate answering of ad-hoc aggregation queries in such databases. Computing a high-quality random sample of the database efficiently in the P2P environment is complicated due to several factors ― the data is distributed (usually in uneven quantities) across many peers, within each peer the data is often highly correlated, and moreover, even collecting a random sample of the peers is difficult to accomplish. To counter these problems, we have developed an adaptive two-phase sampling approach, based on random walks of the P2P graph as well as block-level sampling techniques. We present extensive experimental evaluations to demonstrate the feasibility of our proposed solutio Benjamin Arai, Gautam Das 0001, Dimitrios Gunopulos, Vana Kalogeraki |
ICDE | 4 |
| 2006 | Efficient Online State Tracking Using Sensor NetworksabstractSensor networks are being deployed for tracking events of interest in many environmental or monitoring applications. Because of their distributed nature of operation, a challenging issue is how to accurately identify the aggregate state of the phenomenon that is being observed. This work presents an online mechanism for efficiently determining the overall network status, employing distributed operations that minimize the communication costs. Experiments on real data, suggest that the proposed metholology can be a viable solution for real world systems. Maria Halkidi, Vana Kalogeraki, Dimitrios Gunopulos, Demetris Zeinalipour, Michail Vlachos |
MDM | 2 |
| 2006 | Online Outlier Detection in Sensor Data Using Non-Parametric Models
Sharmila Subramaniam, Themis Palpanas, Vana Kalogeraki, Dimitrios Gunopulos |
VLDB | 4 |
| 2005 | Applying LVQ Techniques to Compress Historical Information in Sensor NetworksabstractSummary form only given. In the emerging area of wireless sensor networks, a typical challenge is to retrieve historical information from the sensor nodes. We propose a new technique, called adaptive learning vector quantization (ALVQ), to compress this historical information. Our technique is based on the following two observations: (1) in sensor networks, the historical information exhibits similar patterns over time; and (2) different measurements are intrinsically correlated. Our algorithm works as follows: first, the codebook is obtained through a LVQ (learning vector quantization), which adjusts the codebook to be nearer to the optimal codebook. Second, ALVQ compresses the codebook update data pieces and transfers the compressed information to the base station. Using 2-level piece-wise regression, ALVQ can compress the updates with high precision while saving more bandwidth for data transmission in order to increase the quality of the approximation. In our experiments we used weather data to compare the performance of the ALVQ algorithm with the recently proposed SBR (self based regression) technique. Our experimental results demonstrate that the LVQ learning process significantly improves the quality of the codebook, thus increasing the regression precision. In addition the use of two-level regression for transmitting the codebook updates further minimizes the required bandwidth. Overall the ALVQ technique can achieve the same precision with SBR while using 75% of the bandwidth. Dimitrios Gunopulos, Stefano Lonardi, Vana Kalogeraki |
DCC | 4 |
| 2005 | MicroHash: An Efficient Index Structure for Flash-Based Sensor Devices
Demetris Zeinalipour, Vana Kalogeraki, Dimitrios Gunopulos, Walid A. Najjar |
FAST | 3 |
| 2005 | Mobile peer-to-peer computing: challenges, metrics and applicationsabstractThe pervasiveness of computers in our current society (transportation, e-commerce, home appliances, medical monitoring, process control) expands our freedom for flexible ad-hoc communication and dynamic collaboration between individuals through a wide variety of devices. With the pervasive deployment of computers, the Peer-to-Peer (P2P) model is increasingly receiving attention for direct and symmetric interaction and to perform critical functions in a decentralized manner. Some of the benefits of a P2P environment include its ability for self-organization, scalability by avoiding dependency on centralized servers, resource aggregation and adaptation to different loads, resiliency to node failures and lower cost of ownership and cost sharing. The goal of this panel is to explore the synergy of these two technologies, identify the challenges along with their shortcomings and explore promising approaches. Vana Kalogeraki |
Mobile Data Management | 1 |
| 2005 | Data dissemination in mobile peer-to-peer networksabstractIn this paper we propose adaptive content-driven routing and data dissemination algorithms for intelligently routing search queries in a peer-to-peer network that supports mobile users. In our mechanism nodes build content synopses of their data and adaptively disseminate them to the most appropriate nodes. Based on the content synopses, a routing mechanism is being built to forward the queries to those nodes that have a high probability of providing the desired results. Our simulation results show that our approach is highly scalable and significantly improves resources usage by saving both bandwidth and processing power. Thomas Repantis, Vana Kalogeraki |
Mobile Data Management | 2 |
| 2005 | Exploiting locality for scalable information retrieval in peer-to-peer networks
Demetris Zeinalipour, Vana Kalogeraki, Dimitrios Gunopulos |
Inf. Syst. | 2 |
| 2002 | A local search mechanism for peer-to-peer networksabstractOne important problem in peer-to-peer (P2P) networks is searching and retrieving the correct information. However, existing searching mechanisms in pure peer-to-peer networks are inefficient due to the decentralized nature of such networks. We propose two mechanisms for information retrieval in pure peer-to-peer networks. The first, the modified Breadth-First Search (BFS) mechanism, is an extension of the current Gnuttela protocol, allows searching with keywords, and is designed to minimize the number of messages that are needed to search the network. The second, the Intelligent Search mechanism, uses the past behavior of the P2P network to further improve the scalability of the search procedure. In this algorithm, each peer autonomously decides which of its peers are most likely to answer a given query. The algorithm is entirely distributed, and therefore scales well with the size of the network. We implemented our mechanisms as middleware platforms. To show the advantages of our mechanisms we present experimental results using the middleware implementation. Vana Kalogeraki, Dimitrios Gunopulos, Demetris Zeinalipour |
CIKM | 1 |