George Pallis 0001

dblp:21/4460 · DBLP profile ↗
← Back
51ranked-venue papers
5as first author
12since 2021 · last 2025
0000-0003-1815-5468ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 20 · 2 first-author · 5 since 2021Systems, architecture and hardware · 16 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 10 · 2 first-author · 2 since 2021Computer networks · 7 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 since 2021Human-computer interaction and ubiquitous computing · 5 · 4 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021
YearPublicationVenuePosition
2025 Adopting Beliefs or Superficial Mimicry? Investigating Nuanced Ideological Manipulation of LLMs
abstract
Large Language Models (LLMs) have transformed natural language processing, but concerns have emerged about their susceptibility to ideological manipulation, particularly in politically sensitive areas. Previous research has largely focused on LLM biases through a binary Left vs. Right framework, often using explicit ideological prompts and fine-tuning with political question-answering datasets. In this work, we move beyond this binary approach to explore the extent to which LLMs can be influenced across a nuanced spectrum of political ideologies, from Progressive-Left to Conservative-Right. We introduce a novel multi-task dataset designed to reflect diverse ideological positions through tasks such as ideological question-answering, statement ranking, manifesto cloze completion, and Congress bill comprehension. By fine-tuning three LLMs—Phi-2, Mistral, and Llama-3—on this dataset, we evaluate their capacity to adopt and express these nuanced ideologies. Our findings indicate that fine-tuning significantly enhances nuanced ideological alignment, while explicit prompts provide only minor refinements. This highlights the models' susceptibility to subtle ideological manipulation, suggesting a need for more robust safeguards to mitigate these risks.
Demetris Paschalides, George Pallis 0001, Marios D. Dikaiakos
ICWSM2
2025 GNN and LLM Insights: Multimodal Cues and Gender Disparities in Video Conversations
abstract
As video content on online platforms continues to increase, understanding the complex aspects of interpersonal communication becomes crucial. Central to this exploration is the pressing issue of gender bias, which manifests in multimodal interactions through visual, vocal, or verbal cues. These interactions present challenges in extracting and interpreting the subtle cues that may point to underlying biases. To tackle these challenges, we introduce a semi-automatic extraction of features and knowledge from user-generated content on video web platforms. Using 1,091 unstructured multi-participant video conversations from Shark Tank, we examine whether the multimodal cues (e.g., emotions) of a conversational participant (e.g., entrepreneur) affect another participant (e.g., investor) differently due to gender biases. Our methodology employs advanced deep learning algorithms for cues extraction and leverages Graph Neural Networks to model the multi-participant conversations. To complement our findings, we utilize textual features extracted through our methodology and employ GPT-4 to simulate decision-making scenarios, thereby assessing its analytical capabilities and potential gender biases.
Dimosthenis Stefanidis, George Pallis 0001, Marios D. Dikaiakos, Nicos Nicolaou
ICWSM2
2024 PARALLAX: Leveraging Polarization Knowledge for Misinformation Detection
Demetris Paschalides, George Pallis 0001, Marios D. Dikaiakos
ASONAM (1)2
2024 Energy modeling of inference workloads with AI accelerators at the Edge: A benchmarking study
abstract
Analyzing and modeling the performance and energy consumption of hybrid Edge Computing systems with embedded devices and Artificial Intelligence (AI) accelerators is crucial, yet challenging due to the lack of systematic methods and tools for measuring and estimating energy consumption. We address this gap by introducing a systematic methodology and a toolset to benchmark AI accelerators and their host devices with inference workloads representing mock Convolutional Neural Network (CNN) models with varying input sizes, network sizes, layer types, and kernel sizes. The primary contributions of this work include the development of the benchmarking methodology, the creation and analysis of a comprehensive dataset comprising power benchmark results, and the development of a predictive model for estimating the energy consumption of ML workloads on the Coral TPU (Tensor Processing Unit) accelerator connected to the edge device. The dataset, generated from extensive testing on the deployed topology, is released and can be used for further studies that seek to enhance the energy efficiency and performance optimization for Edge Computing applications.
Michalis Kasioulis, Moysis Symeonides, Giorgos Ioannou, George Pallis 0001, Marios D. Dikaiakos
IC2E4
2024 Evaluating the utility of human mobility data under local differential privacy
abstract
In this paper, we evaluate the impact of local differential privacy (LDP) on the utility of human mobility data obtained from mobile location services. Specifically, we focus our study on visit data, which consist of user-level information on visited locations. This includes the duration of each visit and its category, such as restaurant or department store. The purpose of LDP is to protect sensitive information in visit records by introducing properly calibrated noise, while still allowing the extraction of useful statistics. To evaluate our approach, we study how different levels of privacy budget ϵ impact the utility of the data. The utility is determined by the estimation accuracy for different statistics of interest, such as the number of visits and the average visit duration for each category. We conduct our evaluation on a visits dataset including records from over 20 million mobile devices. Our findings indicate that the number of visits to popular categories can be accurately estimated even at strong privacy levels (ϵ = 1). The estimation of the average visit duration is generally less precise, but it remains feasible under less stringent privacy levels (ϵ ≥ 2 or ϵ ≥ 4, depending on the application at hand).
Giorgos Ioannou, Thomas Marchioro, Christos Nicolaides, George Pallis 0001, Evangelos P. Markatos
MDM4
2023 Energy-Aware Streaming Analytics Job Scheduling for Edge Computing
abstract
Energy profiling and optimization are expected to be crucial factors impacting the realisation of the Internet of Things (IoT) as more intelligence is deployed at the network extremes to achieve better response times in the proximity of where data are harvested. To improve the performance of streaming analytics jobs, several schedulers have been designed to tackle key challenges in edge computing realms, including resource heterogeneity and highly volatile network links. However, energy-aware scheduling for streaming analytic jobs is at best, not adequately examined. In this article, we introduce PowerStorm, a scheduler for streaming analytic jobs that is designed to explore trade-offs between performance and energy consumption in geodistributed edge computing settings. We implement our scheduler for Apache Storm and show the scheduler’s energy saving capabilities over the Yahoo streaming benchmark with worker nodes featuring heterogeneous power and resource capabilities on both a physical and emulated testbed.
Demetris Trihinas, Moysis Symeonides, Joanna Georgiou, George Pallis 0001, Marios D. Dikaiakos
CloudCom4
2023 SparkEdgeEmu: An Emulation Framework for Edge-Enabled Apache Spark Deployments
Moysis Symeonides, Demetris Trihinas, George Pallis 0001, Marios D. Dikaiakos
Euro-Par3
2023 A cyber-physical management system for medium-scale solar-powered data centers
abstract
Summary The effort to reduce the environmental impact and carbon footprint of data‐center operations has led to the emergence of “green” data centers, which are designed to reduce energy consumption and to increase their use of Renewable Energy Sources (RES). Despite the advances demonstrated by hyper‐scale facilities in energy efficiency and the use of green energy, small and medium‐scale data centers, which contribute to over 50% of the total electricity consumption and carbon emissions of the sector, face significant challenges in the adoption and exploitation of RES. In this article, we present the steps taken to transform a medium‐scale, academic data center into a “green” one that uses solar power. In particular, we describe the design and implementation of: (i) a collocated photovoltaic facility and (ii) a cyber‐physical system comprising IoT sensor devices, a microservices platform, and a visualization and analytics dashboard that supports the configuration and monitoring of the infrastructure. Using data collected from the platform and dashboard, we show the environmental and financial advantages derived from this transformation, and the potential that arises from the availability of integrated operational data.
Marios D. Dikaiakos, Nikolas G. Chatzigeorgiou, Athanasios Tryfonos, Andreas Andreou, Nicholas Loulloudes, George Pallis 0001, George E. Georghiou
Concurr. Comput. Pract. Exp.6
2022 BenchPilot: Repeatable & Reproducible Benchmarking for Edge Micro-DCs
abstract
Micro-Datacenters (DCs) are emerging as key en-ablers for Edge computing and 5G mobile networks by pro-viding processing power closer to IoT devices to extract timely analytic insights. However, the performance evaluation of data stream processing on micro-DCs is a daunting task due to difficulties raised by the time-consuming setup, configuration and heterogeneity of the underlying environment. To address these challenges, we introduce BenchPilot, a modular and highly customizable benchmarking framework for edge micro-DCs. BenchPilot provides a high-level declarative model for describing experiment testbeds and scenarios that automates the bench-marking process on Streaming Distributed Processing Engines (SDPEs). The latter enables users to focus on performance analysis instead of dealing with the complex and time-consuming setup. BenchPilot instantiates the underlying cluster, performs repeatable experimentation, and provides a unified monitoring stack in heterogeneous Micro-DCs. To highlight the usability of BenchPilot, we conduct experiments on two popular streaming engines, namely Apache Storm and Flink. Our experiments compare the engines based on performance, CPU utilization, energy consumption, temperature, and network I/O.
Joanna Georgiou, Moysis Symeonides, Michalis Kasioulis, Demetris Trihinas, George Pallis 0001, Marios D. Dikaiakos
ISCC5
2022 Demo: The RAINBOW Analytics Stack for the Fog Continuum
abstract
With the proliferation of raw Internet of Things (IoTs) data, Fog Computing is emerging as a computing paradigm for delay-sensitive streaming analytics with operators deploying big data distributed engines on Fog resources [1]. Nevertheless, the current (Cloud-based) distributed analytics solutions are unaware of the unique characteristics of Fog realms. For instance, task placement algorithms consider homogeneous underlying resources without considering the Fog nodes' heterogeneity and the non-uniform network connections, resulting in sub-optimal processing performance. Moreover, data quality can play an important role, where corrupted data, and network uncertainty may lead to less useful results. In turn, energy consumption can critically impact the overall cost and liveness of the underlying processing infrastructure. Specifically, scheduling tasks on nodes with energy-hungry profiles or battery-powered devices may temporarily be beneficial for the performance, but it may increase the overall cost, or/and the battery-powered devices may not be available when needed. A Fog-enabled analytics stack must allow users to optimize Fog-specific indicators or trade-offs among them. For instance, users may sacrifice a portion of the execution performance to minimize energy consumption or vice versa. Except for the performance issues raised by Fog, the state-of-the-art distributed processing engines offer only low-level procedural programming interfaces with operators facing a steep learning curve to master them. So, query abstractions are crucial for minimizing the deployment time, errors, and debugging.
Moysis Symeonides, Demetris Trihinas, Joanna Georgiou, Michalis Kasioulis, George Pallis 0001, Marios D. Dikaiakos, Theodoros Toliopoulos, Anna-Valentini Michailidou, Anastasios Gounaris
ISCC5
2021 POLAR: a holistic framework for the modelling of polarization and identification of polarizing topics in news media
abstract
Polarization is an alarming trend in modern societies with serious implications on social cohesion and the democratic process. Typically, polarization manifests itself in the public discourse in politics, governance and ideology. In recent years, however, polarization arises increasingly in a wider range of issues, from identity and culture to healthcare and the environment. As the public and private discourse moves online, polarization feeds in and is fed by phenomena like fake news and hate speech. The identification and analysis of online polarization is challenging because of the massive scale, diversity, and unstructured nature of online content, and the rapid and unpredictable evolution of polarizing issues. Therefore, we need effective ways to identify, quantify, and represent polarization and polarizing topics algorithmically and at scale. In this work, we introduce POLAR - an unsupervised, large-scale framework for modeling and identifying polarizing topics in any domain, without prior domain-specific knowledge. POLAR comprises a processing pipeline that analyzes a corpus of an arbitrary number of news articles to construct a hierarchical knowledge graph that models polarization and identify polarizing topics discussed in the corpus. Our evaluation shows that POLAR is able to identify and rank polarizing topics accurately and efficiently.
Demetris Paschalides, George Pallis 0001, Marios D. Dikaiakos
ASONAM2
2021 Low-Cost Adaptive Monitoring Techniques for the Internet of Things
abstract
Internet-enabled physical devices with “smart” processing capabilities are becoming the tools for understanding the complexity of the global inter-connected world we inhabit. The Internet of Things (IoT) churns tremendous amounts of data flooding from devices scattered across multiple locations to the processing engines of almost all industry sectors. However, as the number of “things” surpasses the population of the technology-enabled world, real-time processing and energy-efficiency are great challenges of the big data era transitioning to IoT. In this article, we introduce a lightweight adaptive monitoring framework suitable for smart IoT devices with limited processing capabilities. Our framework, inexpensively and in place dynamically adjusts the monitoring intensity and the amount of data disseminated through the network based on a low-cost adaptive and probabilistic learning model capable of capturing at runtime the current evolution and variability of the data stream. By accomplishing this, energy consumption and data volume are reduced, allowing IoT devices to preserve battery and ease processing on cloud computing and streaming services. Experiments on real-world data from cloud services, internet security services, wearables and intelligent transportation services, show that our framework achieves a balance between efficiency and accuracy. Specifically, our framework reduces data volume by 74 percent, energy consumption by at least 71 percent, while maintaining accuracy always above 89 percent.
Demetris Trihinas, George Pallis 0001, Marios D. Dikaiakos
IEEE Trans. Serv. Comput.2
2020 Fogify: A Fog Computing Emulation Framework
abstract
Fog Computing is emerging as the dominating paradigm bridging the compute and connectivity gap between sensing devices and latency-sensitive services. However, experimenting and evaluating IoT services is a daunting task involving the manual configuration and deployment of a mixture of geodistributed physical and virtual infrastructure with different resource and network requirements. This results in sub-optimal, costly and error-prone deployments due to numerous unexpected overheads not initially envisioned in the design phase and underwhelming testing conditions not resembling the end environment. In this paper, we introduce Fogify, an emulator easing the modeling, deployment and large-scale experimentation of fog and edge testbeds. Fogify provides a toolset to: (i) model complex fog topologies comprised of heterogeneous resources, network capabilities and QoS criteria; (ii) deploy the modelled configuration and services using popular containerized descriptions to a cloud or local environment; (iii) experiment, measure and evaluate the deployment by injecting faults and adapting the configuration at runtime to test different “what-if” scenarios that reveal the limitations of a service before introduced to the public. In the evaluation, proof-of-concept IoT services with real-world workloads are introduced to show the wide applicability and benefits of rapid prototyping via Fogify.
Moysis Symeonides, Zacharias Georgiou, Demetris Trihinas, George Pallis 0001, Marios D. Dikaiakos
SEC4
2020 Demo: Emulating Geo-Distributed Fog Services
abstract
For more than the better parts of the last decades, we are witnessing the proliferation of IoT devices, as well as an exponential growth in the volume of data generated outside of datacenters. With the generated data at the extremes of the network and the restricted device-to-cloud bandwidth, data mitigation is becoming the major barrier of cloud-based IoT services [1]. To alleviate these challenges, Fog Computing extends the Cloud's capabilities closer to IoT devices.
Moysis Symeonides, Zacharias Georgiou, Demetris Trihinas, George Pallis 0001, Marios D. Dikaiakos
SEC4
2020 MANDOLA: A Big-Data Processing and Visualization Platform for Monitoring and Detecting Online Hate Speech
abstract
In recent years, the increasing propagation of hate speech in online social networks and the need for effective counter-measures have drawn significant investment from social network companies and researchers. This has resulted in the development of many web platforms and mobile applications for reporting and monitoring online hate speech incidents. In this article, we present MANDOLA, a big-data processing system that monitors, detects, visualizes, and reports the spread and penetration of online hate-related speech using big-data approaches. MANDOLA consists of six individual components that intercommunicate to consume, process, store, and visualize statistical information regarding hate speech spread online. We also present a novel ensemble-based classification algorithm for hate speech detection that can significantly improve the performance of MANDOLA’s ability to detect hate speech. To present the functionality and usability of our system, we present a use case scenario of real-life event annotation and data correlation. As shown from the performance of the individual modules, as well as the usability and functionality of the whole system, MANDOLA is a powerful system for reporting and monitoring online hate speech.
Demetris Paschalides, Dimosthenis Stefanidis, Andreas Andreou, Kalia Orphanou, George Pallis 0001, Marios D. Dikaiakos, Evangelos P. Markatos
ACM Trans. Internet Techn.5
2019 Query-Driven Descriptive Analytics for IoT and Edge Computing
abstract
With consumers embracing the prevalence of ubiquitously connected smart devices, Edge Computing is emerging as a principal computing paradigm for latency-sensitive and in-proximity services. However, as the plethora of data generated across connected devices continues to vastly increase, the need to query the "edge" and derive in-time analytic insights is more evident than ever. This paper introduces our vision for a rich and declarative query model abstraction particularly tailored for the unique characteristics of Edge Computing and presents a prototype framework that realizes our vision. Towards this, the declarative query model enables users to express high-level and descriptive analytic insights, while our framework compiles, optimizes and executes the query plan decoupled from the programming model of the underlying data processing engine. Afterwards, we showcase a number of potential use-cases which stand to benefit from the realization of query-driven descriptive analytics for edge computing. We conclude by elaborating on the open challenges that still must be addressed to realize our vision and potential research opportunities for the academic community to further advance the current State-of-the-Art.
Moysis Symeonides, Demetris Trihinas, Zacharias Georgiou, George Pallis 0001, Marios D. Dikaiakos
IC2E4
2019 Check-It: A plugin for Detecting and Reducing the Spread of Fake News and Misinformation on the Web
abstract
Over the past few years, we have been witnessing the rise of misinformation on the Internet. People fall victims of fake news continuously, and contribute to their propagation knowingly or inadvertently. Many recent efforts seek to reduce the damage caused by fake news by identifying them automatically with artificial intelligence techniques, using signals from domain flag-lists, online social networks, etc. In this work, we present Check-It, a system that combines a variety of signals into a pipeline for fake news identification. Check-It is developed as a web browser plugin with the objective of efficient and timely fake news detection, while respecting user privacy. In this paper, we present the design, implementation and performance evaluation of Check-It. Experimental results show that it outperforms state-of-the-art methods on commonly-used datasets.
Demetris Paschalides, Alexandros Kornilakis, Chrysovalantis Christodoulou, Rafael Andreou, George Pallis 0001, Marios D. Dikaiakos, Evangelos P. Markatos
WI5
2019 Two-hop privacy-preserving nearest friend searches
Alexandros Karakasidis 0001, George Pallis 0001, Marios D. Dikaiakos
Knowl. Inf. Syst.2
2019 Edge Computing [Scanning the Issue]
abstract
In recent years, with the proliferation of the Internet of Things (IoT) and the wide penetration of wireless networks, the number of edge devices and the data generated from the edge have been growing rapidly. According to International Data Corporation (IDC) prediction[20], global data will reach 180 zettabytes (ZB), and 70% of the data generated by IoT will be processed on the edge of the network by 2025. IDC also forecasts that more than 150 billion devices will be connected worldwide by 2025. In this case, the centralized processing mode based on cloud computing is not efficient enough to handle the data generated by the edge. The centralized processing model uploads all data to the cloud data center through the network and leverages its supercomputing power to solve the computing and storage problems, which enables the cloud services to create economic benefits. However, in the context of IoT, traditional cloud computing has several shortcomings.
Weisong Shi, George Pallis 0001
Proc. IEEE2
2018 ATMoN: Adapting the "Temporality" in Large-Scale Dynamic Networks
abstract
With the widespread adoption of temporal graphs to study fast evolving interactions in dynamic networks, attention is needed to provide graph metrics in time and at scale. In this paper, we introduce ATMoN, an open-source library developed to computationally offload graph processing engines and ease the communication overhead in dynamic networks over an unprecedented wealth of data. This is achieved, by efficiently adapting, in place and inexpensively, the temporal granularity at which graph metrics are computed based on runtime knowledge captured by a low-cost probabilistic learning model capable of approximating both the metric stream evolution and the volatility of the graph topology. After a thorough evaluation with real-world data from mobile, face-to-face and vehicular networks, results show that ATMoN is able to reduce the compute overhead by at least 76%, data volume by 60% and overall cloud costs by at least 54%, while always maintaining accuracy above 88%.
Demetris Trihinas, Luis F. Chiroque, George Pallis 0001, Antonio Fernández 0001, Marios D. Dikaiakos
ICDCS3
2018 SELECT: A Distributed Publish/Subscribe Notification System for Online Social Networks
abstract
Publish/subscribe (pub/sub) mechanisms constitute an attractive communication paradigm in the design of large-scale notification systems for Online Social Networks (OSNs). To accommodate the large-scale workloads of notifications produced by OSNs, pub/sub mechanisms require thousands of servers distributed on different data centers all over the world, incurring large overheads. To eliminate the pub/sub resources used, we propose SELECT - a distributed pub/sub social notification system over peer-to-peer (P2P) networks. SELECT organizes the peers on a ring topology and provides an adaptive P2P connection establishment algorithm where each peer identifies the number of connections required, based on the social structure and user availability. This allows to propagate messages to the social friends of the users using a reduced number of hops. The presented algorithm is an efficient heuristic to an NP-hard problem which maps workload graphs to structured P2P overlays inducing overall, close to theoretical, minimal number of messages. Experiments show that SELECT reduces the number of relay nodes up to 89% versus the state-of-the-art pub/sub notification systems. Additionally, we demonstrate the advantage of SELECT against socially-aware P2P overlay networks and show that the communication between two socially connected peers is reduced on average by at least 64% hops, while achieving 100% communication availability even under high churn.
Nuno Apolónia, Stefanos Antaris, Sarunas Girdzijauskas, George Pallis 0001, Marios D. Dikaiakos
IPDPS4
2018 Monitoring Elastically Adaptive Multi-Cloud Services
abstract
Automatic resource provisioning is a challenging and complex task. It requires for applications, services and underlying platforms to be continuously monitored at multiple levels and time intervals. The complex nature of this task lays in the ability of the monitoring system to automatically detect runtime configurations in a cloud service due to elasticity action enforcement. Moreover, with the adoption of open cloud standards and library stacks, cloud consumers are now able to migrate their applications or even distribute them across multiple cloud domains. However, current cloud monitoring tools are either bounded to specific cloud platforms or limit their portability to provide elasticity support. In this article, we describe the challenges when monitoring elastically adaptive multi-cloud services. We then introduce a novel automated, modular, multi-layer and portable cloud monitoring framework. Experiments on multiple clouds and real-life applications show that our framework is capable of automatically adapting when elasticity actions are enforced to either the cloud service or to the monitoring topology. Furthermore, it is recoverable from faults introduced in the monitoring configuration with proven scalability and low runtime footprint. Most importantly, our framework is able to reduce network traffic by 41 percent and consequently the monitoring cost, which is both billable and noticeable in large-scale multi-cloud services.
Demetris Trihinas, George Pallis 0001, Marios D. Dikaiakos
IEEE Trans. Cloud Comput.2
2017 ADMin: Adaptive monitoring dissemination for the Internet of Things
abstract
As more knowledge is vastly added to the devices fuelling the Internet of Things (IoT) energy efficiency and real-time data processing are great challenges that must be tackled. In this paper, we introduce ADMin, a low-cost IoT framework that reduces on device energy consumption and the volume of data disseminated across the network. This is achieved by efficiently adapting the rate at which IoT devices disseminate monitoring streams based on run-time knowledge of the stream evolution, variability and seasonal behavior. Rather than transmitting the entire stream, ADMin favors sending updates for its estimation model from which values can be inferred, triggering dissemination only when shifts in the stream evolution are detected. Results on real-life testbeds, show that ADMin is able to reduce energy consumption by at least 83%, data volume by 71%, shift detection delays by 61% while maintaining accuracy above 91% in comparison to other IoT frameworks.
Demetris Trihinas, George Pallis 0001, Marios D. Dikaiakos
INFOCOM2
2017 A cost-effective approach to improving performance of big genomic data analyses in clouds
Christopher Smowton, Andoena Balla, Demetris Antoniades, Crispin J. Miller, George Pallis 0001, Marios D. Dikaiakos
Future Gener. Comput. Syst.5
2016 Online social network evolution: Revisiting the Twitter graph
abstract
In 2010 the popular paper by Kwak et al. [11] presented the first comprehensive study of Twitter as it appeared in 2009, using most of the Twitter network at the time. Since then, Twitter's popularity and usage has exploded, experiencing a 10-fold increase. As of 2015, it has more than 500 million users, out of which 316 million are active, i.e. logging into the service at least once a month.1In this study we revisit the network observed by Kwak et al. to examine the changes exhibited in both the graph and the behavior of the users in it. Our results conclude to a denser network, showing an increase in the number of reciprocal edges, despite the fact that around 12.5% of the 2009 users have now left Twitter. However, the network's largest strongly connected component seems to be significantly decreasing, suggesting a movement of edges towards popular users. Furthermore, we observe numerous changes in the lists of influential Twitter users, having several accounts that where not popular in the past securing a position in the top-20 list as new entries.
Hariton Efstathiades, Demetris Antoniades, George Pallis 0001, Marios D. Dikaiakos, Zoltán Szlávik, Robert-Jan Sips
IEEE BigData3
2015 Identification of Key Locations based on Online Social Network Activity
abstract
Ubiquitous Internet connectivity enables users to update their Online Social Network profile from any location and at any point in time. These, often geo-tagged, data can be used to provide valuable information to closely located users, both in real time and in aggregated form. However, despite the fact that users publish geo-tagged information, only a small number implicitly reports their base location in their Online Social Network profile. In this paper we present a simple yet effective methodology for identifying a user's key locations, namely her home and work places. We evaluate our methodology with Twitter datasets collected from the country of Netherlands, city of London and Los Angeles county. Furthermore, we combine Twitter and LinkedIn information to construct a work location dataset and evaluate our methodology. Results show that our proposed methodology not only outperforms state-of-the-art methods by at least 30% in terms of accuracy, but also cuts the detection radius at least at half the distance from other methods.
Hariton Efstathiades, Demetris Antoniades, George Pallis 0001, Marios D. Dikaiakos
ASONAM3
2015 AdaM: An adaptive monitoring framework for sampling and filtering on IoT devices
abstract
Real-time data processing while the velocity and volume of data generated keep increasing, as well as, energy-efficiency are great challenges of big data streaming which have transitioned to the Internet of Things (IoT) realm. In this paper, we introduce AdaM, a lightweight adaptive monitoring framework for smart battery-powered IoT devices with limited processing capabilities. AdaM, inexpensively and in place dynamically adapts the monitoring intensity and the amount of data disseminated through the network based on the current evolution and variability of the metric stream. Results on real-world testbeds, show that AdaM achieves a balance between efficiency and accuracy. Specifically, AdaM is capable of reducing data volume by 74%, energy consumption by at least 71%, while preserving a greater than 89% accuracy.
Demetris Trihinas, George Pallis 0001, Marios D. Dikaiakos
IEEE BigData2
2015 Analysing Cancer Genomics in the Elastic Cloud
abstract
With the rapidly growing demand for DNA analysis, the need for storing and processing large-scale genome data has presented significant challenges. This paper describes how the Genome Analysis Toolkit (GATK) can be deployed to an elastic cloud, and defines policy to drive elastic scaling of the application. We extensively analyse the GATK to expose opportunities for resource elasticity, demonstrate that it can be practically deployed at scale in a cloud environment, and demonstrate that applying elastic scaling improves the performance to cost tradeoff achieved in a simulated environment.
Christopher Smowton, Crispin J. Miller, Andoena Balla, Demetris Antoniades, George Pallis 0001, Marios D. Dikaiakos
CCGRID6
2015 Clustering Attributed Multi-graphs with Information Ranking
Andreas Papadopoulos, Dimitrios Rafailidis, George Pallis 0001, Marios D. Dikaiakos
DEXA (1)3
2015 Evaluating Cloud Service Elasticity Behavior
abstract
To optimize the cost and performance of complex cloud services under dynamic requirements, workflows and diverse cloud offerings, we rely on different elasticity control processes. An elasticity control process, when being enforced, produces effects in different parts of the cloud service. These effects normally evolve in time and depend on workload characteristics, and on the actions within the elasticity control process enforced. Therefore, understanding the effects on the behavior of the cloud service is of utter importance for runtime decision-making process, when controlling cloud service elasticity. In this paper, we present a novel methodology and a framework for estimating and evaluating cloud service elasticity behaviors. To estimate the elasticity behavior, we collect information concerning service structure, deployment, service runtime, control processes, and cloud infrastructure. Based on this information, we utilize clustering techniques to identify cloud service elasticity behavior, in time, and for different parts of the service. Knowledge about such behavior is utilized within a cloud service elasticity controller to substantially improve the selection and execution of elasticity control processes. These elasticity behavior estimations are successfully being used by our elasticity controller, in order to improve runtime decision quality. We evaluate our framework with three real-world cloud services in different application domains. Experiments show that we are able to estimate the behavior in 89.5% of the cases. Moreover, we have observed improvements in our elasticity controller, which takes better control decisions, and does not exhibit control oscillations.
Georgiana Copil, Hong Linh Truong 0001, Daniel Moldovan, Schahram Dustdar, Demetris Trihinas, George Pallis 0001, Marios D. Dikaiakos
Int. J. Cooperative Inf. Syst.6
2014 JCatascopia: Monitoring Elastically Adaptive Applications in the Cloud
abstract
Over the past decade, Cloud Computing has rapidly become a widely accepted paradigm with core concepts such as elasticity, scalability and on demand automatic resource provisioning emerging as next generation Cloud service-must have-properties. Automatic resource provisioning for Cloud applications is not a trivial task, requiring for both the applications and platform, to be constantly monitored, capturing information at various levels and time granularity. In this paper we describe the challenges that occur when monitoring elastically adaptive Cloud applications and to address these issues we present JCatascopia, a fully automated, multi-layer, interoperable Cloud Monitoring System. Experiments on different production Cloud platforms show that JCatascopia is a Monitoring System capable of supporting a fully automated Cloud resource provisioning system with proven interoperability, scalability and low runtime footprint. Most importantly, JCatascopia is able to adapt in a fully automatic manner when elasticity actions are enforced to an application deployment.
Demetris Trihinas, George Pallis 0001, Marios D. Dikaiakos
CCGRID2
2014 c-Eclipse: An Open-Source Management Framework for Cloud Applications
Chrystalla Sofokleous, Nicholas Loulloudes, Demetris Trihinas, George Pallis 0001, Marios D. Dikaiakos
Euro-Par4
2014 ADVISE - A Framework for Evaluating Cloud Service Elasticity Behavior
Georgiana Copil, Demetris Trihinas, Hong Linh Truong 0001, Daniel Moldovan, George Pallis 0001, Schahram Dustdar, Marios D. Dikaiakos
ICSOC5
2013 Identifying Clusters with Attribute Homogeneity and Similar Connectivity in Information Networks
abstract
With the rapid emergence of the internet world, a lot of information networks become available every day. In many cases, these information networks contain objects connected by multiple links and described by different attributes. In this paper the problem of clustering homogeneous information networks in groups with similar attributes and connections is studied. Clustering such networks is a challenging task due to different importance of links and attributes. In addition, it is not straightforward how to balance the links and attributes information. In this article we describe these challenges and propose a fuzzy clustering model as well as a fuzzy clustering algorithm, HASCOP. Extensive experimentation on real world datasets has shown that HASCOP can be successfully applied in such networks, demonstrating its efficacy and superiority against the state-of-the-art attributed graph clustering methods.
Andreas Papadopoulos, George Pallis 0001, Marios D. Dikaiakos
Web Intelligence2
2012 Automated Tagging for the Retrieval of Software Resources in Grid and Cloud Infrastructures
abstract
A key challenge for Grid and Cloud infrastructures is to make their services easily accessible and attractive to end-users. In this paper we introduce tagging capabilities to the Miner soft system, a powerful tool for software search and discovery in order to help end-users locate application software suitable to their needs. Miner soft is now able to predict and automatically assign tags to software resources it indexes. In order to achieve this, we model the problem of tag prediction as a multi-label classification problem. Using data extracted from production-quality Grid and Cloud computing infrastructures, we evaluate an important number of multi-label classifiers and discuss which one and with what settings is the most appropriate for use in the particular problem.
Ioannis Katakis 0001, George Pallis 0001, Marios D. Dikaiakos, Onisiforos Onoufriou
CCGRID2
2012 The Role of Twitter in YouTube Videos Diffusion
George Christodoulou 0002, Chryssis Georgiou, George Pallis 0001
WISE3
2012 Minersoft: Software retrieval in grid and cloud computing infrastructures
abstract
One of the main goals of Cloud and Grid infrastructures is to make their services easily accessible and attractive to end-users. In this article we investigate the problem of supporting keyword-based searching for the discovery of software files that are installed on the nodes of large-scale, federated Grid and Cloud computing infrastructures. We address a number of challenges that arise from the unstructured nature of software and the unavailability of software-related metadata on large-scale networked environments. We present Minersoft, a harvester that visits Grid/Cloud infrastructures, crawls their file systems, identifies and classifies software files, and discovers implicit associations between them. The results of Minersoft harvesting are encoded in a weighted, typed graph, called the Software Graph. A number of information retrieval (IR) algorithms are used to enrich this graph with structural and content associations, to annotate software files with keywords and build inverted indexes to support keyword-based searching for software. Using a real testbed, we present an evaluation study of our approach, using data extracted from production-quality Grid and Cloud computing infrastructures. Experimental results show that Minersoft is a powerful tool for software search and discovery.
Marios D. Dikaiakos, Asterios Katsifodimos, George Pallis 0001
ACM Trans. Internet Techn.3
2011 Real-time graph visualization tool for vehicular ad-hoc networks: (VIVAGr: VIsualization tool of VAnet graphs in real-time)
abstract
In this work we describe VIVAGr, a graphical-oriented real time visualization tool for vehicular ad-hoc network connectivity graphs. This tool enables the effective synthesis of structural, topological, and dynamic characteristics of VANET graphs, with a variety of parameters that affect the shape and characteristics of a vehicular ad hoc network (wireless range, mobility models, road-network topology, market penetration ratio, and exhibited interference). Our design allows researchers to explore and understand problems and issues related with vehicular ad-hoc networks that face today significant design challenges. The tool is able to present all active connection of the network in real-time mode using mobility traces. A visual encoding syntax is used to represent semantic meanings and highlight the effect of mobility and topology on vehicular network specific properties.
Emmanouil Spanakis, Christodoulos Efstathiades, George Pallis 0001, Marios D. Dikaiakos
ISCC3
2011 Editorial for special issue Internet-based Content Delivery
Giancarlo Fortino, Carlo Mastroianni, George Pallis 0001, Mukaddim Pathan, Athena Vakali
Comput. Networks3
2010 Caching Dynamic Information in Vehicular Ad Hoc Networks
Nicholas Loulloudes, George Pallis 0001, Marios D. Dikaiakos
Euro-Par (2)2
2010 On the Evaluation of Caching in Vehicular Information Systems
abstract
VANETs have been envisioned as an infrastructure for deploying Vehicular Information Systems (VIS) that among others provide drivers with an up-to-date view on the prevailing traffic conditions. In this work we evaluate the benefits of caching vehicular information obtained from such VIS through VITP, a location-aware, application-layer communication protocol that we extend to support caching. We present an evaluation study of our approach conducting extensive simulation on large scale vehicular networks under different realistic urban traffic conditions Our results identify the critical parameters that affect information quality in VANETs as well as demonstrate the viability and effectiveness of the cache-enabled VITP.
Nicholas Loulloudes, George Pallis 0001, Marios D. Dikaiakos
Mobile Data Management2
2010 Searching for Software on the EGEE Infrastructure
George Pallis 0001, Asterios Katsifodimos, Marios D. Dikaiakos
J. Grid Comput.1
2009 Harvesting Large-Scale Grids for Software Resources
abstract
Grid infrastructures are in operation around the world, federating an impressive collection of computational resources and a wide variety of application software. In this context, it is important to establish advanced software discovery services that could help end-users locate software components suitable to their needs. In this paper, we present the design, architecture and implementation of an open-source keyword-based paradigm for the search of software resources in Grid infrastructures, called Minersoft. A key goal of Minersoft is to annotate automatically all the software resources with keyword-rich metadata. Using advanced Information Retrieval techniques, we locate software resources with respect to users queries. Experiments were conducted in EGEE, one of the largest Grid production services currently in operation. Results showed that Minersoft successfully crawled 12.3 million valid files (620 GB size) and sustained, in most sites, high crawling rates.
Asterios Katsifodimos, George Pallis 0001, Marios D. Dikaiakos
CCGRID2
2009 On the structure and evolution of vehicular networks
abstract
Vehicular ad hoc networks have emerged recently as a platform to support intelligent inter-vehicle communication and improve traffic safety and performance. The road-constrained and high mobility of the vehicles, their unbounded power source, and the emergence of roadside wireless infrastructures make VANETs a challenging research topic. A key to the development of protocols for intervehicle communication and services lies in the knowledge of the topological characteristics of the VANET communication graph. This article provides answers to the general question: how does a VANET communication graph look like over time and space? This study is the first one that examines a very large-scale VANET graph and conducts a thorough investigation of its topological characteristics using several metrics, not examined in previous studies. Our work characterizes a VANET graph at the connectivity (link) level, quantifies the notion of ¿qualitative¿ nodes as required by routing and dissemination protocols, and examines the existence and evolution of communities (dense clusters of vehicles) in the VANET. Several latent facts about the VANET graph are revealed and incentives for their exploitation in protocol design are examined.
George Pallis 0001, Dimitrios Katsaros 0001, Marios D. Dikaiakos, Nicholas Loulloudes, Leandros Tassiulas
MASCOTS1
2009 Effective Keyword Search for Software Resources Installed in Large-Scale Grid Infrastructures
abstract
In this paper, we investigate the problem of supporting keyword-based searching for the discovery of software resources that are installed on the nodes of large-scale, federated Grid computing infrastructures. We address a number of challenges that arise from the unstructured nature of software and the unavailability of software-related metadata on Grid sites. We present Minersoft, a Grid harvester that visits Grid sites, crawls their file-systems, identifies and classifies software resources, and discovers implicit associations between them. The results of Minersoft harvesting are encoded in a weighted, typed graph, named the Software Graph. A number of IR algorithms are used to enrich this graph with structural and content associations, to annotate software resources with keywords, and build inverted indexes to support keyword-based searching for software. Using a real testbed, we present an evaluation study of our approach, using data extracted from a production-quality Grid infrastructure. Experimental results show that our approach achieves high search efficiency.
George Pallis 0001, Asterios Katsifodimos, Marios D. Dikaiakos
Web Intelligence1
2009 CDNs Content Outsourcing via Generalized Communities
abstract
Content distribution networks (CDNs) balance costs and quality in services related to content delivery. Devising an efficient content outsourcing policy is crucial since, based on such policies, CDN providers can provide client-tailored content, improve performance, and result in significant economical gains. Earlier content outsourcing approaches may often prove ineffective since they drive prefetching decisions by assuming knowledge of content popularity statistics, which are not always available and are extremely volatile. This work addresses this issue, by proposing a novel self-adaptive technique under a CDN framework on which outsourced content is identified with no a-priori knowledge of (earlier) request statistics. This is employed by using a structure-based approach identifying coherent clusters of "correlated" Web server content objects, the so-called Web page communities. These communities are the core outsourcing unit and in this paper a detailed simulation experimentation has shown that the proposed technique is robust and effective in reducing user-perceived latency as compared with competing approaches, i.e., two communities-based approaches, Web caching, and non-CDN.
Dimitrios Katsaros 0001, George Pallis 0001, Konstantinos Stamos, Athena Vakali, Antonis Sidiropoulos 0001, Yannis Manolopoulos
IEEE Trans. Knowl. Data Eng.2
2008 Prefetching in Content Distribution Networks via Web Communities Identification and Outsourcing
Antonis Sidiropoulos 0001, George Pallis 0001, Dimitrios Katsaros 0001, Konstantinos Stamos, Athena Vakali, Yannis Manolopoulos
World Wide Web2
2007 Validation and interpretation of Web users' sessions clusters
George Pallis 0001, Lefteris Angelis, Athena Vakali
Inf. Process. Manag.1
2006 Integrating Caching Techniques on a Content Distribution Network
Konstantinos Stamos, George Pallis 0001, Athena Vakali
ADBIS2
2006 A similarity based approach for integrated Web caching and content replication in CDNs
abstract
Web caching and content replication techniques emerged to solve performance problems related to the Web. We propose a generic non-parametric heuristic method that integrates both techniques under a CDN. We provide experimentation showing that our method outperforms the so far separate implementations of Web caching and content replication. Moreover, we show that the performance improvement compared with an existing algorithm is significant. We test all these techniques in a simulation environment under a flash crowd event and a workload of a typical light-weighted CDN operation
Konstantinos Stamos, George Pallis 0001, Charilaos Thomos, Athena Vakali
IDEAS2
2005 Model-Based Cluster Analysis for Web Users Sessions
George Pallis 0001, Lefteris Angelis, Athena Vakali
ISMIS1