Georgios Theodoropoulos 0001

dblp:51/7514-1 · also Georgios K. Theodoropoulos 0001 · DBLP profile ↗
← Back
78ranked-venue papers
5as first author
15since 2021 · last 2026
0000-0002-7448-5886ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 35 · 4 since 2021Human-computer interaction and ubiquitous computing · 23 · 3 since 2021Systems, architecture and hardware · 21 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 15 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 15Software engineering, systems software and programming languages · 2 · 1 since 2021Computer networks · 1
YearPublicationVenuePosition
2026 Integrating Heterogeneous Digital Twins in Federated Ecosystems
Christian Vergara, Rami Bahsoon, Nikos Tziritas, Wendy Yanez-Pazmino, Panagiotis Oikonomou, Georgios Theodoropoulos 0001
MDM6
2026 TwinLoop: Simulation-in-the-Loop Digital Twins for Online Multi-Agent Reinforcement Learning
Nan Zhang 0027, Zishuo Wang, Shuyu Huang, Georgios Diamantopoulos, Nikos Tziritas, Panagiotis Oikonomou, Georgios Theodoropoulos 0001
MDM7
2026 QoS-aware placement of interdependent services in energy-harvesting-enabled multi-access edge computing
abstract
The advent of 5G drives the growth of multi-access edge computing (MEC), a revolutionary paradigm that utilises edge resources to enable low-latency mobile access and support complex service execution. Deploying services across geographically distributed edge nodes challenges providers to optimise performance metrics like end-to-end latency and resource efficiency, impacting user experience, operational cost, and environmental footprint. The energy harvesting (EH) technology provides clean and renewable energy at the edge, promoting the MEC system to minimise the impacts on the environment. However, the integration of EH can introduce energy limits and uncertainty to the powered devices. In the context of service scheduling with data flow dependencies, we propose two offline and heuristic-based service placement algorithms that balance minimizing latency and maximizing resource efficiency with fast execution. The two algorithms, evaluated in a simulated environment using state-of-the-art workload benchmarks, achieve significant energy consumption improvements while maintaining comparable latency. Based on the designed algorithms, we take a step further by developing an online dynamic resource scheduling and service offloading approach for MEC systems with EH capabilities. Simulation results demonstrate that the proposed strategy effectively utilise the harvested energy while granting a low user-experienced latency and low operational cost.
Panagiotis Oikonomou, Zhengchang Hua, Nikos Tziritas, Karim Djemame, Nan Zhang 0027, Georgios Theodoropoulos 0001
Future Gener. Comput. Syst.7
2026 Prototype Retrieval-Augmented Federated Learning System for Robust Intrusion Detection
abstract
Detecting malicious attacks is essential for protecting computer systems and ensuring device security. Federated Learning (FL)-based Intrusion Detection Systems (IDS) have emerged as promising solutions, enabling multiple clients (i.e., data owners) to collaboratively train intrusion detection models without sharing private data. However, current FL studies typically assume that each client’s training and test label distribution is identical. This assumption is overly idealistic and rarely holds in real-world scenarios, leading to suboptimal performance when label distribution shifts occur between the training and testing data. To address this challenge, we propose FedPRO, a plug-and-play framework designed to improve the test-time performance of existing FL methods, without modifying their original training pipelines or fine-tuning the trained FL models. Specifically, we develop a unique prototype generation and optimization mechanism to produce semantically meaningful class prototypes. These prototypes constitute a prototype memory bank, serving as an external knowledge repository. At test time, a prototype retrieval-augmented inference strategy is employed to query relevant prototypes and refine predictions on each client, effectively alleviating the label distribution shift issues and boosting prediction accuracy. We evaluate FedPRO by integrating it with various off-the-shelf FL methods on benchmark datasets. Extensive results consistently demonstrate its effectiveness in diverse settings. Notably, applying FedPRO to the state-of-the art method FedDBE improves its test accuracy from 79.25% to 86.66% on the CICIDS-2018 dataset, while introducing only approximately 32KB of additional communication overhead.
Hanlin Zhou, Huiru Yan, Jiawei Nian, Cong Liu 0012, Ying Wang 0001, Georgios Theodoropoulos 0001, Long Cheng 0003
IEEE Trans. Computers7
2026 A Hierarchical GNN-Based Multi-Agent Framework for Workflow Scheduling in Hybrid Clouds Considering Privacy Constraints
Hanlin Zhou, Cong Liu 0012, Fang Fang 0007, Zhiming Zhao, Georgios Theodoropoulos 0001, Long Cheng 0003
IEEE Trans. Serv. Comput.6
2025 A Digital Twin-Based Multi-agent Reinforcement Learning Framework for Vehicle-to-Grid Coordination
Zhengchang Hua, Panagiotis Oikonomou, Karim Djemame, Nikos Tziritas, Georgios Theodoropoulos 0001
ICA3PP (6)5
2024 Dynamic Digital Twins of Blockchain Systems: State Extraction and Mirroring
abstract
Blockchain adoption is reaching an all-time high, with a plethora of blockchain architectures being developed to cover the needs of applications eager to integrate blockchain into their operations. However, blockchain systems suffer from the trilemma trade-off problem, which limits their ability to scale without sacrificing essential metrics such as decentralisation and security. The balance of the trilemma trade-off is primarily dictated by the consensus protocol used. Since consensus protocols are designed to function well under specific system conditions, and consequently, due to the blockchain’s complex and dynamic nature, systems operating under a single consensus protocol are bound to face periods of inefficiency. The work presented in this paper constitutes part of an effort to design a Digital Twin-based blockchain management framework to balance the trilemma trade-off problem, which aims to adapt the consensus process to fit the conditions of the underlying system. Specifically, this work addresses the problems of extracting the blockchain system and mirroring it in its digital twin by proposing algorithms that overcome the challenges posed by blockchains’ decentralised and asynchronous nature and the fundamental problems of global state and synchronisation in such systems. The robustness of the proposed algorithms is experimentally evaluated.
Georgios Diamantopoulos, Nikos Tziritas, Rami Bahsoon, Nan Zhang 0027, Georgios Theodoropoulos 0001
DS-RT5
2024 Distributed Simulation for Digital Twins of Large-Scale Real-World DiffServ-Based Networks
Zhuoyao Huang, Nan Zhang 0027, Jingran Shen, Georgios Diamantopoulos, Zhengchang Hua, Nikos Tziritas, Georgios Theodoropoulos 0001
Euro-Par (3)7
2024 Towards LLM Augmented Discrete Event Simulation of Blockchain Systems
abstract
Despite recent leaps in artificial intelligence and natural language generation, which have led to widespread adoption, the integration of large language models in modelling and simulation has been limited. This work discusses the use of pre-trained large language models for the augmentation of a discrete event blockchain simulation system and their possible implications.
Georgios Diamantopoulos, Georgios Theodoropoulos 0001, Nikos Tziritas, Rami Bahsoon
SIGSIM-PADS2
2024 Federated Digital Twins as an Enabling Technology for Collaborative Decision-Making
abstract
Over the last few years, Digital Twin (DT) has emerged as an innovative concept that integrates multiple technologies to mirror physical assets, systems, and processes. Assisted by data analytics, predictive models, and optimisation techniques, DTs are suitable for enhancing operations in virtual space before transferring information into real-world counterparts. Moreover, initiatives considering DTs as part of a composite complex system might consider federated ecosystems for exchanging insights and relevant information. In this context, the Federated Digital Twin (FDT) concept is a potential solution to address interaction among virtual entities, enabling advanced operations and ensuring collaborative decision-making. This study describes a comprehensive FDT framework inspired by principles and methodologies considered by well-studied federated systems. Furthermore, a reference abstract architecture allowing seamless integration among multiple agent-based DTs is provided as a tool to develop a wide range of DT-based applications.
Christian Vergara, Georgios Theodoropoulos 0001, Rami Bahsoon, Wendy Yánez, Nikos Tziritas
SIGSIM-PADS2
2023 Dynamic Blockchain Reconfiguration: Balancing the Trilemma Trade-off Using Digital Twins
abstract
The trilemma trade-off problem between decentralisation, scalability, and security states that in blockchain systems the above properties are negatively correlated. Infrastructure, node configuration, choice of Consensus Protocol, and complexity of the underlying application are cited among the factors that affect the balance of the trade-off. Given that Blockchains are complex, dynamic systems, a dynamic approach to their management and reconfiguration at runtime is deemed necessary to reflect the changes in the state of the infrastructure and application. This work proposes the use of Digital Twins as the means of optimising the trilemma trade-off of blockchain i.e., re-configuring system parameters such as to maximise scalability, decentralisation and security. Specifically, through a bi-directional feedback loop between the system and the digital twin, simulation, what-if analysis and machine learning techniques will be employed for the computation of an optimal configuration given the current system state. Furthermore, a dynamic update mechanism is proposed to allow for blockchain reconfiguration without violating the decentralisation of the system.
Georgios Diamantopoulos, Nikos Tziritas, Rami Bahsoon, Georgios Theodoropoulos 0001
DS-RT4
2023 Federated Digital Twin
abstract
Digital Twin (DT) is a virtual replica of a physical system that is constantly receiving information from different data sources, enhancing its operations and processes through data analytics, predictions and simulations. The development of DTs relies on advancements in cutting-edge technologies namely IoT, Big data, Cloud computing and Artificial Intelligence; and although it was initially conceived in manufacturing, it is currently contributing to the digital transformation of several fields including aeronautics, healthcare, urban planning and agriculture. The existing body of research suggests that it will be expanded in the next few years with the implementation of sophisticated applications, therefore different proposals to achieve collective work between DTs have been investigated. Nevertheless, much research is needed to develop and validate appropriate mechanisms to ensure its successful deployment in complex real-world cases that require collaboration among individual systems. A Federated Digital Twin (FDT) has been identified as a promising solution for this approach, since it allows the interconnection among autonomous DTs in the virtual space, leveraging their advantages and enabling interaction, collaboration and shared learning. Additionally, since a FDT is envisaged as a network of cooperative DTs, cognitive principles can be applied to assist the overall operations through knowledge acquisition and reasoning, leading to an informed and intelligent decision making. This study aims to expand the FDT concept, develop mechanisms for coordination and synchronization based on well-defined FDT goals and connectionism theory. Furthermore, four architectural styles are provided to enable the integration of collaborative DTs within a federated environment, aiming to improve the operations in complex real-world systems.
Christian Vergara, Rami Bahsoon, Georgios Theodoropoulos 0001, Wendy Yánez, Nikos Tziritas
DS-RT3
2023 SymBChainSim: A Novel Simulation Tool for Dynamic and Adaptive Blockchain Management and its Trilemma Tradeoff
abstract
Despite the recent increase in the popularity of blockchain, the technology suffers from the trilemma trade-off between security decentralisation and scalability prohibiting adoption, and limiting the efficiency and effectiveness of the induced system. Addressing the trilemma trade-off calls for dynamic management and configuration of the blockchain system. In particular, choosing an effective and efficient consensus protocol for balancing the trilemma trade-off when inducing the blockchain-based system is acknowledged to be a challenging problem given the dynamic and complex nature of the blockchain environment. DDDAS approaches are particularly suitable for this challenge, and in previous work, the authors presented a novel DDDAS-based blockchain architecture and demonstrated that it offers a promising approach for dynamically adjusting the parameters of a system and optimising for the trade-off. This paper presents a novel simulation tool that can support and satisfy the DDDAS requirements for a dynamically re-configurable blockchain system. The tool supports the simulation and the dynamic switching of consensus protocols, analysing their trilemma trade-off. The simulator design is modular and allows the implementation and analysis of a wide range of consensus protocols and their implementation scenarios, along with the ability for parallelization. The paper also presents a quantitative evaluation of the tool.
Georgios Diamantopoulos, Rami Bahsoon, Nikos Tziritas, Georgios Theodoropoulos 0001
SIGSIM-PADS4
2022 Online Algorithms for the Interval Scheduling Problem in the Cloud: Affinity Pair Threshold Based Approaches
abstract
In the interval scheduling problem, jobs have known start and end times (referred to as job intervals) and must be assigned to processing nodes for their whole duration. Although the problem originally stems from the resource allocation demands of resident processes in operating systems, it found a renewed interest in the Cloud context, both in IaaS and SaaS, since reservations for virtual machines and services often have known activation intervals. A common objective of interval scheduling is to minimize busy time of machines which relates (among others) to minimizing the number of machines participating in the computation. As a consequence, bin packing techniques have been applied in the past. In this paper we tackle the online version of the problem, whereby future job arrivals are unknown. We propose novel algorithms that work as a pre-processing step to any bin packing scheme by offering recommendations that are enforced in all packing decisions. Job overlaps are used to characterize pairwise job affinity and subsequently provide threshold based job allocation recommendations. Thresholds are calculated using lower bound theoretical analysis upon two extreme workloads (sparse and dense). Experimental evaluation using real world workloads illustrates the merits of our approach against state-of-the-art algorithms.
Panagiotis Oikonomou, Nikos Tziritas, Thanasis Loukopoulos, Georgios Theodoropoulos 0001, Masatoshi Hanai, Samee Ullah Khan
IEEE Trans. Sustain. Comput.4
2021 A Probabilistic Batch Oriented Proactive Workflow Management
abstract
Workflow management is a widely studied research subject due to its criticality for the efficient execution of various processing activities towards concluding innovative applications. The ultimate goal is to eliminate the required time for delivering the final outcome considering the dependencies between workflow’s tasks. In this paper, we enhance the decision making of a scheduler with a batch oriented approach to deal with multiple workflows. A probabilistic data oriented approach combined with an infrastructure oriented scheme is provided to pay attention on dynamic environments where the underlying data are continuously updated trying to minimize the network overhead for migrating data. Workflows are mapped to the available datasets according to their data requirements, then, we combine the outcome with an optimization model upon the time and cost requirements of every placement. The performance of our model is revealed by a high number of experiments depicting the advantages in the network overhead.
Panagiotis Oikonomou, Kostas Kolomvatsos, Christos Anagnostopoulos 0001, Nikos Tziritas, Georgios Theodoropoulos 0001
ICTAI5
2020 Unveiling Ideological Trends Through Data Analytics to Construe National Security Instabilities
abstract
In this paper, a methodology to disclose ideological features using data analytics techniques aimed at interpreting national security instabilities is proposed. The analysis is based on two concepts, namely, authoritarianism and an attribute connected to it, hostility. Different computational techniques are used to address this a problem suchlike natural language processing, machine learning and deep learning models. The methodology proposed in this paper forms part of and enhances a previously reported holistic social media analysis framework for national security. The robustness and effectiveness of our approach are tested on one real-world event related to disruptive activity, protests in Puerto Rico in 2019.
Pedro Cárdenas, Boguslaw Obara, Georgios Theodoropoulos 0001, Ibad Kureshi
IEEE BigData3
2020 Graph-based Approaches for the Interval Scheduling Problem
abstract
One of the fundamental problems encountered by large-scale computing systems, such as clusters and cloud, is to schedule a set of jobs submitted by the users. Each job is characterized by resource demands, as well as start and completion time. Each job must be scheduled to execute on a machine having the required capacity between the start and completion time (referred as interval) of the job. Each machine is defined by a parallelism parameter g that indicates the maximum number of jobs that can be processed by the machine, in parallel. The above problem is referred to as the interval scheduling problem with bounded parallelism. The objective is to minimize the total busy time of all machines. Majority of the solutions proposed in the literature consider homogeneous set of jobs and machines that is a simplified assumption as in practice, heterogeneous jobs and machines are frequently encountered. In this article, we tackle the aforesaid problem with a set of heterogeneous jobs and machines. A major contribution of our work is that the problem is addressed in a novel way by combining a graph-based approach and a dynamic programming approach which is based on a variation of bin packing problem. A greedy algorithm is also proposed by employing only a graph-based approach at the aim to reduce the computational complexity. Experimental results show that the proposed algorithms can significantly reduce the cumulative busy interval over all machines compared with state-of-the-art algorithms proposed in the literature.
Panagiotis Oikonomou, Nikos Tziritas, Georgios Theodoropoulos 0001, Maria G. Koziri, Thanasis Loukopoulos, Samee Ullah Khan
ICPADS3
2020 Uncertainty Driven Workflow Scheduling Using Unreliable Cloud Resources
abstract
The Cloud infrastructure offers to end users a broad set of heterogenous computational resources using the pay-as-you -go model. These virtualized resources can be provisioned using different pricing models like the unreliable model where resources are provided at a fraction of the cost but with no guarantee for an uninterrupted processing. However, the enormous gamut of opportunities comes with a great caveat as resource management and scheduling decisions are increasingly complicated. Moreover, the presented uncertainty in optimally selecting resources has also a negatively impact on the quality of solutions delivered by scheduling algorithms. In this paper, we present a dynamic scheduling algorithm (i.e., the Uncertainty-Driven Scheduling - UDS algorithm) for the management of scientific workflows in Cloud. Our model minimizes both the makespan and the monetary cost by dynamically selecting reliable or unreliable virtualized resources. For covering the uncertainty in decision making, we adopt a Fuzzy Logic Controller (FLC) to derive the pricing model of the resources that will host every task. We evaluate the performance of the proposed algorithm using real workflow applications being tested under the assumption of different probabilities regarding the revocation of unreliable resources. Numerical results depict the performance of the proposed approach and a comparative assessment reveals the position of the paper in the relevant literature.
Panagiotis Oikonomou, Kostas Kolomvatsos, Nikos Tziritas, Georgios Theodoropoulos 0001, Thanasis Loukopoulos, Georgios I. Stamoulis
NCA4
2020 Efficient Direct Agent Interaction in Optimistic Distributed Multi-Agent-System Simulations
abstract
Agent-to-agent communications is an important operation in multi-agent systems and their simulation. Given the data-centric nature of agent-simulations, direct agent-to-agent communication is generally an orthogonal operation to accessing shared data in the simulation. In distributed multi-agent-system simulations in particular, implementing direct agent-to-agent communication may impose serious performance degradation due to potentially large communication and synchronization overheads. In this paper, we propose an efficient agent-to-agent communication method in the context of optimistic distributed simulation of multi-agent systems. An implementation of the proposed method is demonstrated and quantitatively evaluated through its integration into the PDES-MAS simulation kernel.
Masatoshi Hanai, Zhengchang Hua, Nikos Tziritas, Georgios Theodoropoulos 0001
SIGSIM-PADS5
2020 Towards Engineering Cognitive Digital Twins with Self-Awareness
abstract
There has been a recent explosion of interest in digital twins, namely data driven virtual replicas that can provide insights about a physical system and support decision making. This paper deals with cognitive digital twins, namely twins that can exhibit a high level of intelligence that can replicate human cognitive processes and execute conscious actions autonomously. The paper brings together the concepts of digital twins and self-awareness and discusses how the different levels of self-awareness can be harnessed for the design of cognitive-digital twins. A discussion of digital twins in relation to the Dynamic Data Driven Application Systems (DDDAS) paradigm and a classification of digital twins based on their analytics capability are also provided.
Nan Zhang 0027, Rami Bahsoon, Georgios Theodoropoulos 0001
SMC3
2019 On the Use of Neural Text Generation for the Task of Optical Character Recognition
abstract
Optical Character Recognition (OCR), is extraction of textual data from scanned text documents to facilitate their indexing, searching, editing and to reduce storage space. Although OCR systems have improved significantly in recent years, they still suffer in situations where the OCR output does not match the text in the original document. Deep learning models have contributed positively to many problems but their full potential to many other problems are yet to be explored. In this paper we propose a post-processing approach based on the application deep learning to improve the accuracy of OCR system (minimizing the error rate). We report on the use of neural network language models to accomplish the task of correcting incorrectly predicted characters/words by OCR systems. We applied our approach to the IAM handwriting database. Our proposed approach delivers significant accuracy improvement of 20.41% in F-score, 10.86% in character level comparison using Levenshtein distance and 20.69% in document level comparison over previously reported context based OCR empirical results of IAM handwriting database.
Mahnaz Mohammadi, Sardar F. Jaf, A. Stephen McGough, Toby P. Breckon, Peter Matthews, Georgios Theodoropoulos 0001, Boguslaw Obara
AICCSA6
2019 Temporal Neighbourhood Aggregation: Predicting Future Links in Temporal Graphs via Recurrent Variational Graph Convolutions
abstract
Graphs have become a crucial way to represent large, complex and often temporal datasets across a wide range of scientific disciplines. However, when graphs are used as input to machine learning models, this rich temporal information is frequently disregarded during the learning process, resulting in suboptimal performance on certain temporal inference tasks. To combat this, we introduce Temporal Neighbourhood Aggregation (TNA), a novel vertex representation model architecture designed to capture both topological and temporal information to directly predict future graph states. Our model exploits hierarchical recurrence at different depths within the graph to enable exploration of changes in temporal neighbourhoods, whilst requiring no additional features or labels to be present. The final vertex representations are created using variational sampling and are optimised to directly predict the next graph in the sequence. Our claims are supported by experimental evaluation on both real and synthetic benchmark datasets, where our approach demonstrates superior performance compared to competing methods, outperforming them at predicting new temporal edges by as much as 23% on real-world datasets, whilst also requiring fewer overall model parameters.
Stephen Bonner, Amir Atapour Abarghouei, Philip T. G. Jackson, John Brennan, Ibad Kureshi, Georgios Theodoropoulos 0001, A. Stephen McGough, Boguslaw Obara
IEEE BigData6
2019 Analysing Social Media as a Hybrid Tool to Detect and Interpret likely Radical Behavioural Traits for National Security
abstract
The study of National Security and its associated considerations is a sensitive and complex paradigm. It encapsulates both the protection of the territorial integrity and sovereignty of a state, as well as guaranteeing the security of its population. Known as Human Security, human-centred threats arising from radical activities need to be mitigated else they may escalate and have implications on National Security. The modern era has introduced further disruptive challenges, known as Hybrid Threats, that use non-traditional tools (Hybrid Tools) to intensify the impact of a likely threat. Social Media is a clear illustration of such tools, where the stability of the state and its people can be compromised by the dissemination of material. The ability to identify behaviour bordering on criminality within the deregulated world of Social Media is a Human Security imperative for governments. This paper follows on from our earlier work to detect affected National Security variables through the analysis of social media communication and trigger an alert when a likely threat is detected. As a result, a set of crisis interpretation processes are started to construe the event, such as radical behaviour analysis.This paper details the methodological approach to analyse one Hybrid Tool (Social Media) in order to identify likely instability scenarios based on the Human Security spectrum and therefore extract, detect and interpret dissimilar behavioural patterns that outline radical behavioural traits for National Security. The proposed methodology focuses on five steps, namely Instability Scenarios, Entity Extraction, Wordlists Creation, Content Analytics, and Data Interpretation.
Pedro Cárdenas, Boguslaw Obara, Georgios Theodoropoulos 0001, Ibad Kureshi
IEEE BigData3
2019 Island Model Genetic Algorithm for Feature Selection in Non-Traditional Credit Risk Evaluation
abstract
As digital infrastructure expands in new regions of the globe, developing ways to include more diverse information in financial decisions is important. However, making use of novel data sources requires developing methods to evaluate credit with diverse and complex datasets with missing information, dynamic patterns and relationships with decision recommendations, and larger feature sets. Feature selection is one approach that can support the application of machine learning to dynamically build models for credit evaluation with complex data. Genetic algorithms (GAs) have been proved to reach good performance in other research, with high computation cost though. In this paper, we review existing GA approaches and test and develop a novel method based on niching and the use of subpopulations with different data for fitness evaluation. This formulation allows less computation cost, even with better prediction performance in feature selection. In further experiments, we compare the proposed GA-based feature selection approaches in four traditional credit datasets and a novel emerging market dataset from China. The results indicate that the advanced GA-based feature selection methods perform more effectively.
Adam Ghandar, Georgios Theodoropoulos 0001
CEC3
2019 A Metaheuristic Strategy for Feature Selection Problems: Application to Credit Risk Evaluation in Emerging Markets
abstract
As countries develop digital financial infrastructure, a wide range of economic activities expand and grow in importance: from personal loans, to the rapidly developing networked microfinance industry, to mobile telephone services and real estate transactions and so on. Personal credit is also a foundation of trust for facilitation of integrated societal transactions more generally. In emerging markets there is, however, a gap between the requirement for establishing a credit or trust rating and the lack of a credit record. The development of methodologies for greater financial integration of growing economies has the potential to have a significant impact on increasing the GDP of developing economies (4-12% according to a recent McKinsey Global Institute report). In this paper, we develop and test a methodology for feature selection and test its in standard datasets from large institutions in mature market economies, and a recent dataset which illustrates characteristics of emerging markets. The results show performance in classification can be maintained while runtime can be reduced when using a GA for feature selection in a range of machine learning techniques.
Adam Ghandar, Georgios Theodoropoulos 0001
CIFEr3
2019 Exploring the Semantic Content of Unsupervised Graph Embeddings: An Empirical Study
abstract
Graph embeddings have become a key and widely used technique within the field of graph mining, proving to be successful across a broad range of domains including social, citation, transportation and biological. Unsupervised graph embedding techniques aim to automatically create a low-dimensional representation of a given graph, which captures key structural elements in the resulting embedding space. However, to date, there has been little work exploring exactly which topological structures are being learned in the embeddings, which could be a possible way to bring interpretability to the process. In this paper, we investigate if graph embeddings are approximating something analogous to traditional vertex-level graph features. If such a relationship can be found, it could be used to provide a theoretical insight into how graph embedding approaches function. We perform this investigation by predicting known topological features, using supervised and unsupervised methods, directly from the embedding space. If a mapping between the embeddings and topological features can be found, then we argue that the structural information encapsulated by the features is represented in the embedding space. To explore this, we present extensive experimental evaluation with five state-of-the-art unsupervised graph embedding techniques, across a range of empirical graph datasets, measuring a selection of topological features. We demonstrate that several topological features are indeed being approximated in the embedding space, allowing key insight into how graph embeddings create good representations.
Stephen Bonner, Ibad Kureshi, John Brennan, Georgios Theodoropoulos 0001, A. Stephen McGough, Boguslaw Obara
Data Sci. Eng.4
2019 Distributed Edge Partitioning for Trillion-edge Graphs
abstract
We propose Distributed Neighbor Expansion (Distributed NE), a parallel and distributed graph partitioning method that can scale to trillion-edge graphs while providing high partitioning quality. Distributed NE is based on a new heuristic, called parallel expansion, where each partition is constructed in parallel by greedily expanding its edge set from a single vertex in such a way that the increase of the vertex cuts becomes local minimal. We theoretically prove that the proposed method has the upper bound in the partitioning quality. The empirical evaluation with various graphs shows that the proposed method produces higher-quality partitions than the state-of-the-art distributed graph partitioning algorithms. The performance evaluation shows that the space efficiency of the proposed method is an order-of-magnitude better than the existing algorithms, keeping its time efficiency comparable. As a result, Distributed NE can partition a trillion-edge graph using only 256 machines within 70 minutes.
Masatoshi Hanai, Toyotaro Suzumura, Wen Jun Tan, Elvis S. Liu, Georgios Theodoropoulos 0001, Wentong Cai 0001
Proc. VLDB Endow.5
2018 Temporal Graph Offset Reconstruction: Towards Temporally Robust Graph Representation Learning
abstract
Graphs are a commonly used construct for representing relationships between elements in complex high dimensional datasets. Many real-world phenomenon are dynamic in nature, meaning that any graph used to represent them is inherently temporal. However, many of the machine learning models designed to capture knowledge about the structure of these graphs ignore this rich temporal information when creating representations of the graph. This results in models which do not perform well when used to make predictions about the future state of the graph - especially when the delta between time stamps is not small. In this work, we explore a novel training procedure and an associated unsupervised model which creates graph representations optimised to predict the future state of the graph. We make use of graph convo-lutional neural networks to encode the graph into a latent representation, which we then use to train our temporal offset reconstruction method, inspired by auto-encoders, to predict a later time point - multiple time steps into the future. Using our method, we demonstrate superior performance for the task of future link prediction compared with none-temporal state-of-the-art baselines. We show our approach to be capable of outperforming non-temporal baselines by 38% on a real world dataset.
Stephen Bonner, John Brennan, Ibad Kureshi, Georgios Theodoropoulos 0001, A. Stephen McGough, Boguslaw Obara
IEEE BigData4
2018 Defining an Alert Mechanism for Detecting likely threats to National Security
abstract
The paper presents an Alert Mechanism for analysing and detecting National Security threats using Social Media posts as the primary source of information. This mechanism is meant to be an early warning system that can identify situations where a critical mass of individuals feel attracted towards a disruptive cause, based on emotions and Human Security aspects, which may lead to societal tipping points. Comprehensive experiments on real-world events related to disruptive and non-disruptive cases demonstrate both the robustness and effectiveness of the proposed mechanism.
Pedro Cárdenas, Boguslaw Obara, Georgios Theodoropoulos 0001, Ibad Kureshi
IEEE BigData3
2018 TMIXT: A process flow for Transcribing MIXed handwritten and machine-printed Text
abstract
Handling large corpuses of documents is of significant importance in many fields, no more so than in the areas of crime investigation and defence, where an organisation may be presented with a large volume of scanned documents which need to be processed in a finite time. However, this problem is exacerbated both by the volume, in terms of scanned documents and the complexity of the pages, which need to be processed. Often containing many different elements, which each need to be processed and understood. Text recognition, which is a primary task of this process, is usually dependent upon the type of text, being either handwritten or machine-printed. Accordingly, the recognition involves prior classification of the text category, before deciding on the recognition method to be applied. This poses a more challenging task if a document contains both handwritten and machine-printed text. In this work, we present a generic process flow for text recognition in scanned documents containing mixed handwritten and machine-printed text without the need to classify text in advance. We realize the proposed process flow using several open-source image processing and text recognition packages. The evaluation is performed using a specially developed variant, presented in this work, of the IAM handwriting database, where we achieve an average transcription accuracy of nearly 80% for pages containing both printed and handwritten text.
Fady Medhat, Mahnaz Mohammadi, Sardar F. Jaf, Chris G. Willcocks, Toby P. Breckon, Peter Matthews, A. Stephen McGough, Georgios Theodoropoulos 0001, Boguslaw Obara
IEEE BigData8
2018 Minimizing Network Traffic for Distributed Joins Using Lightweight Locality-Aware Scheduling
Long Cheng 0003, John Murphy 0001, Qingzhi Liu, Chunliang Hao, Georgios Theodoropoulos 0001
Euro-Par5
2017 Evaluating the quality of graph embeddings via topological feature reconstruction
abstract
In this paper we study three state-of-the-art, but competing, approaches for generating graph embeddings using unsupervised neural networks. Graph embeddings aim to discover the `best' representation for a graph automatically and have been applied to graphs from numerous domains, including social networks. We evaluate their effectiveness at capturing a good representation of a graph's topological structure by using the embeddings to predict a series of topological features at the vertex level. We hypothesise that an `ideal' high quality graph embedding should be able to capture key parts of the graph's topology, thus we should be able to use it to predict common measures of the topology, for example vertex centrality. This could also be used to better understand which topological structures are truly being captured by the embeddings. We first review these three graph embedding techniques and then evaluate how close they are to being `ideal'. We provide a framework, with extensive experimental evaluation on empirical and synthetic datasets, to assess the effectiveness of several approaches at creating graph embeddings which capture detailed topological structure.
Stephen Bonner, John Brennan, Ibad Kureshi, Georgios Theodoropoulos 0001, A. Stephen McGough, Boguslaw Obara
IEEE BigData4
2017 Improving the robustness and performance of parallel joins over distributed systems
Long Cheng 0003, Spyros Kotoulas, Tomás Ward, Georgios Theodoropoulos 0001
J. Parallel Distributed Comput.4
2016 Deep topology classification: A new approach for massive graph classification
abstract
The classification of graphs is a key challenge within many scientific fields using graphs to represent data and is an active area of research. Graph classification can be critical in identifying and labelling unknown graphs within a dataset and has seen application across many scientific fields. Graph classification poses two distinct problems: the classification of elements within a graph and the classification of the entire graph. Whilst there is considerable work on the first problem, the efficient and accurate classification of massive graphs into one or more classes has, thus far, received less attention. In this paper we propose the Deep Topology Classification (DTC) approach for global graph classification. DTC extracts both global and vertex level topological features from a graph to create a highly discriminate representation in feature space. A deep feed-forward neural network is designed and trained to classify these graph feature vectors. This approach is shown to be over 99% accurate at discerning graph classes over two datasets. Additionally, it is shown to be more accurate than current state of the art approaches both in binary and multi-class graph classification tasks.
Stephen Bonner, John Brennan, Georgios Theodoropoulos 0001, Ibad Kureshi, A. Stephen McGough
IEEE BigData3
2016 GFP-X: A parallel approach to massive graph comparison using spark
abstract
The problem of how to compare empirical graphs is an area of great interest within the field of network science. The ability to accurately but efficiently compare graphs has a significant impact in such areas as temporal graph evolution, anomaly detection and protein comparison. The comparison problem is compounded when working with massive graphs containing millions of vertices and edges. This paper introduces a parallel feature extraction based approach for the efficient comparison of large unlabelled graph datasets using Apache Spark. The approach acts by producing a `Graph Fingerprint' which represents both vertex level and global level topological features from a graph. By using Spark we are able to efficiently compare graphs considered unmanageably large to other approaches. The runtime of the approach is shown to scale sub-linearly with the size and complexity of the graphs being fingerprinted. Importantly, the approach is shown to not only be comparable to existing approaches, but on when comparing topology and size, more sensitive at detecting variation between graphs.
Stephen Bonner, John Brennan, Georgios Theodoropoulos 0001, Ibad Kureshi, A. Stephen McGough
IEEE BigData3
2016 Fast Compression of Large Semantic Web Data Using X10
abstract
The Semantic Web comprises enormous volumes of semi-structured data elements. For interoperability, these elements are represented by long strings. Such representations are not efficient for the purposes of applications that perform computations over large volumes of such information. A common approach to alleviate this problem is through the use of compression methods that produce more compact representations of the data. The use of dictionary encoding is particularly prevalent in Semantic Web database systems for this purpose. However, centralized implementations present performance bottlenecks, giving rise to the need for scalable, efficient distributed encoding schemes. In this paper, we propose an efficient algorithm for fast encoding large Semantic Web data. Specially, we present the detailed implementation of our approach based on the state-of-art asynchronous partitioned global address space (APGAS) parallel programming model. We evaluate performance on a cluster of up to 384 cores and datasets of up to 11 billion triples (1.9 TB). Compared to the state-of-art approach, we demonstrate a speed-up of$2.6 - 7.4\times$and excellent scalability. In the meantime, these results also illustrate the significant potential of the APGAS model for efficient implementation of dictionary encoding and contributes to the engineering of more efficient, larger scale Semantic Web applications.
Long Cheng 0003, Avinash Malik, Spyros Kotoulas, Tomás Ward, Georgios Theodoropoulos 0001
IEEE Trans. Parallel Distributed Syst.5
2015 Data quality assessment and anomaly detection via map/reduce and linked data: A case study in the medical domain
abstract
Recent technological advances in modern healthcare have lead to the ability to collect a vast wealth of patient monitoring data. This data can be utilised for patient diagnosis but it also holds the potential for use within medical research. However, these datasets often contain errors which limit their value to medical research, with one study finding error rates ranging from 2.3%-26.9% in a selection of medical databases. Previous methods for automatically assessing data quality normally rely on threshold rules, which are often unable to correctly identify errors, as further complex domain knowledge is required. To combat this, a semantic web based framework has previously been developed to assess the quality of medical data. However, early work, based solely on traditional semantic web technologies, revealed they are either unable or inefficient at scaling to the vast volumes of medical data. In this paper we present a new method for storing and querying medical RDF datasets using Hadoop Map / Reduce. This approach exploits the inherent parallelism found within RDF datasets and queries, allowing us to scale with both dataset and system size. Unlike previous solutions, this framework uses highly optimised (SPARQL) joining strategies, intelligent data caching and the use of a super-query to enable the completion of eight distinct SPARQL lookups, comprising over eighty distinct joins, in only two Map / Reduce iterations. Results are presented comparing both the Jena and a previous Hadoop implementation demonstrating the superior performance of the new methodology. The new method is shown to be five times faster than Jena and twice as fast as the previous approach.
Stephen Bonner, A. Stephen McGough, Ibad Kureshi, John Brennan, Georgios Theodoropoulos 0001, Laura Moss, David Corsar, Grigoris Antoniou
IEEE BigData5
2015 Towards an Info-Symbiotic Decision Support System for Disaster Risk Management
abstract
This paper outlines a framework for an info-symbiotic modelling system using cyber-physical sensors to assist in decision-making. Using a dynamic data-driven simulation approach, this system can help with the identification of target areas and resource allocation in emergency situations. Using different natural disasters as exemplars, we will show how cyber-physical sensors can enhance ground level intelligence and aid in the creation of dynamic models to capture the state of human casualties. Using a virtual command & control centre communicating with sensors in the field, up-to-date information of the ground realities can be incorporated in a dynamic feedback loop. Using other information (e.g. Weather models) a complex and rich model can be created. The framework adaptively manages the heterogeneous collection of data resources and uses agent-based models to create what-if scenarios in order to determine the best course of action.
Ibad Kureshi, Georgios Theodoropoulos 0001, Eleni E. Mangina, Gregory M. P. O'Hare, John Roche
DS-RT2
2015 Exact-Differential Large-Scale Traffic Simulation
abstract
Analyzing large-scale traffics by simulation needs repeating execution many times with various patterns of scenarios or parameters. Such repeating execution brings about big redundancy because the change from a prior scenario to a later scenario is very minor in most cases, for example, blocking only one of roads or changing the speed limit of several roads. In this paper, we propose a new redundancy reduction technique, called exact-differential simulation, which enables to simulate only changing scenarios in later execution while keeping exactly same results as in the case of whole simulation. The paper consists of two main efforts: (i) a key idea and algorithm of the exact-differential simulation, (ii) a method to build large-scale traffic simulation on the top of the exact-differential simulation. In experiments of Tokyo traffic simulation, the exact-differential simulation shows 7.26 times as much elapsed time improvement in average and 2.26 times improvement even in the worst case as the whole simulation.
Masatoshi Hanai, Toyotaro Suzumura, Georgios Theodoropoulos 0001, Kalyan S. Perumalla
SIGSIM-PADS3
2015 Simulation in the era of Big Data: Trends and Challenges
abstract
The emergence of extreme scale computing systems and the data explosion have presented an unprecedented opportunity for the analysis of systems at a rapidly increasing scale, complexity and granularity. This paradigm shift calls for an intermingling of 'what-if' and data analytics approaches, however the worlds of Simulation and Big Data have so far been largely separate. The talk will focus on the interplay between simulation, data and emerging computational platforms, identifying gaps and opportunities and discussing some concrete examples of interacting scalable data infrastructures and agent-based simulations.
Georgios Theodoropoulos 0001
SIGSIM-PADS1
2014 Automated Dynamic Resource Provisioning and Monitoring in Virtualized Large-Scale Datacenter
abstract
Infrastructure as a Service (IaaS) is a pay-as-you go based cloud provision model which on demand outsources the physical servers, guest virtual machine (VM) instances, storage resources, and networking connections. This article reports the design and development of our proposed innovative symbiotic simulation based system to support the automated management of IaaS-based distributed virtualized data enter. To make the ideas work in practice, we have implemented an Open Stack based open source cloud computing platform. A smart benchmarking application "Cloud Rapid Experimentation and Analysis Tool (aka CBTool)" is utilized to mark the resource allocation potential of our test cloud system. The real-time benchmarking metrics of cloud are fed to a distributed multi-agent based intelligence middleware layer. To optimally control the dynamic operation of prototype data enter, we predefine some custom policies for VM provisioning and application performance profiling within a versatile cloud modeling and simulation toolkit "CloudSim". Both tools for our prototypes' implementation can scale up to thousands of VMs, therefore, our devised mechanism is highly scalable and flexibly be interpolated at large-scale level. Autonomic characteristics of agents aid in streamlining symbiosis among the simulation system and IaaS cloud in a closed feedback control loop. The practical worth and applicability of the multiagent-based technology lies in the fact that this technique is inherently scalable hence can efficiently be implemented within the complex cloud computing environment. To demonstrate the efficacy of our approach, we have deployed an intelligible lightweight representative scenario in the context of monitoring and provisioning virtual machines within the test-bed. Experimental results indicate notable improvement in the resource provision profile of virtualized data enter on incorporating our proposed strategy.
Sameera Abar, Pierre Lemarinier, Georgios Theodoropoulos 0001, Gregory M. P. O'Hare
AINA3
2014 Efficiently Handling Skew in Outer Joins on Distributed Systems
abstract
Outer joins are ubiquitous in databases and big data systems. The question of how best to execute outer joins in large parallel systems is particularly challenging as real world datasets are characterized by data skew leading to performance issues. Although skew handling techniques have been extensively studied for inner joins, there is little published work solving the corresponding problem for parallel outer joins. Conventional approaches to this problem such as ones based on hash redistribution often lead to load balancing problems while duplication-based approaches incurs significant overhead in terms of network communication. In this paper, we propose a new algorithm, query with counters (QC), for directly handling skew in outer joins on distributed architectures. We present an efficient implementation of our approach based on the asynchronous partitioned global address space (APGAS) parallel programming model. We evaluate the performance of our approach on a cluster of 192 cores (16 nodes) and datasets of 1 billion tuples with different skew. Experimental results show that our method is scalable and, in cases of high skew, faster than the state-of-the-art.
Long Cheng 0003, Spyros Kotoulas, Tomás Ward, Georgios Theodoropoulos 0001
CCGRID4
2014 Robust and Skew-resistant Parallel Joins in Shared-Nothing Systems
abstract
The performance of joins in parallel database management systems is critical for data intensive operations such as querying. Since data skew is common in many applications, poorly engineered join operations result in load imbalance and performance bottlenecks. State-of-the-art methods designed to handle this problem offer significant improvements over naive implementations. However, performance could be further improved by removing the dependency on global skew knowledge and broadcasting. In this paper, we propose PRPQ (partial redistribution & partial query), an efficient and robust join algorithm for processing large-scale joins over distributed systems. We present the detailed implementation and a quantitative evaluation of our method. The experimental results demonstrate that the proposed PRPQ algorithm is indeed robust and scalable under a wide range of skew conditions. Specifically, compared to the state-of-art PRPD method, we achieve 16% - 167% performance improvement and 24% - 54% less network communication under different join workloads.
Long Cheng 0003, Spyros Kotoulas, Tomás Ward, Georgios Theodoropoulos 0001
CIKM4
2014 Robust and Efficient Large-Large Table Outer Joins on Distributed Infrastructures
Long Cheng 0003, Spyros Kotoulas, Tomás Ward, Georgios Theodoropoulos 0001
Euro-Par4
2014 Design and evaluation of parallel hashing over large-scale data
abstract
High-performance analytical data processing systems often run on servers with large amounts of memory. A common data structure used in such environment is the hash tables. This paper focuses on investigating efficient parallel hash algorithms for processing large-scale data. Currently, hash tables on distributed architectures are accessed one key at a time by local or remote threads while shared-memory approaches focus on accessing a single table with multiple threads. A relatively straightforward “bulk-operation” approach seems to have been neglected by researchers. In this work, using such a method, we propose a high-level parallel hashing framework, Structured Parallel Hashing, targeting efficiently processing massive data on distributed memory. We present a theoretical analysis of the proposed method and describe the design of our hashing implementations. The evaluation reveals a very interesting result - the proposed straightforward method can vastly outperform distributed hashing methods and can even offer performance comparable with approaches based on shared memory supercomputers which use specialized hardware predicates. Moreover, we characterize the performance of our hash implementations through extensive experiments, thereby allowing system developers to make a more informed choice for their high-performance applications.
Long Cheng 0003, Spyros Kotoulas, Tomás Ward, Georgios Theodoropoulos 0001
HiPC4
2014 A reconfigurable, regular-topology cluster/datacenter network using commodity optical switches
Diego Lugones, Kostas Katrinis, Georgios Theodoropoulos 0001, Martin Collier
Future Gener. Comput. Syst.3
2014 Editorial: Special Issue on Extreme Scale Parallel Architectures and Systems
Georgios Theodoropoulos 0001, Kostas Katrinis, Rolf Riesen, Shoukat Ali
Future Gener. Comput. Syst.1
2012 MWGrid: A System for Distributed Agent-Based Simulation in the Digital Humanities
abstract
Digital Humanities offer a new exciting domain for agent-based distributed simulation. In historical studies interpretation rarely rises above the level of unproven assertion and is rarely tested against a range of evidence. Agent-based simulation can provide an opportunity to break these cycles of academic claim and counterclaim. The MWGrid framework utilises distributed agent based simulation to study medieval military logistics. As a use-case, it has focused on the logistical analysis of the Byzantine army's march to the battle of Manzikert (AD 1071), a key event in medieval history. It integrates an agent design template, a transparent, layered mechanism to translate model-level agents' actions to time stamped events and the PDES-MAS distributed simulation kernel. The paper presents an overview of the MWGrid system and a quantitative evaluation of its performance.
Bart G. W. Craenen, Phil Murgatroyd, Georgios Theodoropoulos 0001, Vincent Gaffney, Vinoth Suryanarayanan
DS-RT3
2012 Parallel Simulation Models for the Evaluation of Future Large-Scale Datacenter Networks
abstract
The recent trend for Software as a Service and other types of cloud services is driving demand for data centers of ever increasing scale. This will require the scaling up and scaling out of existing data center architectures and will ultimately need new architectures featuring new topologies for the data center networks that interconnect its constituent nodes. The preferred tool for the detailed evaluation of such designs is simulation, as it gives finer detail than analysis at lower cost than test bed evaluation. Realistic simulation times require that the simulation itself be capable of being scaled out, i.e., it should be amenable to parallelization. We describe a simulation framework, together with a methodology for partitioning the relevant simulation models to allow their parallel implementation, and demonstrate its validity by applying it to the simulation of a hybrid optical/electrical network architecture using a cluster of high-end servers. We report on results capturing the performance of our simulator and discuss how these are associated with the underlying hardware hosting the simulation.
Diego Lugones, Kostas Katrinis, Martin Collier, Georgios Theodoropoulos 0001
DS-RT4
2012 SParTSim: A Space Partitioning Guided by Road Network for Distributed Traffic Simulations
abstract
Traffic simulation can be very computationally intensive, especially for microscopic simulations of large urban areas (tens of thousands of road segments, hundreds of thousands of agents) and when real-time or better than real-time simulation is required. For instance, running a couple of what-if scenarios for road management authorities/police during a road incident: time is a hard constraint and the size of the simulation is relatively high. Hence the need for distributed simulations and for optimal space partitioning algorithms, ensuring an even distribution of the load and minimal communication between computing nodes. In this paper we describe a distributed version of SUMO, a simulator of urban mobility, and SParTSim, a space partitioning algorithm guided by road network for distributed simulations. It outperforms classical uniform space partitioning in terms of road segment cuts and load-balancing.
Anthony Ventresque, Quentin Bragard, Elvis S. Liu, Dawid Nowak, Liam Murphy 0001, Georgios Theodoropoulos 0001
DS-RT6
2012 Still Alive: Extending Keep-Alive Intervals in P2P Overlay Networks
Richard Price, Peter Tiño, Georgios Theodoropoulos 0001
Mob. Networks Appl.3
2011 A Parallel Interest Matching Algorithm for Distributed-Memory Systems
abstract
As the scale of Distributed Virtual Environments (DVEs) grows in terms of participants and virtual entities, using interest management schemes to reduce bandwidth consumption becomes increasingly common for DVE development. The interest matching process is essential for most of the interest management schemes which determines what data should be sent to the participants as well as what data should be filtered. However, if the computational overhead of interest matching is too high, it would be unsuitable for real-time DVEs for which runtime performance is important. This paper presents a new approach of interest matching which divides the workload of matching process among a cluster of computers. Experimental evidence shows that our approach is an effective solution for the real-time applications.
Elvis S. Liu, Georgios Theodoropoulos 0001
DS-RT2
2011 A Multi-scale Agent-Based Distributed Simulation Framework for Groundwater Pollution Management
abstract
Groundwater is like dark matter - we know very little apart from the fact that it is hugely important. Given the scarcity of data, mathematical modelling can come to the rescue but existing groundwater models are mainly restricted to simulate the transport and degradation of contaminants on the scale of whole contaminated field sites by averaging out the effect of spatial heterogeneity on the availability of the pollutant to the degrading organisms. These coarse-scale mean-field models therefore tend to rely on fitting to data rather than being predictive. Also, they are less suited to incorporate spatial variability and non-linear kinetics and feedbacks. We propose to solve the two mutually exacerbating problems of environmental patchiness and data scarcity by developing a flexible and robust distributed simulation framework that uses an ensemble of small scale simulations running on different processors/computers to scale-up, i.e. to feed the effect of small-scale patchiness into a concurrent site-scale simulation of the dynamics of groundwater pollutant degradation. Our scaling approach solves problem #1 by simulating dynamics also on the small scale where some of the patchiness resides, and problem #2 by enabling rigorous validation of our small-scale model and scaling approach with laboratory data, which are high quality at low cost.
Susanne I. Schmidt, Cristian Picioreanu, Bart G. W. Craenen, Rae Mackay, Jan-Ulrich Kreft, Georgios Theodoropoulos 0001
DS-RT6
2010 Synchronised Range Queries in Distributed Simulations of Multi-agent Systems
abstract
Range-Query is an important associative form of data access in distributed simulations and Distributed Virtual Environments. This paper discusses the problem of Range-Queries in the context of distributed simulation of multi-agent systems. An algorithm is presented for performing instantaneous Queries within an optimistic synchronisation framework and in the presence of dynamic migration of the simulation state. A quantitative evaluation of the effectiveness of the algorithm under different conditions is also presented.
Vinoth Suryanarayanan, Bart G. W. Craenen, Georgios Theodoropoulos 0001
DS-RT3
2010 An Integrated Approach to Learning Object Sequencing
abstract
The use of learning objects (LOs) to support learning processes is considered a key factor in the deployment of e-learning frameworks. Although their discrete and self-contained nature offers many advantages LOs have presented courseware designers with two fundamental issues. The first issue concerns the adequate identification of the learning objects and the second relates to their integration into a suitable learning programme. This paper is concerned with the presentation of a framework that addresses and reconciles these requirements by capitalising on the metadata of the learning objects and on the profiles of the learners. The learning process is mediated by a core component, a learning management system (LMS) that, through a registry, accesses the metadata of the learning objects published by providers and acts as a repository of metadata. In addition, it enables learners to construct learning paths (LP) from a set of relevant LOs in accordance with their profile. The LMS can also generate automatically a learning path on behalf of a learner and determine the schedule of the learning process. The implementation of the framework takes advantage of the flexibility of Web Services.
Battur Tugsgerel, Rachid Anane, Georgios Theodoropoulos 0001
ICALT3
2010 Synchronization in federation community networks
Dan Chen 0001, Stephen John Turner, Wentong Cai 0001, Georgios Theodoropoulos 0001, Muzhou Xiong, Michael Lees
J. Parallel Distributed Comput.4
2009 An Approach for Parallel Interest Matching in Distributed Virtual Environments
abstract
Interest management is essential for real-time large-scale distributed virtual environments (DVEs) which seeks to filter irrelevant messages on the network. Many existing interest management schemes such as HLA DDM focus on providing precise message filtering mechanisms. However, this leads to a second problem: the computational overhead of the interest matching process. If the CPU cost of interest matching is too high, it would be unsuitable for real-time applications such as multiplayer online games for which runtime performance is important. This paper evaluates the performance of existing interest matching algorithms and proposes a new algorithm based on parallel processing. The new algorithm is expected to have better computational efficiency than existing algorithms and maintain the same accuracy of message filtering as them. Experimental evidence shows that our approach works well in practice.
Elvis S. Liu, Georgios Theodoropoulos 0001
DS-RT2
2009 Synchronised Range Queries
abstract
In this paper, we present and evaluate a system for performing logical-time synchronised Range Queries over data in the context of parallel and distributed simulations of Multi-Agent Systems (MAS). MAS are often extremely complex and simulation is commonly used to understand their behaviour or investigate the implications of alternative agent architectures. Range Queries are widely used in various fields such as Peer to Peer systems, Wireless communications or Database systems. They are key to many MAS models as they are commonly used to represent the spatial perceptive abilities of the agents in the MAS. PDES-MAS (Parallel and Discrete Event Simulation for Multi-Agent Systems) is a decentralised, discrete event simulation (DES) system which can be used to distribute and run a large scale MAS simulation over a parallel computation architecture. This paper presents a design for Logical-Time synchronised Range Queries and the implementation and evaluation of this design within the PDES-MAS system.
Vinoth Suryanarayanan, Rob Minson, Georgios Theodoropoulos 0001
DS-RT3
2009 Analysing probabilistically constrained optimism
abstract
Abstract In previous work we presented the DTRD algorithm, an optimistic synchronization algorithm for parallel discrete event simulation of multi‐agent systems, and showed that it outperforms Time Warp and time windows on a range of test cases. DTRD uses a decision‐theoretic model of rollback to derive an optimal time to delay read event so as to maximize the rate of LVT progression. The algorithm assumes that the inter‐arrival times (both virtual and real) of events are normally distributed. In this paper we present a more detailed evaluation of the DTRD algorithm, and specifically how the performance of the algorithm is affected when the inter‐arrival times do not follow the assumed distributions. Our analysis suggests that the performance of the algorithm is relatively insensitive to events whose inter‐arrival times are not normally distributed. However, as the variance of event inter‐arrival times increases, its performance degrades to that of Time Warp. The evaluation approach we present is generally applicable, and we sketch how a similar analysis may be performed for two other decision‐theoretic optimistic synchronization algorithms. Copyright © 2009 John Wiley & Sons, Ltd.
Michael Lees, Brian Logan 0001, Georgios Theodoropoulos 0001
Concurr. Comput. Pract. Exp.3
2008 Evaluating Large Scale Distributed Simulation of P2P Networks
abstract
P2P systems have witnessed phenomenal development in recent years. Evaluating and analyzing new and existing algorithms and techniques is a key issue for developers of P2P systems. In this context, simulation is an important tool for P2P developers. However, such systems are often very large and few existing simulators offer the ability to execute simulations with an Internet scale. In this paper we utilize parallel discrete event simulation simulation techniques for executing large scale simulation of P2P systems which scale effectively, only limited by the amount of computational resource available (memory and CPU). We show results from a number of P2P protocols, indicating good scalability both in terms of size (memory) and execution time (CPU). The results demonstrate how the differences of these protocols and which of the underlying factors affect the performance of the distributed simulation infrastructure.
Tien Tuan Anh Dinh, Georgios Theodoropoulos 0001, Rob Minson
DS-RT2
2008 Load Skew in Cell-Based Interest Management Systems
abstract
In large, real-time interactive distributed systems such as distributed simulations and multiplayer games, interest management (IM) is often implemented using a cell-based paradigm. In such a paradigm the subscription patterns of interactive clients are mapped on to some set of disjoint regions or cells which are typically hosted within a routing network made up either of dedicated machines or of the clients themselves. These systems often incorporate some mechanism for balancing the load placed on this routing network, on the assumption that interests over this population of cells will be non-uniform. Using a set of reference models for cell-based IM systems found in the research corpus, we evaluate the extent to which this phenomenon takes place. We also evaluate what effects an adaptive algorithm from previous work by Minson, R. and Theodoropoulos, G. (2007) has on this phenomenon.
Rob Minson, Georgios Theodoropoulos 0001
DS-RT2
2008 Push-Pull Interest Management for Virtual Worlds
abstract
Several approaches for scalable interest management (IM) within real-time distributed virtual environments (DVEs) have been proposed based upon some division of the data-space in to disjoint volumes or cells. Any such approach, however, must implement some mechanism for propagating the query and update messages around the distributed system. The efficiency of this process can greatly effect the scalability of such systems. In this paper we evaluate an adaptive approach to this problem.
Rob Minson, Georgios Theodoropoulos 0001
ISORC2
2008 Large Scale Distributed Simulation of p2p Networks
abstract
P2P systems have witnessed phenomenal development in recent years, evaluating and analysing new and existing new algorithms and techniques is a key issue for developers of p2p systems. In this context Simulation is an important tool for p2p developers. However, such systems are often very large and few existing simulators offer the ability to execute systems of real world size. In this paper we present a tool for executing large scale simulation of p2p systems which scale effectively, only limited by the amount of computational resource available (memory and CPU). This is achieved through the application of parallel discrete event simulation techniques to an existing, already scalable simulator, peersim. We show results from a case study using the chord p2p protocol, indicating good scalability both in terms of size (memory) and execution time (CPU).
Tien Tuan Anh Dinh, Michael Lees, Georgios Theodoropoulos 0001, Rob Minson
PDP3
2008 Distributing RePast agent-based simulations with HLA
abstract
Abstract Large, experimental multi‐agent system (MAS) simulations are highly demanding tasks, both computationally and developmentally. Agent toolkits provide reliable templates for the design of even the largest MAS simulations, without offering a solution to computational limitations. Conversely, distributed simulation architectures offer performance benefits, but the introduction of parallel logic can complicate the design process significantly. The motivations of distribution are not limited to this question of processing power. True interoperation of sequential agent‐simulation platforms would allow agents designed using different toolkits to transparently interact in common abstract domains. This paper discusses the design and implementation of a system capable of harnessing the computational power of a distributed simulation infrastructure with the design efficiency of an agent toolkit. The system permits integration, through a higher‐level architecture (HLA) federation, of multiple instances of the Java‐based lightweight agent‐simulation toolkit RePast. This paper defines abstractly the engineering process necessary in creating such middleware, and reports on the experience in the specific case of the RePast toolkit. The paper also presents performance results that illustrate that significant speedup can be achieved through the integration of RePast with HLA. Copyright © 2008 John Wiley & Sons, Ltd.
Rob Minson, Georgios Theodoropoulos 0001
Concurr. Comput. Pract. Exp.2
2008 Large scale agent-based simulation on the grid
Dan Chen 0001, Georgios Theodoropoulos 0001, Stephen John Turner, Wentong Cai 0001, Rob Minson, Yi Zhang 0004
Future Gener. Comput. Syst.2
2008 Data access in distributed simulations of multi-agent systems
Dan Chen 0001, Roland Ewald, Georgios Theodoropoulos 0001, Rob Minson, Ton Oguara, Michael Lees, Brian Logan 0001, Adelinde M. Uhrmacher
J. Syst. Softw.3
2007 An Evaluation of Push-Pull Algorithms in Support of Cell-Based Interest Management
abstract
Several approaches for scalable interest management (IM) within real-time distributed virtual environments (DVEs) have been proposed based upon some division of the data-space in to disjoint volumes or cells. Responsibility for the entire space can then be distributed. Any such approach, however, must implement some mechanism for propagating the query and update messages around the distributed system. The efficiency of this process can greatly effect the scalability of such systems. In this paper we evaluate an adaptive approach to this problem, designed for use in cell-based systems in general.
Rob Minson, Georgios Theodoropoulos 0001
DS-RT2
2006 Large Scale Distributed Simulation on the Grid
Georgios Theodoropoulos 0001, Yi Zhang 0004, Dan Chen 0001, Rob Minson, Stephen John Turner, Wentong Cai 0001, Brian Logan 0001
CCGRID1
2006 A Simulation Approach to Facilitate Parallel and Distributed Discrete-Event Simulator Development
abstract
Efficiently simulating discrete-event models in a parallel and distributed manner is a challenging endeavour. On one hand, various factors, such as hardware infrastructure or model characteristics, have to be considered. On the other hand, there is a wide variety of algorithms which address subproblems of parallel and distributed simulation and whose performance depends on the application at hand. We illustrate the resulting difficulties with respect to the development of parallel and distributed simulation systems and argue that the simulation of distributed simulation systems is a feasible approach to alleviate them. To underpin this, we introduce SIMSIM, a sequential simulator for parallel and distributed simulation systems. SIMSIM's pertinency is illustrated by the development of a load balancing algorithm for PDEVS. The algorithm's performance is analysed using SIMSIM and the predicted performance is compared to the performance of its implementation in the simulation system JAMES II
Roland Ewald, Jan Himmelspach, Adelinde M. Uhrmacher, Dan Chen 0001, Georgios Theodoropoulos 0001
DS-RT5
2006 Analysing Probabilistically Constrained Optimism
abstract
In previous work we presented the DTRD algorithm, an optimistic synchronisation algorithm for parallel discrete event simulation of multi-agent systems, and showed that it outperforms time warp and time windows on range of test cases. DTRD uses a decision theoretic model of rollback to derive an optimal time to delay read event so as to maximise the rate of LVT progression. The algorithm assumes that the inter-arrival times (both virtual and real) of events are normally distributed. In this paper we present a more detailed evaluation of the DTRD algorithm, and specifically how the performance of the algorithm is affected when the inter-arrival times do not follow the assumed distributions. Our analysis suggests that the performance of the algorithm is relatively insensitive to events whose inter-arrival times are not normally distributed. However as the variance of the input events increases its performance degrades to that of Time Warp. Our approach to evaluation is general, and we outline how the analysis may be applied to other decision theoretic algorithms
Michael Lees, Brian Logan 0001, Dan Chen 0001, Ton Oguara, Georgios Theodoropoulos 0001
DS-RT5
2006 Adaptive Interest Management via Push-Pull Algorithms
abstract
Interest management in large-scale distributed applications aims to reduce the amount of extraneous broadcast communication between nodes in the system with the aim of increasing responsiveness and scalability. We present a middleware-layer interest management framework based on pattern prediction to inform the oscillation of the system's protocol for the processing of state updates between two competing modes. This framework is transparent to the application itself. We discuss various algorithms for performing the prediction and experimentally evaluate the effectiveness of these algorithms against each other and a set of optimal and sub-optimal baselines
Rob Minson, Georgios Theodoropoulos 0001
DS-RT2
2005 Decision-Theoretic Throttling for Optimistic Simulations of Multi-Agent Systems
abstract
In this paper we present a throttling mechanism for optimistic simulations of multi-agent systems, which delays read accesses to the shared simulation state that are likely to be rolled back. We develop a decision-theoretic model of rollback and show how this can be used to derive the optimal time to delay a read event so as to minimize the expected overall execution time of the simulation. We briefly describe an implementation of this approach in ASSK, a distributed simulation kernel developed to investigate synchronization mechanisms for MAS simulation, and report the results of preliminary experiments to evaluate the effectiveness of our approach.
Michael Lees, Brian Logan 0001, Dan Chen 0001, Ton Oguara, Georgios Theodoropoulos 0001
DS-RT5
2005 An Adaptive Load Management Mechanism for Distributed Simulation of Multi-agent Systems
abstract
The paper presents a load management mechanism for distributed simulations of multi-agent systems. The mechanism minimizes the cost of accessing the shared state in the distributed simulation by dynamically redistributing shared state variables according to the access pattern of the simulation model. To evaluate the effectiveness and performance of the mechanism, a series of benchmark experiments were performed using the PDES-MAS framework for distributed simulation of multi-agent systems. Although preliminary, the results indicate that the proposed mechanism significantly reduces the overall access cost of the system.
Ton Oguara, Dan Chen 0001, Georgios Theodoropoulos 0001, Brian Logan 0001, Michael Lees
DS-RT3
2005 Revisiting Distributed Simulation and the Grid: A Panel
abstract
Summary form only given. The grid, or grid computing, provides a new and unrivalled technology for large scale distributed simulation as it enables collaboration and the use of distributed computing resources. Last year at DS-RT 2004 a panel was convened to consider the impact of the grid on distributed simulation. Four members presented their views of this area and together they tried to identify the main research issues involved in applying grid technology to distributed simulation and the key future challenges that need to be solved to achieve this goal. These challenges included not only technical ones, but also social ones such as management methodology and the development of standards. This year we revisit this fast changing technology and ask the questions: 1) What major changes has grid computing had over the past year?; 2) How can distributed simulation and related applications benefit from grid computing?; 3) What are the barriers to this?; 4) Are there alternative new technologies that would be a better ROI for distributed simulation than grid computing?.
Simon J. E. Taylor, Geoffrey C. Fox, Richard M. Fujimoto, J. Mark Pullen, David J. Roberts 0001, Georgios Theodoropoulos 0001
DS-RT6
2004 Distributed Simulation of MAS
Michael Lees, Brian Logan 0001, Rob Minson, Ton Oguara, Georgios Theodoropoulos 0001
MABS5
2002 Distributed Simulation of Asynchronous Hardware: The Program Driven Synchronization Protocol
Georgios Theodoropoulos 0001
J. Parallel Distributed Comput.1
2001 Simulating asynchronous hardware on multiprocessor platforms: the case of AMULET1
abstract
Abstract Synchronous VLSI design is approaching a critical point, with clock distribution becoming an increasingly costly and complicated issue and power consumption rapidly emerging as a major concern. Hence, recently, there has been a resurgence of interest in asynchronous digital design techniques which promise to liberate digital design from the inherent problems of synchronous systems. This activity has revealed a need for modelling and simulation techniques suitable for the asynchronous design style. The concurrent process algebra communicating sequential processes (CSP) and its executable counterpart, occam, are increasingly advocated as particularly suitable for this purpose. This paper focuses on issues related to the execution of CSP/occam models of asynchronous hardware on multiprocessor machines, including I/O, monitoring and debugging, partition, mapping and load balancing. These issues are addressed in the context of occarm, an occam simulation model of the AMULET1 asynchronous microprocessor; however, the solutions devised are more general and may be applied to other systems too. Copyright © 2001 John Wiley & Sons, Ltd.
Georgios Theodoropoulos 0001
Concurr. Comput. Pract. Exp.1
2001 The distributed simulation of multiagent systems
abstract
Agent based systems are increasingly being applied in a wide range of areas including telecommunications, business process modeling, computer games, control of mobile robots, and military simulations. Such systems are typically extremely complex and it is often useful to be able to simulate an agent based system to learn more about its behavior or investigate the implications of alternative architectures. The authors discuss the application of distributed discrete event simulation techniques to the simulation of multiagent systems. We identify the efficient distribution of the agents' environment as a key problem in the simulation of agent based systems and present an approach to the decomposition of the environment that facilitates load balancing.
Brian Logan 0001, Georgios Theodoropoulos 0001
Proc. IEEE2