VLDB 2026 Research / reviewers in the wild / expert
Vasileios Argyriou
dblp:32/487
· DBLP profile ↗
74ranked-venue papers
14as first author
26since 2021 · last 2026
0000-0003-4679-8049ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 39 · 10 first-author · 8 since 2021Artificial intelligence and machine learning · 31 · 7 first-author · 11 since 2021Computer networks · 6 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Systems, architecture and hardware · 2 · 2 since 2021Security and privacy · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Training-Free Multi-Concept LoRA Composition with Prompt-Aware Weighting
Georgios Tsoumplekas, Stella Bounareli, Vasileios Argyriou |
FG | 3 |
| 2026 | TiC-MuFormer: Time-Aware Caption-Integrated Multimodal Transformers for User-Level Mental Health Modeling
Georgios Tsoumplekas, Yannis Spyridis, Vasileios Argyriou |
LREC | 3 |
| 2026 | Visual hand gesture recognition with deep learning: A comprehensive review of methods, datasets, challenges and future research directionsabstractThe rapid evolution of deep learning (DL) models and the ever-increasing size of available datasets have raised the interest of the research community in the always-important field of visual hand gesture recognition (VHGR), and delivered a wide range of applications, such as sign language understanding and human-computer interaction using cameras. Despite the large volume of research works in the field, a structured and complete survey on VHGR is still missing, leaving researchers to navigate through hundreds of papers in order to find the current state-of-the-art (SOTA). The current survey aims to fill this gap by presenting a comprehensive overview of this computer vision field. With a systematic research methodology that identifies the SOTA works and a structured presentation of the various methods, datasets, and evaluation metrics, this review aims to constitute a useful guideline for researchers, helping them to propose improvements. Specifically, this survey focuses on four fundamental questions: what are the main VHGR aspects, what are the current SOTA methods, what comparative insights can be drawn across methods and tasks, and which challenges shape future research. Starting with the methodology used to locate the related literature, the survey identifies and organizes the key VHGR approaches in a taxonomy-based format, and presents the various dimensions that affect the final method choice, such as input modality, task type, and application domain. The SOTA techniques are grouped across three primary VHGR tasks: static, isolated dynamic and continuous gesture recognition. For each task, the architectural trends and learning strategies are listed. To support the experimental evaluation of future methods in the field, the study reviews commonly used datasets and presents the standard performance metrics. Our survey concludes by identifying the major challenges in VHGR, including both general computer vision issues and domain-specific obstacles, and outlines promising directions for future research. Konstantinos Foteinos, Manousos Linardakis, Panagiotis I. Radoglou-Grammatikis, Vasileios Argyriou, Panagiotis G. Sarigiannidis, Iraklis Varlamis, Georgios Th. Papadopoulos |
Neurocomputing | 4 |
| 2026 | Introducing Energy Efficient Routing in UAV-Satellite NTNs for Dynamic 6G InterconnectivityabstractThe integration of Unmanned Aerial Vehicles (UAVs) and Low-Earth Orbit (LEO) satellites as aerial nodes in non-terrestrial networks (NTNs) presents both opportunities and challenges for on-demand 6G interconnectivity. This paper presents a new Composite Cost Metric (CCM) which improves energy-efficient routing performance in combined UAV-satellite constellations. We consider incorporating cumulative Free Space Path Loss (FSPL) and residual energy into the route selection process for both proactive and reactive protocols, our approach refines the routing decisions of classical protocols. The proposed CCM-driven modifications and protocol-specific integration typologies can improve overall route stability, reduce energy consumption per delivered packet, and optimize network reliability by dynamically selecting relays with lower attenuation and higher energy availability. We develop an NS-3-based simulation framework that integrates realistic satellite orbital mechanics, UAV mobility models, and a hybrid energy model that includes solar energy harvesting for satellites. Simulation results demonstrate that our enhancements can indeed outperform baseline implementations in packet delivery ratio, energy efficiency, and end-to-end delay which makes them viable for next-generation NTN-supported 6G networks, at the expense of some additional control overhead. With this set of developments we aim to pave the way for global-optimum and energy-aware emergency and disaster relief communications. George Amponis, Thomas Lagkas, Pavlos S. Bouzinis, Panagiotis I. Radoglou-Grammatikis, Antonios Sarigiannidis, Panagiotis G. Sarigiannidis, Vasileios Argyriou |
IEEE Trans. Commun. | 7 |
| 2025 | DiffusionAct: Controllable Diffusion Autoencoder for One-shot Face ReenactmentabstractVideo-driven neural face reenactment aims to synthesize realistic facial images that successfully preserve the identity and appearance of a source face, while transferring the target head pose and facial expressions. Existing GAN-based methods suffer from either distortions and visual artifacts or poor reconstruction quality, i.e., the background and several important appearance details, such as hair style/color, glasses and accessories, are not faithfully reconstructed. Recent advances in Diffusion Probabilistic Models (DPMs) enable the generation of high-quality realistic images. To this end, in this paper we present DiffusionAct, a novel method that leverages the photo-realistic image generation of diffusion models to perform neural face reenactment. Specifically, we propose to control the semantic space of a Diffusion Autoencoder (DiffAE), in order to edit the facial pose of the input images, defined as the head pose orientation and the facial expressions. Our method allows one-shot, self, and cross-subject reenactment, without requiring subject-specific fine-tuning. We compare against state-of-the-art GAN-, StyleGAN2-, and diffusion-based methods, showing better or on-par reenactment performance. Project page: https://stelabou.github.io/diffusionact/ Stella Bounareli, Christos Tzelepis, Vasileios Argyriou, Ioannis Patras, Georgios Tzimiropoulos |
FG | 3 |
| 2025 | Malware Detection in Docker Containers: An Image is Worth a Thousand LogsabstractMalware detection is increasingly challenged by evolving techniques like obfuscation and polymorphism, limiting the effectiveness of traditional methods. Meanwhile, the widespread adoption of software containers has introduced new security challenges, including the growing threat of malicious software injection, where a container, once compromised, can serve as entry point for further cyberattacks. In this work, we address these security issues by introducing a method to identify compromised containers through machine learning analysis of their file systems. We cast the entire software containers into large RGB images via their tarball representations, and propose to use established Convolutional Neural Network architectures on a streaming, patchbased manner. To support our experiments, we release the COSOCO dataset-the first of its kind-containing 3364 largescale RGB images of benign and compromised software containers at https://huggingface.co/datasets/k3ylabs/cosoco-imagedataset. Our method detects more malware and achieves higher F1 and Recall scores than all individual and ensembles of VirusTotal engines, demonstrating its effectiveness and setting a new standard for identifying malware-compromised software containers. Akis Nousias, Efklidis Katsaros, Evangelos Syrmos, Panagiotis I. Radoglou-Grammatikis, Thomas Lagkas, Vasileios Argyriou, Ioannis D. Moscholios, Evangelos Markakis 0002, Sotirios K. Goudos, Panagiotis G. Sarigiannidis |
ICC | 6 |
| 2025 | Fusion Grad-CAM: A Methodology for Generating Unified Attention Maps in Majority Voting ClassifiersabstractIn this work, we introduce a novel methodology called Fusion Grad-CAM for generating attention maps in ensemble classification systems that utilize majority voting. Currently, there is no established technique for producing a consolidated attention map in such systems. Traditional approaches require inspecting individual attention maps from each contributing model, which is inefficient. Our proposed method integrates Grad-CAM produced attention maps by averaging or weighted averaging them based on both the predicted class and the confidence level of each model’s decision. This approach offers a holistic representation of the ensemble’s decision-making process. We validate our methodology using various combinations of well-known convolutional neural networks (CNNs) pretrained on ImageNet, demonstrating the effectiveness and clarity of the resulting attention maps. Nikolaos Ntampakis, Konstantinos I. Diamantaras, Vasileios Argyriou, Panagiotis Sarigianndis |
IPAS | 3 |
| 2025 | NeuroXVocal: Detection and Explanation of Alzheimer's Disease Through Non-Invasive Analysis of Picture-Prompted Speech
Nikolaos Ntampakis, Konstantinos I. Diamantaras, Ioanna Chouvarda, Magda Tsolaki, Panagiotis Sarigianndis, Vasileios Argyriou |
MICCAI (14) | 6 |
| 2025 | Enhancing 3D object detection in autonomous vehicles based on synthetic virtual environment analysisabstractAutonomous Vehicles (AVs) rely on real-time processing of natural images and videos for scene understanding and safety assurance through proactive object detection. Traditional methods have primarily focused on 2D object detection, limiting their spatial understanding. This study introduces a novel approach by leveraging 3D object detection in conjunction with augmented reality (AR) ecosystems for enhanced real-time scene analysis. Our approach pioneers the integration of a synthetic dataset, designed to simulate various environmental, lighting, and spatiotemporal conditions, to train and evaluate an AI model capable of deducing 3D bounding boxes. This dataset, with its diverse weather conditions and varying camera settings, allows us to explore detection performance in highly challenging scenarios. The proposed method also significantly improves processing times while maintaining accuracy, offering competitive results in conditions previously considered difficult for object recognition. The combination of 3D detection within the AR framework and the use of synthetic data to tackle environmental complexity marks a notable contribution to the field of AV scene analysis. • A multimodal architecture for real-time 3D object detection in AV systems. • Efficient 3D bounding box prediction extrapolated from 2D images. • Novel synthetic dataset simulates diverse environmental conditions for AVs. • Comparative evaluation against state-of-the-art techniques for object detection. Vladislav Li, Ilias Siniosoglou, Thomai Karamitsou, Anastasios Lytos, Ioannis D. Moscholios, Sotirios K. Goudos, Jyoti S. Banerjee, Panagiotis G. Sarigiannidis, Vasileios Argyriou |
Image Vis. Comput. | 9 |
| 2025 | Is it worth the energy? An in-depth study on the energy efficiency of data augmentation strategies for finetuning-based low/few-shot object detectionabstractCurrent methods for low- and few-shot object detection have primarily focused on enhancing model performance for detecting objects. One common approach to achieve this is by combining model finetuning with data augmentation strategies. However, little attention has been given to the energy efficiency of these approaches in data-scarce regimes. This paper seeks to conduct a comprehensive empirical study that examines both model performance and energy efficiency of custom data augmentations and automated data augmentation selection strategies when combined with a lightweight object detector. The methods are evaluated in four different benchmark datasets in terms of their performance and energy consumption, providing valuable insights regarding reaching an optimal tradeoff between these two objectives. Additionally, to better quantify this tradeoff, we propose a novel metric named modified Efficiency Factor that combines both of these conflicting objectives in a single metric and thus enables gaining insights into the effectiveness of the examined models and data augmentation strategies when considering both performance and efficiency. Consequently, it is shown that while some broader guidelines regarding appropriate data augmentation selections can be provided based on the obtained performance and energy efficiency results, in many cases, the performance gains of data augmentation strategies are overshadowed by their increased energy usage, necessitating the development of more energy-efficient data augmentation strategies to address data scarcity. Vladislav Li, Georgios Tsoumplekas, Ilias Siniosoglou, Panagiotis G. Sarigiannidis, Vasileios Argyriou |
J. Syst. Archit. | 5 |
| 2025 | StatAvg: Mitigating Data Heterogeneity in Federated Learning for Intrusion Detection SystemsabstractFederated learning (FL) enables devices to collaboratively build a shared machine learning (ML) or deep learning (DL) model without exposing raw data. Its privacy-preserving nature has made it popular for intrusion detection systems (IDS) in the field of cybersecurity. However, data heterogeneity across participants poses challenges for FL-based IDS. This paper proposes statistical averaging (StatAvg) method to alleviate non-independently and identically (non-iid) distributed features across local clients’ data in FL. In particular, StatAvg allows the FL clients to share their individual local data statistics with the server. These statistics include the mean and variance of each client’s feature vector. The server then aggregates this information to produce global statistics, which are shared with the clients and used for universal data normalization, i.e., common scaling of the input features by all clients. It is worth mentioning that StatAvg can seamlessly integrate with any FL aggregation strategy, as it occurs before the actual FL training process. The proposed method is evaluated against well-known baseline approaches that rely on batch and layer normalization, such as FedBN, and address the non-iid features issue in FL. Experiments were conducted using the TON-IoT and CIC-IoT-2023 datasets, which are relevant to the design of host and network IDS, respectively. The experimental results demonstrate the efficiency of StatAvg in mitigating non-iid feature distributions across the FL clients compared to the baseline methods, offering a gain in IDS accuracy ranging from 4% to 17%. Pavlos S. Bouzinis, Panagiotis I. Radoglou-Grammatikis, Ioannis Makris, Thomas Lagkas, Vasileios Argyriou, Georgios Th. Papadopoulos, Panagiotis G. Sarigiannidis, George K. Karagiannidis |
IEEE Trans. Netw. Serv. Manag. | 5 |
| 2024 | AAG: Adversarial Attack Generator for evaluating the robustness of Machine Learning Models against Adversarial AttacksabstractWith the ongoing integration of machine learning models into critical infrastructure, the resilience of these systems against adversarial attacks is important for all domains. This paper introduces an adversarial attack generator framework against a network dataset that is part of OCPP Dataset using CI-CFlowMeter parser. We conduct a comprehensive evaluation of various prominent adversarial attacks, including FGSMA, JSMA, PGD, C&W, and more to assess their efficacy on the OCCP dataset. The Adversarial Generator is meticulously evaluated, demonstrating a significant impact in the models performance to detect potential perturbations. The results showcased the impact of the different type of adversarial attacks, contributing to a critical advancement in future defense strategies that need to be utilised in order to protect industrial control systems. Dimitrios Christos Asimopoulos, Panagiotis I. Radoglou-Grammatikis, Thomas Lagkas, Vasileios Argyriou, Ioannis D. Moscholios, Jorgen Cani, Georgios Th. Papadopoulos, Evangelos Markakis 0002, Panagiotis G. Sarigiannidis |
IEEE Big Data | 4 |
| 2024 | One-Shot Neural Face Reenactment via Finding Directions in GAN's Latent SpaceabstractAbstract In this paper, we present our framework for neural face/head reenactment whose goal is to transfer the 3D head orientation and expression of a target face to a source face. Previous methods focus on learning embedding networks for identity and head pose/expression disentanglement which proves to be a rather hard task, degrading the quality of the generated images. We take a different approach, bypassing the training of such networks, by using (fine-tuned) pre-trained GANs which have been shown capable of producing high-quality facial images. Because GANs are characterized by weak controllability, the core of our approach is a method to discover which directions in latent GAN space are responsible for controlling head pose and expression variations. We present a simple pipeline to learn such directions with the aid of a 3D shape model which, by construction, inherently captures disentangled directions for head pose, identity, and expression. Moreover, we show that by embedding real images in the GAN latent space, our method can be successfully used for the reenactment of real-world faces. Our method features several favorable properties including using a single source image (one-shot) and enabling cross-person reenactment. Extensive qualitative and quantitative results show that our approach typically produces reenacted faces of notably higher quality than those produced by state-of-the-art methods for the standard benchmarks of VoxCeleb1 & 2. Stella Bounareli, Christos Tzelepis, Vasileios Argyriou, Ioannis Patras, Georgios Tzimiropoulos |
Int. J. Comput. Vis. | 3 |
| 2023 | Surveying Cyber Threat Intelligence and Collaboration: A Concise Analysis of Current Landscape and TrendsabstractThe evolution of cyberattacks has been significantly impacted by the rise of Artificial Intelligence (AI). In particular, AI-driven attacks leverage Machine Learning (ML) and Deep Learning (DL) methods to automate tasks like identifying vulnerabilities, crafting convincing phishing emails, and evading conventional security measures. These cyberattacks can adapt in real time, making them more elusive and challenging to detect. Furthermore, AI has enabled the development of AI-powered malware that can learn and evolve, making it even more dangerous. As AI continues to evolve, both attackers and defenders are engaged in a relentless arms race, with cybersecurity professionals striving to harness AI for threat detection and response while cybercriminals seek to exploit AI’s capabilities for their malicious purposes. This ongoing battle underscores the need for proactive and adaptive cybersecurity strategies to mitigate the evolving threats posed by AI-driven cyberattacks. Based on the aforementioned remarks, it is evident that efficient and adaptable countermeasures are necessary. In this paper, we focus our attention on Cyber Threat Intelligence (CTI) mechanisms. CTI is the process of collecting, analysing, and sharing information about potential cybersecurity threats to help organisations proactively defend against cyberattacks. In particular, after providing an overview of the CTI use cases, a brief analysis of existing solutions follows, highlighting the current trends and directions for future work in this research field. Panagiotis I. Radoglou-Grammatikis, Elisavet Kioseoglou, Dimitrios Christos Asimopoulos, Miltiadis G. Siavvas, Ioannis Nanos, Thomas Lagkas, Vasileios Argyriou, Kostas E. Psannis, Sotirios K. Goudos, Panagiotis G. Sarigiannidis |
CloudCom | 7 |
| 2023 | StyleMask: Disentangling the Style Space of StyleGAN2 for Neural Face ReenactmentabstractIn this paper we address the problem of neural face reenactment, where, given a pair of a source and a target facial image, we need to transfer the target's pose (defined as the head pose and its facial expressions) to the source image, by preserving at the same time the source's identity characteristics (e.g., facial shape, hair style, etc), even in the challenging case where the source and the target faces belong to different identities. In doing so, we address some of the limitations of the state-of-the-art works, namely, a) that they depend on paired training data (i.e., source and target faces have the same identity), b) that they rely on labeled data during inference, and c) that they do not preserve identity in large head pose changes. More specifically, we propose a framework that, using unpaired randomly generated facial images, learns to disentangle the identity characteristics of the face from its pose by incorporating the recently introduced style space S [1] of StyleGAN2 [2], a latent representation space that exhibits remarkable disentanglement properties. By capitalizing on this, we learn to successfully mix a pair of source and target style codes using supervision from a 3D model. The resulting latent code, that is subsequently used for reenactment, consists of latent units corresponding to the facial pose of the target only and of units corresponding to the identity of the source only, leading to notable improvement in the reenactment performance compared to recent state-of-the-art methods. In comparison to state of the art, we quantitatively and qualitatively show that the proposed method produces higher quality results even on extreme pose variations. Finally, we report results on real images by first embedding them on the latent space of the pretrained generator. We make the code and the pretrained models publicly available at: https://github.com/StelaBou/StyleMask. Stella Bounareli, Christos Tzelepis, Vasileios Argyriou, Ioannis Patras, Georgios Tzimiropoulos |
FG | 3 |
| 2023 | HyperReenact: One-Shot Reenactment via Jointly Learning to Refine and Retarget FacesabstractIn this paper, we present our method for neural face reenactment, called HyperReenact, that aims to generate realistic talking head images of a source identity, driven by a target facial pose. Existing state-of-the-art face reenactment methods train controllable generative models that learn to synthesize realistic facial images, yet producing reenacted faces that are prone to significant visual artifacts, especially under the challenging condition of extreme head pose changes, or requiring expensive few-shot fine-tuning to better preserve the source identity characteristics. We propose to address these limitations by leveraging the photorealistic generation ability and the disentangled properties of a pretrained StyleGAN2 generator, by first inverting the real images into its latent space and then using a hypernetwork to perform: (i) refinement of the source identity characteristics and (ii) facial pose re-targeting, eliminating this way the dependence on external editing methods that typically produce artifacts. Our method operates under the one-shot setting (i.e., using a single source frame) and allows for cross-subject reenactment, without requiring any subject-specific fine-tuning. We compare our method both quantitatively and qualitatively against several state-of-the-art techniques on the standard benchmarks of VoxCeleb1 and VoxCeleb2, demonstrating the superiority of our approach in producing artifact-free images, exhibiting remarkable robustness even under extreme head pose changes. We make the code and the pretrained models publicly available at: https://github.com/StelaBou/HyperReenact. Stella Bounareli, Christos Tzelepis, Vasileios Argyriou, Ioannis Patras, Georgios Tzimiropoulos |
ICCV | 3 |
| 2023 | Post-Processing Fairness Evaluation of Federated Models: An Unsupervised Approach in HealthcareabstractModern Healthcare cyberphysical systems have begun to rely more and more on distributed AI leveraging the power of Federated Learning (FL). Its ability to train Machine Learning (ML) and Deep Learning (DL) models for the wide variety of medical fields, while at the same time fortifying the privacy of the sensitive information that are present in the medical sector, makes the FL technology a necessary tool in modern health and medical systems. Unfortunately, due to the polymorphy of distributed data and the shortcomings of distributed learning, the local training of Federated models sometimes proves inadequate and thus negatively imposes the federated learning optimization process and in extend in the subsequent performance of the rest Federated models. Badly trained models can cause dire implications in the healthcare field due to their critical nature. This work strives to solve this problem by applying a post-processing pipeline to models used by FL. In particular, the proposed work ranks the model by finding how fair they are by discovering and inspecting micro-Manifolds that cluster each neural model's latent knowledge. The produced work applies a completely unsupervised both model and data agnostic methodology that can be leveraged for general model fairness discovery. The proposed methodology is tested against a variety of benchmark DL architectures and in the FL environment, showing an average 8.75% increase in Federated model accuracy in comparison with similar work. Ilias Siniosoglou, Vasileios Argyriou, Panagiotis G. Sarigiannidis, Thomas Lagkas, Antonios Sarigiannidis, Sotirios K. Goudos, Shaohua Wan 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2022 | Finding Directions in GAN's Latent Space for Neural Face Reenactment
Stella Bounareli, Vasileios Argyriou, Georgios Tzimiropoulos |
BMVC | 2 |
| 2022 | False Data Injection Attacks against Low Voltage Distribution SystemsabstractThe transformation of the conventional electrical grid into a digital ecosystem brings significant benefits, such as two-way communication between energy consumers and utilities, self-monitoring and pervasive controls. However, the advent of the smart electrical grid raises severe cybersecurity and privacy concerns, given the presence of legacy systems and communications protocols. This paper focuses on False Data Injection (FDI) cyberattacks against a low-voltage distribution system, taking full advantage of Man In The Middle (MITM) actions. The first cyberattack targets the communication between a smart meter and an Active Distribution Management System (ADMS), while the second FDI cyberattack targets the communication between a smart inverter and ADMS. In both cases, the cyberattacks affect the operation of the distribution transformer, thus resulting in devastating consequences. Moreover, this paper provides an Artificial Intelligence (AI)-based Intrusion Detection System (IDS), detecting and mitigating the above cyberattacks in a timely manner. The evaluation results demonstrate the efficiency of the proposed IDS. Panagiotis I. Radoglou-Grammatikis, Christos Dalamagkas, Thomas Lagkas, Magda Zafeiropoulou, Maria Atanasova, Pencho Zlatev, Alexandros-Apostolos A. Boulogeorgos, Vasileios Argyriou, Evangelos Markakis 0002, Ioannis D. Moscholios, Panagiotis G. Sarigiannidis |
GLOBECOM | 8 |
| 2022 | Dynamic Risk Assessment and Certification in the Power Grid: A Collaborative ApproachabstractThe digitisation of the typical electrical grid introduces valuable services, such as pervasive control, remote monitoring and self-healing. However, despite the benefits, cybersecurity and privacy issues can result in devastating effects or even fatal accidents, given the interdependence between the energy sector and other critical infrastructures. Large-scale cyber attacks, such as Indostroyer and DragonFly have already demonstrated the weaknesses of the current electrical grid with disastrous consequences. Based on the aforementioned remarks, both academia and industry have already designed various cybersecurity standards, such as IEC 62351. However, dynamic risk assessment and certification remain crucial aspects, given the sensitive nature of the electrical grid. On the one hand, dynamic risk assessment intends to re-compute the risk value of the affected assets and their relationships in a dynamic manner based on the relevant security events and alarms. On the other hand, based on the certification process, new approach for the dynamic management of the security need to be defined in order to provide adaptive reaction to new threats. This paper presents a combined approach, showing how both aspects can be applied in a collaborative manner in the smart electrical grid. Thanasis Liatifis, Pedro Ruzafa Alcazar, Panagiotis I. Radoglou-Grammatikis, Dimitrios Papamartzivanos, Sofia-Anna Menesidou, Thomas Krousarlis, Alberto Molinuevo Martín, Iñaki Angulo, Antonios Sarigiannidis, Thomas Lagkas, Vasileios Argyriou, Antonio F. Skarmeta, Panagiotis G. Sarigiannidis |
NetSoft | 11 |
| 2022 | Modeling, Detecting, and Mitigating Threats Against Industrial Healthcare Systems: A Combined Software Defined Networking and Reinforcement Learning ApproachabstractThe rise of the Internet of Medical Things introduces the healthcare ecosystem in a new digital era with multiple benefits, such as remote medical assistance, realtime monitoring, and pervasive control.However, despite the valuable healthcare services, this progression raises significant cybersecurity and privacy concerns.In this article, we focus our attention on the IEC 60 870-5-104 protocol, which is widely adopted in industrial healthcare systems.First, we investigate and assess the severity of the IEC 60 870-5-104 cyberattacks by providing a quantitative threat model, which relies on Attack Defence Trees and Common Vulnerability Scoring System v3.1.Next, we introduce an intrusion detection and prevention system (IDPS), which is capable of discriminating and mitigating automatically the IEC 60 870-5-104 cyberattacks.The proposed IDPS takes full advantage of the machine learning (ML) and software defined networking (SDN) technologies.ML is used to detect the IEC 60 870-5-104 cyberattacks, utilizing 1) Transmission Control Protocol/Internet Protocol network flow statistics and 2) IEC 60 870-5-104 payload flow statistics. Panagiotis I. Radoglou-Grammatikis, Konstantinos Rompolos, Panagiotis G. Sarigiannidis, Vasileios Argyriou, Thomas Lagkas, Antonios Sarigiannidis, Sotirios K. Goudos, Shaohua Wan 0001 |
IEEE Trans. Ind. Informatics | 4 |
| 2021 | Unsupervised Ethical Equity Evaluation of Adversarial Federated NetworksabstractWhile the technology of Deep Learning (DL) is a powerful tool when properly trained for image analysis and classification applications, some factors for its optimization rely solely on the training data and their environment. In an effort to tackle the problem of knowledge bias created during the training process of a Deep Neural Network (DNN) and specifically Adversarial Networks for image augmentation, this work presents an entirely unsupervised methodology for discovering the unfairness level of Deep Learning (DL) models and in extend, its wrongly accumulated or biased classes. Fdi, the proposed evaluation metric for quantizing the level of unfairness of a model is introduced, along with the method of weighting the model’s knowledge and producing its weakest aspects in a data-agnostic way. Ilias Siniosoglou, Vasileios Argyriou, Stamatia Bibi, Thomas Lagkas, Panagiotis G. Sarigiannidis |
ARES | 2 |
| 2021 | Similarity Model using Gradient Images to Compare Human and AI AgentsabstractThe objective is to determine how close an Artificial Intelligence agent is, in comparison to a human player, using only game play images. Identifying Artificial Intelligence agents during game play is typically done through the analysis and collection of bio-metric data, such as keyboard, mouse and other controller interfaces. This document presents a model of an Auto Encoder architecture, with Long Short Term Memory layers. Gradient and Non-Augmented Images have been evaluated and compared. Three distinct personalities of agents are evaluated, human players, a trained Artificial Intelligence agent, and a basic Artificial Intelligence agent. Through testing the Gradient Image augmentation sets present promising results, with the model successfully identifying the human as the closest similarity to the baseline, followed by the trained Artificial Intelligence agent. Gordon Johnson, Vasileios Argyriou, Christos Politis |
DCOSS | 2 |
| 2021 | Fourier Transformation Autoencoders for Anomaly DetectionabstractAnomaly detection is a challenging problem, mainly due to the lack of a sufficient set of abnormal samples that represents every possible anomaly. Therefore unsupervised methods are employed to model normality and anomaly is detected as an outlier to such models. This paper introduces Fourier Trans-forms into AutoEncoders to demonstrate how the inclusion of a frequency domain presents less noisy features for a deep learning network to detect anomalies. Comparing our results to the state of the art on a variety of datasets, we show how the proposed method can provide competitive results. Demetris Lappas, Vasileios Argyriou, Dimitrios Makris 0001 |
ICASSP | 2 |
| 2021 | A Cyber Resilience Framework for NG-IoT Healthcare Using Machine Learning and BlockchainabstractInternet of Things (IoT) technology such as intelligent devices, sensors, actuators and wearables have been integrated in the healthcare industry, thus contributing in the creation of smart hospitals and remote assistance environments. Ensuring the eHealth network adopts the appropriate security measures in order to effectively protect sensitive patient data against malicious attempts is a tough challenge. Devices composing eHealth infrastructure are considered to be easily exploitable. To that end, a solution monitoring the intelligent healthcare environment is of essence. In addition, by digitalising all health records, appropriate measures need to be implemented in order for patient records to be accessible by authorized personnel only. Furthermore, creating interoperable systems, capable of being integrated by multiple organizations such as hospitals and insurance companies, while maintaining a General Data Protection Regulation-friendly posture, providing access to health data is a great importance for optimal patient assistance. To address both concerns, we present a framework featuring a multi-layer tool for providing a highly effective security solution specifically designed to address the eHealth requirements, and a blockchain access control component, based on smart contracts to provide access control for authorized users to patient records and health data in a distributed way. Vasiliki Kelli, Panagiotis G. Sarigiannidis, Vasileios Argyriou, Thomas Lagkas, Vasileios Vitsas |
ICC | 3 |
| 2021 | Semi-Grant-Free Non-Orthogonal Multiple Access for Tactile Internet of ThingsabstractUltra-low latency connections for a massive number of devices are one of the main requirements of the next-generation tactile Internet-of-Things (TIoT). Grant-free non-orthogonal multiple access (GF-NOMA) is a novel paradigm that leverages the advantages of grant-free access and non-orthogonal transmissions, to deliver ultra-low latency connectivity. In this work, we present a joint channel assignment and power allocation solution for semi-GF-NOMA systems, which provides access to both grant-based (GB) and grant-free (GF) devices, maximizes the network throughput, and is capable of ensuring each device’s throughput requirements. In this direction, we provide the mathematical formulation of the aforementioned problem. After explaining that it is not convex, we propose a solution strategy based on the Lagrange multipliers and subgradient method. To evaluate the performance of our solution, we carry out system-level Monte Carlo simulations. The simulation results indicate that the proposed solution can optimize the total system throughput and achieve a high association rate, while taking into account the minimum throughput requirements of both GB and GF devices. Dimitrios Pliatsios, Alexandros-Apostolos A. Boulogeorgos, Thomas Lagkas, Vasileios Argyriou, Ioannis D. Moscholios, Panagiotis G. Sarigiannidis |
PIMRC | 4 |
| 2020 | Accurate Deep Net Crowd Counting for Smart IoT Video acquisition devicesabstractA novel deep neural network is proposed, for accurate and robust crowd counting. Crowd counting is a complex task, as it strongly depends on the deployed camera characteristics and, above all, the scene perspective. Crowd counting is essential in security applications where Internet of Things (IoT) cameras are deployed to help with crowd management tasks. The complexity of a scene varies greatly, and a medium to large scale security system based on IoT cameras must cater for changes in perspective and how people appear from different vantage points. To address this, our deep architecture extracts multi-scale features with a pyramid contextual module to provide long-range contextual information and enlarge the receptive field. Experiments were run on three major crowd counting datasets, to test our proposed method. Results demonstrate our method supersedes the performance of state-of-the-art methods. Anish R. Khadka, Vasileios Argyriou, Paolo Remagnino |
DCOSS | 2 |
| 2020 | Synthetic Crowd and Pedestrian Generator for Deep Learning ProblemsabstractDeep Neural networks (DNN) dominate the state of art results in computer vision (CV) and other fields. One of the primary reasons why DNN outperform existing algorithms is that these produce superior results when more labelled data are used, unlike classic CV techniques. Nonetheless, it is well known that DNN requires a very large amount of data to generalise well. Collecting and labelling these datasets are expensive, time-consuming and sometimes impossible. Therefore, researchers tried to use alternative techniques, such as graphics simulators to automatically generate labelled datasets. However, these techniques are still expensive and require domain knowledge to produce good datasets. In this paper, therefore, a graphics simulator is presented which automatically generates multi-model datasets in real-time providing the corresponding ground truth and annotation. The tool concentrates on pedestrian and crowd analysis including 3D human pose estimation, pedestrian detection as well as crowd density and flow estimation. Anish R. Khadka, Paolo Remagnino, Vasileios Argyriou |
ICASSP | 3 |
| 2020 | Latent Bernoulli AutoencoderabstractIn this work, we pose the question whether it is possible to design and train an autoencoder model in an end-to-end fashion to learn representations in the multivariate Bernoulli latent space, and achieve performance comparable with the state-of-the-art variational methods. Moreover, we investigate how to generate novel samples and perform smooth interpolation and attributes modification in the binary latent space. To meet our objective, we propose a simplified, deterministic model with a straight-through gradient estimator to learn the binary latents and show its competitiveness with the latest VAE methods. Furthermore, we propose a novel method based on a random hyperplane rounding for sampling and smooth interpolation in the latent space. Our method performs on a par or better than the current state-of-the-art methods on common CelebA, CIFAR-10 and MNIST datasets. Jiri Fajtl, Vasileios Argyriou, Dorothy Ndedi Monekosso, Paolo Remagnino |
ICML | 2 |
| 2020 | A modular CNN-based building detector for remote sensing imagesabstractConvolutional neural networks (CNNs) have resurged lately due to their state-of-the-art performance in various disciplines, such as computer vision , audio and text processing. However, CNNs have not been widely employed for remote sensing applications. In this paper, we propose a CNN architecture , named Modular-CNN, to improve the performance of building detectors that employ Histogram of Oriented Gradients (HOG) and Local Binary Patterns (LBP) in a remote sensing dataset. Additionally, we propose two improvements to increase the classification accuracy of Modular-CNN. The first improvement combines the power of raw and normalised features, while the second one concerns the Euler transformation of feature vectors. We demonstrate the effectiveness of our proposed Modular-CNN and the novel improvements in remote sensing and other datasets in a comparative study with other state-of-the-art methods. Dimitrios Konstantinidis, Vasileios Argyriou, Tania Stathaki, Nikolaos Grammalidis |
Comput. Networks | 2 |
| 2020 | Improving Dataset Volumes and Model Accuracy With Semi-Supervised Iterative Self-LearningabstractWithin this work a novel semi-supervised learning technique is introduced based on a simple iterative learning cycle together with learned thresholding techniques and an ensemble decision support system. State-of-the-art model performance and increased training data volume are demonstrated, through the use of unlabelled data when training deeply learned classification models. The methods presented work independently from the model architectures or loss functions, making this approach applicable to a wide range of machine learning and classification tasks. Evaluation of the proposed approach is performed on commonly used datasets when evaluating semi-supervised learning techniques as well as a number of more challenging image classification datasets (CIFAR-100 and a 200 class subset of ImageNet). Rob Dupre, Jiri Fajtl, Vasileios Argyriou, Paolo Remagnino |
IEEE Trans. Image Process. | 3 |
| 2019 | Learning how to analyse crowd behaviour using synthetic dataabstractDeep learning has greatly improved pattern recognition algorithms, however its strength is also its weakness: training data must be plentiful, varied and, above all, annotated. Annotating data for crowd analysis in unrealistic: for instance, consider the task of accurately counting people in a dense crowd. Such task would require hours of painstaking manual marking where people are in a crowd that might have been captured from far field. We propose a method that builds controlled simulations of people moving in a real environment; people are modelled as synthetic humanoids with realistic appearance, while the background is the actual image of a scene. Controlling illumination, dynamics of people and density in the scene, we can virtually generate an infinite number of simulations and very large annotated datasets. This paper demonstrates how the performance of a conventional deep network, used for a crowd analysis task, can be improved by using simulated datasets. Anish R. Khadka, Mahdi Maktabdar Oghaz, W. Matta, M. Cosentino, Paolo Remagnino, Vasileios Argyriou |
CASA | 6 |
| 2019 | Smart Monitoring of Crops Using Generative Adversarial Networks
Hamideh Kerdegari, Manzoor Razaak, Vasileios Argyriou, Paolo Remagnino |
CAIP (1) | 3 |
| 2019 | Scene and Environment Monitoring Using Aerial Imagery and Deep LearningabstractUnmanned Aerial vehicles (UAV) are a promising technology for smart farming related applications. Aerial monitoring of agriculture farms with UAV enables key decision-making pertaining to crop monitoring. Advancements in deep learning techniques have further enhanced the precision and reliability of aerial imagery based analysis. The capabilities to mount various kinds of sensors (RGB, spectral cameras) on UAV allows remote crop analysis applications such as vegetation classification and segmentation, crop counting, yield monitoring and prediction, crop mapping, weed detection, disease and nutrient deficiency detection and others. A significant amount of studies are found in the literature that explores UAV for smart farming applications. In this paper, a review of studies applying deep learning on UAV imagery for smart farming is presented. Based on the application, we have classified these studies into five major groups including: vegetation identification, classification and segmentation, crop counting and yield predictions, crop mapping, weed detection and crop disease and nutrient deficiency detection. An in depth critical analysis of each study is provided. Mahdi Maktabdar Oghaz, Manzoor Razaak, Hamideh Kerdegari, Vasileios Argyriou, Paolo Remagnino |
DCOSS | 4 |
| 2019 | Smart IoT Cameras for Crowd Analysis based on augmentation for automatic pedestrian detection, simulation and annotationabstractSmart video sensors for applications related to surveillance and security are IOT-based as they use Internet for various purposes. Such applications include crowd behaviour monitoring and advanced decision support systems operating and transmitting information over internet. The analysis of crowd and pedestrian behaviour is an important task for smart IoT cameras and in particular video processing. In order to provide related behavioural models, simulation and tracking approaches have been considered in the literature. In both cases ground truth is essential to train deep models and provide a meaningful quantitative evaluation. We propose a framework for crowd simulation and automatic data generation and annotation that supports multiple cameras and multiple targets. The proposed approach is based on synthetically generated human agents, augmented frames and compositing techniques combined with path finding and planning methods. A number of popular crowd and pedestrian data sets were used to validate the model, and scenarios related to annotation and simulation were considered. Antoine Rimboux, Rob Dupre, Eldriona Daci, Thomas Lagkas, Panagiotis G. Sarigiannidis, Paolo Remagnino, Vasileios Argyriou |
DCOSS | 7 |
| 2019 | Multi-scale Feature Fused Single Shot Detector for Small Object Detection in UAV Images
Manzoor Razaak, Hamideh Kerdegari, Vasileios Argyriou, Paolo Remagnino |
ICVS | 3 |
| 2019 | A human and group behavior simulation evaluation framework utilizing composition and video analysisabstractAbstract In this work, we present the modular crowd simulation evaluation through composition framework, which provides a quantitative comparison between different pedestrian and crowd simulation approaches. Evaluation is made based on the comparison of source footage against synthetic video created through novel composition techniques. The proposed framework seeks to reduce the complexity of simulation evaluation and provide a platform from which the comparison of differing simulation algorithms and parametric tuning can be conducted to improve simulation accuracy or provide measures of similarity between crowd simulation algorithms and source data. Through the use of features designed to mimic the human visual system, specific simulation properties can be evaluated relative to sample footage. Validation was performed on a number of popular crowd data sets and through comparisons of multiple pedestrian and crowd simulation algorithms. Rob Dupre, Vasileios Argyriou |
Comput. Animat. Virtual Worlds | 2 |
| 2019 | Phase Amplified Correlation for Improved Sub-Pixel Motion EstimationabstractPhase correlation (PC) is widely employed by several sub-pixel motion estimation techniques in an attempt to accurately and robustly detect the displacement between two images. To achieve sub-pixel accuracy, these techniques employ interpolation methods and function-fitting approaches on the cross-correlation function derived from the PC core. However, such motion estimation techniques still present a lower bound of accuracy that cannot be overcome. To allow room for further improvements, we propose in this paper the enhancement of the sub-pixel accuracy of motion estimation techniques by employing a completely different approach: the concept of motion magnification. To this end, we propose the novel phase amplified correlation (PAC) that integrates motion magnification between two compared images inside the phase correlation part of frequencybased motion estimation algorithms and thus directly substitutes the PC core. The experimentation on magnetic resonance (MR) images and real video sequences demonstrates the ability of the proposed PAC core to make subtle motions highly distinguishable and improve the sub-pixel accuracy of frequency-based motion estimation techniques. Dimitrios Konstantinidis, Tania Stathaki, Vasileios Argyriou |
IEEE Trans. Image Process. | 3 |
| 2018 | AMNet: Memorability Estimation With AttentionabstractIn this paper we present the design and evaluation of an end-to-end trainable, deep neural network with a visual attention mechanism for memorability estimation in still images. We analyze the suitability of transfer learning of deep models from image classification to the memorability task. Further on we study the impact of the attention mechanism on the memorability estimation and evaluate our network on the SUN Memorability and the LaMem datasets. Our network outperforms the existing state of the art models on both datasets in terms of the Spearman's rank correlation as well as the mean squared error, closely matching human consistency. Jiri Fajtl, Vasileios Argyriou, Dorothy Ndedi Monekosso, Paolo Remagnino |
CVPR | 2 |
| 2018 | Deep Residual Network with Subclass Discriminant Analysis for Crowd Behavior RecognitionabstractIn this work, we extract rich representations of crowd behavior from video using a fine-tuned deep convolutional neural residual network. Using spatial partitioning trees we create subclasses within the feature maps from each of the crowd behavior attributes (classes). Features from these subclasses are then regularized using an eigen modeling scheme. This enables to model the variance appearing from the intra-subclass information. Low dimensional discriminative features are extracted after using the total subclass scatter information. Dynamic time warping is used on the cosine distance measure to find the similarity measure between videos. A 1-nearest neighbor (NN) classifier is used to find the respective crowd behavior attribute classes from the normal videos. Experimental results on large crowd behavior video database show the superior performance of our proposed framework as compared to the baseline and current state-of-the-art methodologies for the crowd behavior recognition task. Bappaditya Mandal, Jiri Fajtl, Vasileios Argyriou, Dorothy Ndedi Monekosso, Paolo Remagnino |
ICIP | 3 |
| 2018 | Superframes, A Temporal Video SegmentationabstractThe goal of video segmentation is to turn video data into a set of concrete motion clusters that can be easily interpreted as building blocks of the video. There are some works on similar topics like detecting scene cuts in a video, but there is few specific research on clustering video data into the desired number of compact segments. It would be more intuitive, and more efficient, to work with perceptually meaningful entity obtained from a low-level grouping process which we call it `superframe'. This paper presents a new simple and efficient technique to detect superframes of similar content patterns in videos. We calculate the similarity of content-motion to obtain the strength of change between consecutive frames. With the help of existing optical flow technique using deep models, the proposed method is able to perform more accurate motion estimation efficiently. We also propose two criteria for measuring and comparing the performance of different algorithms on various databases. Experimental results on the videos from benchmark databases have demonstrated the effectiveness of the proposed method. Hajar Sadeghi Sokeh, Vasileios Argyriou, Dorothy Ndedi Monekosso, Paolo Remagnino |
ICPR | 2 |
| 2017 | Large Pose 3D Face Reconstruction from a Single Image via Direct Volumetric CNN Regressionabstract3D face reconstruction is a fundamental Computer Vision problem of extraordinary difficulty. Current systems often assume the availability of multiple facial images (sometimes from the same subject) as input, and must address a number of methodological challenges such as establishing dense correspondences across large facial poses, expressions, and non-uniform illumination. In general these methods require complex and inefficient pipelines for model building and fitting. In this work, we propose to address many of these limitations by training a Convolutional Neural Network (CNN) on an appropriate dataset consisting of 2D images and 3D facial models or scans. Our CNN works with just a single 2D facial image, does not require accurate alignment nor establishes dense correspondence between images, works for arbitrary facial poses and expressions, and can be used to reconstruct the whole 3D facial geometry (including the non-visible parts of the face) bypassing the construction (during training) and fitting (during testing) of a 3D Morphable Model. We achieve this via a simple CNN architecture that performs direct regression of a volumetric representation of the 3D facial geometry from a single 2D image. We also demonstrate how the related task of facial landmark localization can be incorporated into the proposed framework and help improve reconstruction quality, especially for the cases of large poses and facial expressions. Code and models will be made available at http://aaronsplace.co.uk. Aaron S. Jackson, Adrian Bulat, Vasileios Argyriou, Georgios Tzimiropoulos |
ICCV | 3 |
| 2017 | Frequency domain subpixel registration using HOG phase correlation
Vasileios Argyriou, Georgios Tzimiropoulos |
Comput. Vis. Image Underst. | 1 |
| 2017 | Linear latent low dimensional space for online early action recognition and prediction
Victoria Bloom, Vasileios Argyriou, Dimitrios Makris 0001 |
Pattern Recognit. | 2 |
| 2016 | Risk assessment for RGBD scans in real timeabstractIn this paper we address the notion of risk assessment of three dimensional scenes. Furthermore through the use of local feature recognition techniques and machine learning we perform this analysis on real time point cloud recordings. We provide a definition of risk and potential hazards that incorporates different elements but mainly focuses on intrinsic risk related properties of an object (e.g sharpness). A 3D Voxel HOG descriptor is utilised that aims to classify and recognise the presence of hazardous characteristics and features of objects present in a given scene. Additionally we utilise and extend the 3D Risk Scenes Dataset (3DRS) designed for risk evaluation in scene analysis. The effectiveness of our method is tested on captured point cloud sequences containing hazardous and non hazardous data with a high degree of accuracy across all tested data. Rob Dupre, Vasileios Argyriou |
ICASSP | 2 |
| 2016 | Gradient schemes for robust FFT-based motion estimationabstractIn this work, the focus is on gradient-based correlation schemes which constitute an alternative to phase correlation for FFT-based motion estimation. In particular, our contribution is threefold. First, we present an analysis which highlights the key features of gradient schemes. Second, we introduce an illumination invariant to gradient correlation. Third, we provide a comparison of gradient schemes in the application of block matching for video coding and draw several useful observations and conclusions. Georgios Tzimiropoulos, Vasileios Argyriou |
ICASSP | 2 |
| 2016 | Gaze estimation using EEG signals for HCI in augmented and virtual reality headsetsabstractAugmented and virtual reality have evolved significantly over the last few years providing new ways of entertainment and interaction with the environment. Although many systems and solutions are currently available, still there is much left unsettled and some technologies are missing from many VR/AR devices, such as foveated rendering and HCI. In this paper, a novel approach for coarse gaze estimation using EEG sensors with applications in items selection for HCI or foveated rendering for VR/AR devices is proposed. The suggested method requires only few electroencephalogram sensors that can be easily added to the current virtual and augmented reality headsets. A supervised machine leaning approach was suggested utilising novel features, based on quaternions allowing gaze estimation. Experiments were performed to evaluate the proposed method and a new dataset was designed and captured. Finally, the introduced learning framework was compared with other similar techniques demonstrating further the gain of the proposed descriptors. Juan Manuel Fernandez Montenegro, Vasileios Argyriou |
ICPR | 2 |
| 2016 | Hierarchical transfer learning for online recognition of compound actions
Victoria Bloom, Vasileios Argyriou, Dimitrios Makris 0001 |
Comput. Vis. Image Underst. | 2 |
| 2016 | Risk analysis for smart homes and domestic robots using robust shape and physics descriptors, and complex boosting techniques
Rob Dupre, Vasileios Argyriou, Georgios Tzimiropoulos, Darrel Greenhill |
Inf. Sci. | 2 |
| 2015 | A 3D Scene Analysis Framework and Descriptors for Risk EvaluationabstractIn this paper we evaluate the notion of scene analysis with regard to risk. We consider the problem of evaluating risk and potential hazards in an environment and providing a quantified risk score. A definition of risk is given incorporating two elements, Firstly scene stability, where Newtonian Physics are introduced into the scene analysis process, evaluating object stability within a scene. The effectiveness of which is demonstrated by conducting experiments on several scenes including a variety of stability levels. Secondly the analysis of the intrinsic risk related properties of an object, which is estimated using learning techniques and the utilisation of the 3D Voxel HOG descriptor, analysed against the state-of-the-art descriptors. Finally a new dataset is provided that is designed for scene analysis focusing on risk evaluation. Rob Dupre, Vasileios Argyriou, Darrel Greenhill, Georgios Tzimiropoulos |
3DV | 2 |
| 2015 | Corrigendum to 'Homage to Professor Maria Petrou' [ Pattern Recognition Letters 48 (2014) 2-7]
Xavier Lladó, Atsushi Imiya, David Mason, Constantino Carlos Reyes-Aldasoro, Kazuaki Aoki, Mineichi Kudo, Yu-Jin Zhang, Vasileios Argyriou |
Pattern Recognit. Lett. | 8 |
| 2014 | Clustered Spatio-temporal Manifolds for Online Action RecognitionabstractIn this paper, a novel method is presented for low-latency online action recognition from skeleton data. The introduction of pose based features has reduced viewpoint and anthropometric variations, so differing execution rates and personal styles are the major sources of classification error. Previous work for online action recognition fails to adequately address both execution rate and personal style. To overcome these limitations a compression and fusion of offline action recognition approaches has transpired. Specifically, clustered action manifolds are proposed for low computational latency and template fragment matching with peak key poses are introduced for low observational latency. The style invariance of spatio-temporal manifolds is combined with the execution rate invariance of Dynamic Time Warping (DTW). Experimental results on two publicly available datasets demonstrate the high accuracy of the proposed method. Victoria Bloom, Dimitrios Makris 0001, Vasileios Argyriou |
ICPR | 3 |
| 2014 | Optimal illumination directions for faces and rough surfaces for single and multiple light imaging using class-specific prior knowledge
Vasileios Argyriou, Stefanos Zafeiriou, Maria Petrou |
Comput. Vis. Image Underst. | 1 |
| 2014 | Cast shadows estimation and synthesis using the Walsh transform
Ferdinand Redelinghuys, Vasileios Argyriou, Maria Petrou |
Pattern Recognit. Lett. | 2 |
| 2013 | A sparse representation method for determining the optimal illumination directions in Photometric Stereo
Vasileios Argyriou, Stefanos Zafeiriou, Barbara Villarini, Maria Petrou |
Signal Process. | 1 |
| 2013 | Guest Editorial Introduction to the Special Issue on Modern Control for Computer GamesabstractA typical gaming scenario, as developed in the past 20 years, involves a player interacting with a game using a specialized input device, such as a joystic, a mouse, a keyboard, etc. Recent technological advances and new sensors (for example, low cost commodity depth cameras) have enabled the introduction of more elaborated approaches in which the player is now able to interact with the game using his body pose, facial expressions, actions, and even his physiological signals. A new era of games has already started, employing computer vision techniques, brain-computer interfaces systems, haptic and wearable devices. The future lies in games that will be intelligent enough not only to extract the player's commands provided by his speech and gestures but also his behavioral cues, as well as his/her emotional states, and adjust their game plot accordingly in order to ensure more realistic and satisfactory gameplay experience. This special issue on modern control for computer games discusses several interdisciplinary factors that influence a user's input to a game, something directly linked to the gaming experience. These include, but are not limited to, the following: behavioral affective gaming, user satisfaction and perception, motion capture and scene modeling, and complete software frameworks that address several challenges risen in such scenarios. Vasileios Argyriou, Irene Kotsia, Stefanos Zafeiriou, Maria Petrou |
IEEE Trans. Cybern. | 1 |
| 2013 | Face Recognition and Verification Using Photometric Stereo: The Photoface Database and a Comprehensive EvaluationabstractThis paper presents a new database suitable for both 2-D and 3-D face recognition based on photometric stereo (PS): the Photoface database. The database was collected using a custom-made four-source PS device designed to enable data capture with minimal interaction necessary from the subjects. The device, which automatically detects the presence of a subject using ultrasound, was placed at the entrance to a busy workplace and captured 1839 sessions of face images with natural pose and expression. This meant that the acquired data is more realistic for everyday use than existing databases and is, therefore, an invaluable test bed for state-of-the-art recognition algorithms. The paper also presents experiments of various face recognition and verification algorithms using the albedo, surface normals, and recovered depth maps. Finally, we have conducted experiments in order to demonstrate how different methods in the pipeline of PS (i.e., normal field computation and depth map reconstruction) affect recognition and verification performance. These experiments help to 1) demonstrate the usefulness of PS, and our device in particular, for minimal-interaction face recognition, and 2) highlight the optimal reconstruction and recognition algorithms for use with natural-expression PS data. The database can be downloaded from http://www.uwe.ac.uk/research/Photoface. Stefanos Zafeiriou, Gary A. Atkinson, Mark F. Hansen, William A. P. Smith, Vasileios Argyriou, Maria Petrou, Melvyn L. Smith, Lyndon N. Smith |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2012 | Toward a Two-Handed Gesture-Based Visual 3D Interactive Object-Oriented Environment for Software DevelopmentabstractThis paper presents the conceptual and architectural design of a multilayered framework aiming to provide a two handed gesture based visual interactive 3D object oriented environment for software development. We argue that this is a viable, intuitive and attractive approach for software development facilitating natural human computer interaction, thus supporting tasks related to software development in a more effective and efficient way. We illustrate the value of the proposed environment through a data analysis from early evaluations of one prototype system in use. The results and implications of this study are useful for designing effective and efficient gesture based 3D interactive software development environments. Raúl A. Herrera Acuña, Christos Fidas, Vasileios Argyriou, Sergio A. Velastin |
Intelligent Environments | 3 |
| 2011 | Sub-Hexagonal Phase Correlation for Motion EstimationabstractWe present a novel frequency-domain motion estimation technique, which operates on hexagonal images and employs the hexagonal Fourier transform. Our method involves image sampling on a hexagonal lattice followed by a normalised hexagonal cross-correlation in the frequency domain. The term subpixel (or subcell) is defined on a hexagonal grid in order to achieve floating point registration. Experiments using both artificially induced motion and actual motion demonstrate that the proposed method outperforms the state-of-the-art in frequency-domain motion estimation operating on a square lattice, in the shape of phase correlation, in terms of subpixel accuracy for a range of test material and motion scenarios. Vasileios Argyriou |
IEEE Trans. Image Process. | 1 |
| 2011 | Subpixel Registration With Gradient CorrelationabstractWe address the problem of subpixel registration of images assumed to be related by a pure translation. We present a method which extends gradient correlation to achieve subpixel accuracy. Our scheme is based on modeling the dominant singular vectors of the 2-D gradient correlation matrix with a generic kernel which we derive by studying the structure of gradient correlation assuming natural image statistics. Our kernel has a parametric form which offers flexibility in modeling the functions obtained from various types of image data. We estimate the kernel parameters, including the unknown subpixel shifts, using the Levenberg-Marquardt algorithm. Experiments with LANDSAT and MRI data show that our scheme outperforms recently proposed state-of-the-art phase correlation methods. Georgios Tzimiropoulos, Vasileios Argyriou, Tania Stathaki |
IEEE Trans. Image Process. | 2 |
| 2010 | Photometric stereo with an arbitrary number of illuminants
Vasileios Argyriou, Maria Petrou, Svetlana Barsky |
Comput. Vis. Image Underst. | 1 |
| 2010 | Robust FFT-Based Scale-Invariant Image Registration with Image GradientsabstractWe present a robust FFT-based approach to scale-invariant image registration. Our method relies on FFT-based correlation twice: once in the log-polar Fourier domain to estimate the scaling and rotation and once in the spatial domain to recover the residual translation. Previous methods based on the same principles are not robust. To equip our scheme with robustness and accuracy, we introduce modifications which tailor the method to the nature of images. First, we derive efficient log-polar Fourier representations by replacing image functions with complex gray-level edge maps. We show that this representation both captures the structure of salient image features and circumvents problems related to the low-pass nature of images, interpolation errors, border effects, and aliasing. Second, to recover the unknown parameters, we introduce the normalized gradient correlation. We show that, using image gradients to perform correlation, the errors induced by outliers are mapped to a uniform distribution for which our normalized gradient correlation features robust performance. Exhaustive experimentation with real images showed that, unlike any other Fourier-based correlation techniques, the proposed method was able to estimate translations, arbitrary rotations, and scale factors up to 6. Georgios Tzimiropoulos, Vasileios Argyriou, Stefanos Zafeiriou, Tania Stathaki |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2009 | FFT-based estimation of large motions in images: A robust gradient-based approachabstractA fast and robust gradient-based motion estimation technique which operates in the frequency domain is presented. The algorithm combines the natural advantages of a good feature selection offered by gradient-based methods with the robustness and speed provided by FFT-based correlation schemes. Experimentation with real images taken from a popular database showed that, unlike any other Fourier-based techniques, the method was able to estimate translations, arbitrary rotations and scale factors in the range 4-6. Georgios Tzimiropoulos, Vasileios Argyriou, Tania Stathaki |
ICASSP | 2 |
| 2008 | Generalisation of Photometric Stereo technique to Q-illuminantsabstractWe present a generalisation of the 4-light photometric stereo technique to an arbitrary number Q of illuminants. We assume that the surface reflectance can be approximated by the Lambertian model plus a reflectance component. The algorithm works in a recursive manner eliminating the pixel intensities affected by shadows or highlights based on a least squares error technique, retaining only the information coming from illumination directions that can be used for photometric stereo reconstruction of the normal of the corresponding surface patch. We report results for both simulated and real surfaces and compare them with the results of other state of the art photometric stereo algorithms. 1 Vasileios Argyriou, Svetlana Barsky, Maria Petrou |
BMVC | 1 |
| 2008 | A Frequency Domain Approach to Roto-translation Estimation using Gradient Cross-CorrelationabstractA novel frequency domain approach to roto-translation estimation is presented. The baseline gradient cross-correlation method is extended to handle rotations. A key feature of the proposed scheme is the ability to achieve good performance in the presence of both large translations and rotations as well as noise, a scenario for which other Fourier-based methods typically fail. Robustness and accuracy in conjunction with computational efficiency, offered by the frequency domain formulation, make the algorithm useful in a number of image processing tasks such as image registration. 1 Georgios Tzimiropoulos, Vasileios Argyriou, Tania Stathaki |
BMVC | 2 |
| 2008 | Recursive photometric stereo when multiple shadows and highlights are presentabstractWe present a recursive algorithm for 3D surface reconstruction based on photometric stereo in the presence of highlights, and self and cast shadows. We assume that the surface reflectance outside the highlights can be approximated by the Lambertian model. The algorithm works with as few as three light sources, and it can be generalised for N without any difficulties. Furthermore, this reconstruction method is able to identify areas where the majority of the lighting directions result in unreliable pixel intensities, providing the capability to adjust a reconstruction algorithm and improve its performance avoiding the unreliable sources. We report results for both artificial and real images and compare them with the results of other state of the art photometric stereo algorithms. Vasileios Argyriou, Maria Petrou |
CVPR | 1 |
| 2008 | Symmetry detection using frequency domain motion estimation techniquesabstractA frequency domain approach for the detection of symmetries in real images is presented. Our framework is based on recent state-of-the-art research where motion estimation techniques are employed to sequentially determine all the associated parameters. In particular, we introduce several modifications regarding the order of symmetry estimation and the detection of the axes of possible bilateral symmetry. Preliminary results demonstrate the efficiency of our approach. Georgios Tzimiropoulos, Vasileios Argyriou, Tania Stathaki |
ICASSP | 2 |
| 2008 | Gradient-Adaptive Normalized ConvolutionabstractSignal estimation for sparsely and irregularly sampled signals can be carried out using either noniterative methods, or iterative methods or methods that deal with irregular samples and their uncertainty, through normalized convolution. The latter is a general method for filtering incomplete or uncertain data and is based on the separation of both data and operator into a signal part and a certainty part. It has been proven that normalized convolution yields a local description which is optimal both in an algebraic and a least-squares sense. In this letter, we employ the normalized convolution concept to formulate a novel reconstruction method for irregularly sampled signals, utilizing an anisotropic, rotated applicability filter. Our experimental results demonstrate performance gains in a least-squares sense, retaining edge and contour information, especially in sparsely sampled areas on the image plane. Vasileios Argyriou, Theodore Vlachos, Roberta Piroddi |
IEEE Signal Process. Lett. | 1 |
| 2007 | Quad-Tree Motion Estimation in the Frequency Domain Using Gradient CorrelationabstractWe propose a new motion estimation scheme particularly suitable for broadcast-quality digital video applications due to its performance-complexity characteristics. The proposed scheme is based on the principle of gradient correlation and its computational efficiency is due to the fact that it operates in the frequency domain. The scheme involves the quad-tree decomposition of a frame thus providing a better level of adaptation to scene contents compared to fixed block size approaches. Quad-tree decompositions are obtained by using the motion compensated prediction error to control the partition of a parent block to four children quadrants. The partition criterion is applied iteratively until a target number of motion vectors or a target level of motion compensated prediction error is achieved or, ultimately, until no more than a single motion component can be identified. The partition criterion also guarantees a monotonic decrease of the motion compensated prediction error with an increasing number of iterations making our scheme suitable for progressive transmission and embedded coding applications. Our results show that our scheme outperforms fixed block size phase correlation as well as quad-tree motion estimation based on phase correlation in terms of rate-distortion characteristics yielding smoother motion vector fields as well as more compressible motion-compensated prediction residuals. Vasileios Argyriou, Theodore Vlachos |
IEEE Trans. Multim. | 1 |
| 2006 | A Study of Sub-pixel Motion Estimation using Phase CorrelationabstractWe propose a method for obtaining high-accuracy sub-pixel motion estimates using phase correlation. Our method is motivated by recently published analysis according to which the Fourier inverse of the normalized cross-power spectrum of pairs of images which have been mutually shifted by a fractional amount can be approximated by a two-dimensional sinc function. We use a modified version of such a function to obtain a sub-pixel estimate of motion by means of variable-separable fitting in the vicinity of the maximum peak of the phase correlation surface. We demonstrate that our method outperforms, in terms of sub-pixel accuracy, not only other surface fitting techniques but also the state-of-the-art in motion estimation using phase correlation including the technique that motivated our work in the first place. Furthermore our method performs particularly well in the presence of artificially induced additive white Gaussian noise and also offers better motion vector coherence in terms of zero-order entropy. Vasileios Argyriou, Theodore Vlachos |
BMVC | 1 |
| 2005 | Motion estimation using quad-tree phase correlationabstractWe propose a quad-tree scheme for obtaining subpixel estimates of interframe motion in the frequency domain. Our scheme is based on phase correlation and uses key features of the phase correlation surface to control the partition of a parent block to four children quadrants. The partition criterion is applied iteratively until a target number of motion vectors or a target level of motion compensated prediction error is achieved or, ultimately, if no more than a single motion component per block can be identified. Our results show that our scheme provides a better level of adaptation to scene contents and outperforms fixed block size phase correlation in terms of total motion compensated prediction error for the same number of motion vectors and also in terms of number of motion vectors for the same level of motion compensated prediction error. Vasileios Argyriou, Theodore Vlachos |
ICIP (1) | 1 |
| 2005 | Quad-Tree Motion Estimation in the Frequency DomainabstractWe propose a quad-tree scheme for obtaining sub-pixel estimates of interframe motion in the frequency domain. Our scheme is based on phase correlation and uses motion compensated prediction error to control the partition of a parent block to four children quadrants. This criterion guarantees a monotonic decrease of the motion compensated prediction error with an increasing number of iterations making our scheme suitable for embedded coding applications. Our results show that our scheme provides a better level of adaptation to scene contents and outperforms fixed block size phase correlation in terms of total motion compensated prediction error for the same number of motion vectors and also in terms of number of motion vectors for the same level of motion compensated prediction error Vasileios Argyriou, Theodore Vlachos |
ICME | 1 |
| 2005 | A User-Oriented Multimodal-Interface Framework for General Content-Based Multimedia RetrievalabstractA user-oriented multimodal interface (MMI) framework is proposed. Considering the complexities of media connotations and uncertainties of the user’s demands, content-based retrieval has intrinsic requirements for MMI for effective media-content interactions. Through integration of knowledge based conduction, learning of semantic concepts, natural language processing and analysis of users’ profiles, our framework can establish a solid basis for design and implementation of general CBR systems satisfying extensibility, condensability and inter-operability. Jinchang Ren, Theodore Vlachos, Vasileios Argyriou |
ICME | 3 |
| 2004 | Using gradient correlation for sub-pixel motion estimation of video sequencesabstractA highly accurate and computationally efficient method is presented suitable for the estimation of motion in video sequences. The method is based on the maximization of the spatial gradient cross-correlation function, which is computed in the frequency domain and therefore can be implemented by fast transformation algorithms. We present enhancements to the baseline gradient-correlation algorithm, which further improve performance, especially in the presence of manually induced additive Gaussian noise. We also present a comparative performance analysis, which demonstrates that the proposed method outperforms the state-of-the-art in frequency-domain motion estimation, in the shape of phase correlation. Vasileios Argyriou, Theodore Vlachos |
ICASSP (3) | 1 |