Christos Papaioannidis

dblp:254/8183 · DBLP profile ↗
← Back
16ranked-venue papers
6as first author
13since 2021 · last 2026
0000-0002-3839-4514ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 5 first-author · 8 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Computer networks · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Extreme weakly supervised binary semantic image segmentation via one-pixel supervision
abstract
Despite recent advancements, Unsupervised Semantic Segmentation (USS) methods still exhibit a significant performance deficit compared to supervised approaches, particularly in binary semantic segmentation. This limitation arises because, without supervision, USS methods struggle to distinguish foreground from background image regions, particularly when the foreground contains small or uncommon objects. This issue is addressed by our proposed Extremely Weakly Supervised Binary Semantic Segmentation (EWS) framework. EWS expects minimal supervision, consisting only of a small set of one-pixel annotations explicitly belonging to the foreground class across the entire image dataset. Our approach leverages these one-pixel annotations and employs two contrastive losses to map visual transformer features into well-separated foreground and background feature clusters. Additionally, we propose a novel loss function to eliminate the need for hyperparameter tuning of the contrastive loss threshold, by dynamically computing it based on the similarity between the input image features. Even if we employ a single one-pixel annotation, EWS achieves competitive results in binary segmentation tasks while maintaining low computational costs, making it an efficient solution for critical segmentation applications. GitHub Repo: https://github.com/matJTzimas/EWS
Matthaios Dimitrios Tzimas, Vasileios Mygdalis, Christos Papaioannidis, Ioannis Pitas
Pattern Recognit.3
2025 Padnet: a Patch-Based Anomaly Detection Framework for Industrial Pipeline Damage Detection
abstract
Industrial pipeline inspection in petrochemical refineries is dangerous, expensive, time-consuming and prone to errors. Anomaly detection can play a crucial role towards its automation. Damages in this type of infrastructure are few and can be considered as anomalies (essentially outliers). This paper proposes a novel patch-based Anomaly Detection Network (PADNet), that employs deep learning for detecting insulated pipe damages. It consists of three main components: a) a pipeline segmentation module, b) an image patch proposal module, and c) an anomaly detection module. These components work sequentially first to localize insulated pipelines in the input UAV or ground camera images or video frames and then analyze image patches to detect and localize any damages. Importantly, the anomaly detection module can be trained using undamaged pipeline image data only, hence eliminating the need for costly damaged pipeline image annotation. Experimental results demonstrate the effectiveness of the proposed PADNet method in detecting pipeline damages, making it a promising solution for autonomous industrial infrastructure inspection.
Erofili Alexaki, Christos Papaioannidis, Vasileios Mygdalis, Ioannis Pitas
ICASSP2
2025 Divide-and-Summarize: Enhancing Deep Neural Video Summarization
abstract
Sequence-based neural architectures, such as Long Short-Term Memory (LSTM) networks and Transformers, have driven advances in supervised video summarization by modeling inter-frame dependencies. However, existing methods assume that long-range dependencies are essential for summary generation, which may lead to unnecessary computational overhead. To address this, we propose a Field of View (FOV) adjustment strategy, Divide-and-Summarize (DIV-SUM). By partitioning input videos into smaller fragments of predefined size, our approach explicitly models short-range inter-frame relationships, enabling a fully parallelizable end-to-end video summarization pipeline. Furthermore, most prior work formulates neural video summarization as a frame-wise score regression task. We introduce a simple yet effective target space quantization module, which discretizes the regression targets into classes, introducing a tolerance margin that improves performance. Our approach offers two key benefits: (1) we achieve state-of-the-art performance on the SumMe benchmark while remaining competitive on TVSum, and (2) we significantly reduce the computational cost of inference, improving efficiency without sacrificing quality.
Evangelos Charalampakis, Christos Papaioannidis, Ioannis Pitas
ICIP2
2025 Blaze: A Dataset For Wildfire And Burnt Area UAV Image Classification And Segmentation
abstract
The easy and cost-effective deployment of Unmanned Aerial Vehicles (UAVs) that can fly over an area, and collect data for wildfire detection/assessment or for determining the extent of the disaster (e.g., measuring the area of burnt forest regions) using Deep Neural Network (DNN) algorithms, highlights the importance of UAVs and Artificial Intelligence (AI) algorithms for Natural Disaster Management (NDM). However, there is a limited availability of properly annotated data for the abovementioned tasks, which are typically necessary for training accurate and reliable DNN-based AI algorithms. In this direction, this paper introduces the BLAZE dataset, comprising approximately 5.4K annotated RGB images depicting both urban and non-urban areas before, during and after a wildfire. Moreover, using the proposed dataset, several baseline DNN-based algorithms have been trained and evaluated for the wildfire image classification and burnt areas segmentation tasks. Experimental results show that increased accuracy can be achieved for both tasks, thus proving the usefulness of the developed BLAZE dataset in the NDM domain. Data can be found here https://aiia.csd.auth.gr/blaze-fire-classification-segmentation-dataset/.
Michael Siavrakas, Christos Papaioannidis, Ioannis Pitas
ICIP2
2025 Distilling Structural Knowledge: Teaching Representations in Multi-DNN Agent Systems
abstract
Recent advancements in multi-agent systems have highlighted the potential of enabling efficient knowledge exchange among Deep Neural Network (DNN) agents to address complex tasks. This paper investigates the dynamics of DNN teacher-student interactions, with a focus on distilling specific knowledge from teacher DNNs to student DNNs to enhance retrieval performance. We propose an approach that optimizes the feature structure of the student DNN agent, enabling the distillation of detailed representation knowledge. Utilizing the concept of triplets, our method captures data correlations and transfers structural knowledge, aiming to compress the knowledge of representations and their structural data dependencies from larger to smaller DNN agents while preserving performance accuracy. Our triplet-based knowledge distillation strategy guides the student DNN agent to learn optimal representations for image retrieval in a multi-DNN agent system. Experimental results demonstrate enhancements in the student agent’s efficiency, showing improvements in performance across various DNN architectures and datasets.
Ioanna Valsamara, Christos Papaioannidis, Ioannis Pitas
ISCC2
2024 A Unified DNN-Based System for Industrial Pipeline Segmentation
abstract
This paper presents a unified system tailored for autonomous pipe segmentation within an industrial setting. To this end, it is designed to analyze RGB images captured by Unmanned Aerial Vehicle (UAV)-mounted cameras to predict binary pipe segmentation maps. The overall proposed system consists of three main components: a) a Convolutional Neural Network (CNN) that is used to obtain initial estimates of the pipe segmentation maps, b) a point extraction module that acts on the outputs of the CNN to propose strong pipe class representatives in the input image space, and c) a foundation segmentation model, utilized to refine the initial estimations based on the proposed pipe class representatives. The architecture of the proposed system was specifically designed to ensure increased generalization ability in different, unknown environments, offering an effective solution to a well-known limitation of typical segmentation CNNs, at least in the pipe segmentation task. The effectiveness of the proposed system in this particular setting is evaluated by utilizing two pipe segmentation datasets, originating from two different industrial sites, which were manually annotated with the corresponding pipe segmentation maps. Experimental results demonstrate that the proposed system outperforms the baseline segmentation CNNs, demonstrating its remarkable generalization capabilities.
Dimitrios Psarras, Christos Papaioannidis, Vasileios Mygdalis, Ioannis Pitas
ICASSP2
2024 Domain Expertise Assessment for Multi-DNN Agent Systems
abstract
Recently, multi-agent systems that facilitate knowledge sharing among Deep Neural Network (DNN) agents, have gained increasing attention. This paper explores the dynamics of multi-agent systems that support Teacher-Student DNN interactions, where knowledge is distilled from Teachers to Students. Within such systems, selecting the most compatible Teacher for a given task is far from trivial and can lead to low-quality decisions. Hence, the need arises for accurate domain knowledge evaluation. In that context, we propose including an OOD detection module in each DNN agent to enable effective agent expertise evaluation and precise identification of suitable Teachers. This setup allows Student agents to distill knowledge from the most knowledgeable Teachers within a specific domain, ensuring optimal system performance. To effectively utilize OOD detection in this context, we address key challenges such as determining the minimum data cardinality required to ensure optimal performance and reliable inferences of the OOD detectors.
Ioanna Valsamara, Christos Papaioannidis, Ioannis Pitas
ISCC2
2023 Evaluating Deep Neural Network-based Fire Detection for Natural Disaster Management
abstract
Recently, climate change has led to more frequent extreme weather events, introducing new challenges for Natural Disaster Management (NDM) organizations. This fact makes the employment of modern technological tools such as Deep Neural Networks-based fire detectors a necessity, as they can assist such organizations manage these extreme events more effectively. In this work, we argue that the mean Average Precision (mAP) metric that is commonly used to evaluate typical object detection algorithms can not be trusted for the fire detection task, due to its high dependence on the employed data annotation strategy. This means that the mAP score of a fire detection algorithm may be low even when it predicts fire bounding boxes that accurately enclose the depicted fires. In this direction, a new evaluation metric for fire detection is proposed, denoted as Image-level mean Average Precision (ImAP), which reduces the dependence on the bounding box annotation strategy by rewarding/penalizing bounding box predictions on image level, rather than on bounding box level. Experiments using different object detection algorithms have shown that the proposed ImAP metric reveals the true fire detection capabilities of the tested algorithms more effectively.
Matthaios Dimitrios Tzimas, Christos Papaioannidis, Vasileios Mygdalis, Ioannis Pitas
BDCAT2
2023 Fast Single-Person 2D Human Pose Estimation Using Multi-Task Convolutional Neural Networks
abstract
This paper presents a novel neural module for enhancing existing fast and lightweight 2D human pose estimation CNNs, in order to increase their accuracy. A baseline stem CNN is augmented by a collateral module, which is tasked to encode global spatial and semantic information and provide it to the stem network during inference. The latter one outputs the final 2D human pose estimations. Since global information encoding is an inherent subtask of 2D human pose estimation, this particular setup allows the stem network to better focus on the local details of the input image and on precisely localizing each body joint, thus increasing overall 2D human pose estimation accuracy. Furthermore, the collateral module is designed to be lightweight, adding negligible runtime computational cost, so that the unified architecture retains the fast execution property of the stem network. Evaluation of the proposed method on public 2D human pose estimation datasets shows that it increases the accuracy of different baseline stem CNNs, while outperforming all competing fast 2D human pose estimation methods.
Christos Papaioannidis, Ioannis Mademlis, Ioannis Pitas
ICASSP1
2023 Fast CNN-Based Single-Person 2D Human Pose Estimation for Autonomous Systems
abstract
This paper presents a novel Convolutional Neural Network (CNN) architecture for 2D human pose estimation from RGB images that balances between high 2D human pose/skeleton estimation accuracy and rapid inference. Thus, it is suitable for safety-critical embedded AI scenarios in autonomous systems, where computational resources are typically limited and fast execution is often required, but accuracy cannot be sacrificed. The architecture is composed of a shared feature extraction backbone and two parallel heads attached on top of it: one for 2D human body joint regression and one for global human body structure modelling through Image-to-Image Translation (I2I). A corresponding multitask loss function allows training of the unified network for both tasks, through combining a typical 2D body joint regression with a novel I2I term. Along with enhanced information flow between the parallel neural heads via skip synapses, this strategy is able to extract both ample semantic and rich spatial information, while using a less complex CNN; thus it permits fast execution. The proposed architecture is evaluated on public 2D human pose estimation datasets, achieving the best accuracy-speed ratio compared to the state-of-the-art. Additionally, it is evaluated on a pedestrian intention recognition task for self-driving cars, leading to increased accuracy and speed in comparison to competing approaches.
Christos Papaioannidis, Ioannis Mademlis, Ioannis Pitas
IEEE Trans. Circuits Syst. Video Technol.1
2022 An Efficient Framework for Human Action Recognition Based on Graph Convolutional Networks
abstract
This paper presents a novel framework for skeleton-based Human Action Recognition (HAR) based on Graph Convolution Networks (GCNs). The proposed framework aims to increase human action recognition performance of GCN-based methods by incorporating a missing-joint-handling pre-processing step and a novel adjacency matrix construction method in a single human action recognition pipeline. The missing-joint-handling pre-processing step is utilized to infer missing data in the input sequence, which may occur due to imperfect skeleton extraction, based on imputation methods. The novel adjacency matrix construction method is executed offline to compute an improved weighted adjacency matrix specifically designed for HAR, which is utilized in every layer of the employed GCN. Moreover, both the pre-processing step and the adjacency construction method can be utilized along with any GCN architecture, allowing any GCN-based HAR method to be employed in the proposed framework. Experimental evaluation on two public datasets indicate favorable human action classification scores compared to the employed baseline and all competing methods both for 2D and 3D skeleton-based human action recognition, while using a GCN architecture with less learnable parameters.
Nikolaos Kilis, Christos Papaioannidis, Ioannis Mademlis, Ioannis Pitas
ICIP2
2022 Fast Semantic Image Segmentation for Autonomous Systems
abstract
Fast semantic image segmentation is crucial for autonomous systems, as it allows an autonomous system (e.g., self-driving car, drone, etc.) to interpret its environment on-the-fly and decide on necessary actions by exploiting dense semantic maps. The speed of semantic segmentation on embedded computational hardware is as important as its accuracy. Thus, this paper proposes a novel framework for semantic image segmentation that is both fast and accurate. It augments existing real-time semantic image segmentation architectures by an auxiliary, parallel neural branch that is tasked to predict semantic maps in an alternative manner by utilizing Generative Adversarial Networks (GANs). Additional attention-based neural synapses linking the two branches allow information to flow between them during both the training and the inference stage. Extensive experiments on three public datasets for autonomous driving and for aerial-perspective image analysis indicate non-negligible gains in segmentation accuracy, without compromises on inference speed.
Christos Papaioannidis, Ioannis Mademlis, Ioannis Pitas
ICIP1
2021 Autonomous UAV Safety by Visual Human Crowd Detection Using Multi-Task Deep Neural Networks
abstract
Camera-equipped UAVs, or drones, are increasingly employed in a wide range of applications. Thus, ensuring their safe flight in areas containing people is a top priority. In this paper, a deep neural network-based method is proposed for the task of visual human crowd detection from UAV footage, allowing a drone to rapidly extract semantic segmentation maps from captured video frames during flight. These maps can be exploited (e.g., by a path planner) to define no-fly zones over, or near human crowds and, hence, enhance UAV flight safety. To this end, a novel neural architecture for binary (crowd/non- crowd) semantic segmentation from single RGB images is proposed, based on Convolutional Neural Networks (CNNs). It consists of a semantic segmentation and an image-to-image translation (I2I) neural branch. The overall network is trained using a novel multi-task loss function that addresses both tasks by processing the output of the corresponding branch. During inference, information flows across branches through additional skip synapses to further assist the crowd detection task. In order to evaluate the proposed method, we introduce a real and a synthetic human crowd RGB image dataset. The proposed method outperforms previous aerial crowd detection methods by a large margin and without any post-processing. Moreover, it demonstrates increased generalization ability, while running at real-time and near-real-time speeds on a ground computer and on embedded AI hardware, respectively.
Christos Papaioannidis, Ioannis Mademlis, Ioannis Pitas
ICRA1
2020 3D Object Pose Estimation Using Multi-Objective Quaternion Learning
abstract
In this paper, a framework is proposed for object recognition and pose estimation from color images using convolutional neural networks (CNNs). 3D object pose estimation along with object recognition has numerous applications, such as robot positioning versus a target object and robotic object grasping. Previous methods addressing this problem relied on both color and depth (RGB-D) images to learn low-dimensional viewpoint descriptors for object pose retrieval. In the proposed method, a novel quaternion-based multi-objective loss function is used, which combines manifold learning and regression to learn 3D pose descriptors and direct 3D object pose estimation, using only color (RGB) images. The 3D object pose can then be obtained either by using the learned descriptors in the nearest neighbor (NN) search or by direct neural network regression. An extensive experimental evaluation has proven that such descriptors provide greater pose estimation accuracy than the state-of-the-art methods. In addition, the learned 3D pose descriptors are almost object-independent and, thus, generalizable to unseen objects. Finally, when the object identity is not of interest, the 3D object pose can be regressed directly from the network, by overriding the NN search, thus significantly reducing the object pose inference time.
Christos Papaioannidis, Ioannis Pitas
IEEE Trans. Circuits Syst. Video Technol.1
2020 Domain-Translated 3D Object Pose Estimation
abstract
Synthetic 3D object models have been proven crucial in object pose estimation, as they are utilized to generate a huge number of accurately annotated data. The object pose estimation problem is usually solved for images originating from the real data domain by employing synthetic images for training data enrichment, without fully exploiting the fact that synthetic and real images may have different data distributions. In this work, we argue that 3D object pose estimation problem is easier to solve for images originating from the synthetic domain, rather than the real data domain. To this end, we propose a 3D object pose estimation framework consisting of a two-step process, where a novel pose-oriented image-to-image translation step is first employed to translate noisy real images to clean synthetic ones and then, a 3D object pose estimation method is applied on the translated synthetic images to finally predict the 3D object poses. A novel pose-oriented objective function is employed for training the image-to-image translation network, which enforces that pose-related object image characteristics are preserved in the translated images. As a result, the pose estimation network does not require real data for training purposes. Experimental evaluation has shown that the proposed framework greatly improves the 3D object pose estimation performance, when compared to state-of-the-art methods.
Christos Papaioannidis, Vasileios Mygdalis, Ioannis Pitas
IEEE Trans. Image Process.1
2019 Adversarial Face De-Identification
abstract
Recently, much research has been done on how to secure personal data, notably facial images. Face de-identification is one example of privacy protection that protects person identity by fooling intelligent face recognition systems, while typically allowing face recognition by human observers. While many face de-identification methods exist, the generated de-identified facial images do not resemble the original ones. This paper proposes the usage of adversarial examples for face de-identification that introduces minimal facial image distortion, while fooling automatic face recognition systems. Specifically, it introduces P-FGVM, a novel adversarial attack method, which operates on the image spatial domain and generates adversarial de-identified facial images that resemble the original ones. A comparison between P-FGVM and other adversarial attack methods shows that P-FGVM both protects privacy and preserves visual facial image quality more efficiently.
Efstathios Chatzikyriakidis, Christos Papaioannidis, Ioannis Pitas
ICIP2