VLDB 2026 Research / reviewers in the wild / expert
Konstantinos Ioannidis
dblp:35/2092
· DBLP profile ↗
27ranked-venue papers
0as first author
20since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 21 · 15 since 2021Artificial intelligence and machine learning · 11 · 9 since 2021Databases, data management, data science and information retrieval · 5 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HDD-Unet: A Unet-based architecture for low-light image enhancementabstractLow-light imaging has become a popular topic in image processing, with the quality enhancement of low light images being as a significant challenge, due to the difficulty in retaining colors, patterns, texture and style when generating a normal light image. Our objectives are mainly to firstly better preserve texture regions in image enhancement, while, secondly, preserving colors via color histogram blocks and, finally, to enhance the quality of image through dense denoising blocks. Our proposed novel framework, namely HDD-Unet, is a double Unet based on photorealistic style transfer for low-light image enhancement. The proposed low-light image enhancement method combines color histogram-based fusion, Haar wavelet pooling, dense-denoising blocks and U-net as a backbone architecture to enhance the contrast, reduce noise, and improve the visibility of low light images. Experimental results demonstrate that our proposed method outperforms existing methods in terms of PSNR and SSIM quantitative evaluation metrics, reaching or outperforming state-of-the-art accuracy, but with less resources. We also conduct an ablation study to investigate the impact of our approach on overexposed images, and systematic analysis on the late fusion weighting parameters. Multiple experiments were conducted with artificial noise inserted to accomplish more efficient comparison. The results show that the proposed framework enhances accurately images with various gamma corrections. The proposed method represents a significant advance in the field of low light image enhancement and has the potential to address several challenges associated with low light imaging. • Low-light image enhancement based on a combination of two U-Net branches with wavelet transformations, Adaptive Instance Normalization, color histogram blocks and denoising dense blocks. • Analytic framework creation for the HDD-Unet architecture. • Evaluation of the proposed method in a qualitative and a quantitative way. Elissavet Batziou, Konstantinos Ioannidis, Ioannis Patras, Stefanos Vrochidis, Ioannis Kompatsiaris |
Image Vis. Comput. | 2 |
| 2026 | Unsupervised Object Localization driven by self-supervised foundation models: A comprehensive reviewabstractObject localization is a fundamental task in computer vision that traditionally requires labeled datasets for accurate results. Recent progress in self-supervised learning has enabled unsupervised object localization, reducing reliance on manual annotations. Unlike supervised encoders, which depend on annotated training data, self-supervised encoders learn semantic representations directly from large collections of unlabeled images. This makes them the natural foundation for unsupervised object localization, as they capture object-relevant features while eliminating the need for costly manual labels. These encoders produce semantically coherent patch embeddings. Grouping these embeddings reveals sets of patches that correspond to objects in an image. These patch sets can be converted into object masks or bounding boxes, enabling tasks such as single-object discovery, multi-object detection, and instance segmentation. By applying off-line mask clustering or using pre-trained vision-language models, unsupervised localization methods can assign semantic labels to discovered objects. This transforms initially class-agnostic objects (objects without class labels) into class-aware ones (objects with class labels), aligning these tasks with their supervised counterparts. This paper provides a structured review of unsupervised object localization methods in both class-agnostic and class-aware settings. In contrast, previous surveys have focused only on class-agnostic localization. We discuss state-of-the-art object discovery strategies based on self-supervised features and provide a detailed comparison of experimental results across a wide range of tasks, datasets, and evaluation metrics. Sotirios Papadopoulos, Emmanouil Patsiouras, Konstantinos Ioannidis, Stefanos Vrochidis, Ioannis Kompatsiaris, Ioannis Patras |
Image Vis. Comput. | 3 |
| 2025 | Vision-Language Pretraining for Variable-Shot Image Classification
Sotirios Papadopoulos, Konstantinos Ioannidis, Stefanos Vrochidis, Ioannis Kompatsiaris, Ioannis Patras |
MMM (4) | 2 |
| 2024 | Enhanced Defect Detection in Airport Runway Infrastructure Using Image-Text PairingabstractMaintaining runway infrastructure is vital for air transport safety, with defects like cracks and tire marks posing significant risks to take-off and landing. Researchers have proposed various methods for automatic surface defect detection using computer vision and machine learning. However, they often require explicitly annotated datasets that demand significant workload and field expertise. Additionally, detection outcomes usually follow the low-level training labels scheme to describe the detected defects, requiring post-processing for high-level semantic information extraction such as damage severity. In this work, we present a novel method for defect detection and damage severity estimation on runway surfaces, leveraging the Contrastive Language-Image Pre-training (CLIP) architecture for imagetext pairing. Our model processes runway images and attaches text descriptions mentioning detected defects and severity level, identifying three defect types (crack, joint, and tire mark) and categorizing damage severity into three levels (low, medium, and high). Utilizing natural language annotations simplifies the labeling process, eliminating the need for labor-intensive lowlevel image-based annotations. The model exploits the natural language labels for direct estimation of damage severity and delivers high-level semantic information to the end-user as text, providing a comprehensive runway condition assessment tool. The proposed method demonstrates high performance across various test sets, posing a valuable human-centric approach for efficient defect detection and damage estimation on runway surfaces. Marios Krestenitis, E. V. Badeka, Ilias Koulalis, Konstantinos Ioannidis, Stefanos Vrochidis |
CBMI | 4 |
| 2024 | A Framework for Vision-Based 3D Inspections for Maintenance Activities and Digital Twin IntegrationabstractVision-based monitoring methods have been actively studied in the construction industry as they can automatically generate information related to progress, productivity, and safety. 3D reconstruction is key in such monitoring techniques, allowing the inference of job-site context, the creation of digital counterparts of physical spaces, and the comparisons between asdesigned and as-built conditions. However, 3D applications in construction currently produce large volumes of unstructured data and unusable point clouds, which are time-consuming to convert into an interactive environment for Building Information Modelling (BIM) or Digital Twins. While radiance field rendering methods are increasingly gaining traction, the adoption of generated Neural Radiance Fields or Gaussian Splatting models by digital construction technology is still tentative. This study introduces a framework that uses Neural Radiance Fields (NeRFs) to improve 3D inspection and maintenance on construction sites. It merges NeRF's high-resolution, real-time 3D modelling with an interactive platform, facilitating detailed remote site analysis and defect detection. The framework incorporates a custom version of TurboNeRF, tailored specifically for construction site inspections. Through this paper, we aim to highlight the potential of combining 3D imaging technology, the use of drone imagery and ontological models to improve construction site management practices. Panagiotis Vrachnos, Carlos Ramonell, Ilias Koulalis, Konstantinos Ioannidis, Irina Stipanovic, Stefanos Vrochidis |
CBMI | 4 |
| 2024 | Incorporating Social Media Sensing and Computer Vision Technologies to Support Wildfire MonitoringabstractSocial media have evolved into a major source of communication and information sharing, and gradually become impactful in monitoring natural disasters such as wildfires, complementing traditional wildfire monitoring technologies. This paper proposes a comprehensive social media-sensing framework for early wildfire detection, encompassing functionalities like social media crawling, visual analytics, and geolocation, for the analysis of social media posts from the X (former Twitter) platform. Upon analysis, a fire event detection module clusters collected posts into fire events, generating relevant alerts. The framework synergizes with a computer vision algorithm, based on a YOLOv8 architecture, performing object detection on UAV imagery for the detection of individuals in danger in affected areas. The collaborative utilization of social media data and UAV imagery improves situational awareness, by providing information for both the fire incidents and the affected subjects in the area, allowing for a more informed decision making. Emmanouil Michail, Aristeidis Bozas, Dimitrios Stefanopoulos, Stavros Paspalakis, Georgios Orfanidis, Anastasia Moumtzidou, Ilias Gialampoukidis, Konstantinos Ioannidis, Stefanos Vrochidis, Ioannis Kompatsiaris |
IGARSS | 8 |
| 2024 | A Framework for 3D Modeling of Construction Sites Using Aerial Imagery and Semantic NeRFs
Panagiotis Vrachnos, Marios Krestenitis, Ilias Koulalis, Konstantinos Ioannidis, Stefanos Vrochidis |
MMM (4) | 4 |
| 2023 | Tweaking EfficientDet for frugal trainingabstractObject detection appears to be omnipresent nowadays with detectors being available for every problem available, covering solutions from extra-light to ultra resource demanding models. Yet, the vast majority of these approaches are based on large datasets to provide the required feature diversity. This work focuses on object detection solutions which do not rely heavily on abundant training datasets but rather on medium-sized data collections. It uses Efficientdet object detector as base for the application of novel modifications which achieve better performance both in efficiency as well in effectiveness. The focus on medium-sized datasets aim at representing more commonplace datasets which can be accumulated and compiled with relative ease. Georgios Orfanidis, Konstantinos Ioannidis, Anastasios Tefas, Stefanos Vrochidis, Ioannis Kompatsiaris |
ICMR | 2 |
| 2023 | Low-Light Image Enhancement Based on U-Net and Haar Wavelet Pooling
Elissavet Batziou, Konstantinos Ioannidis, Ioannis Patras, Stefanos Vrochidis, Ioannis Kompatsiaris |
MMM (2) | 2 |
| 2023 | Comparison of Deep Learning Techniques for Video-Based Automatic Recognition of Greek Folk Dances
Georgios Loupas, Theodora Pistola, Sotiris Diplaris, Konstantinos Ioannidis, Stefanos Vrochidis, Ioannis Kompatsiaris |
MMM (2) | 4 |
| 2023 | Graph-Based Data Association in Multiple Object Tracking: A Survey
Despoina Touska, Konstantinos Gkountakos, Theodora Tsikrika, Konstantinos Ioannidis, Stefanos Vrochidis, Ioannis Kompatsiaris |
MMM (2) | 4 |
| 2023 | Artistic neural style transfer using CycleGAN and FABEMD by adaptive information selection
Elissavet Batziou, Konstantinos Ioannidis, Ioannis Patras, Stefanos Vrochidis, Ioannis Kompatsiaris |
Pattern Recognit. Lett. | 2 |
| 2022 | Sentiment analysis on 2D images of urban and indoor spaces using deep learning architecturesabstractThis paper focuses on the determination of the evoked sentiments to people by observing outdoor and indoor spaces, aiming to create a tool for designers and architects that can be utilized for sophisticated designs. Since sentiment is subjective, the design process can be facilitated by an ancillary automated tool for sentiment extraction. Simultaneously, a dataset containing both real and virtual images of vacant architectural spaces is introduced, while the SUN attributes are also extracted from the images in order to be included throughout training. The dataset is annotated towards both valence and arousal, while five established and two custom architectures, one which has never been used before in classifying abstract concepts, are evaluated on the collected data. Konstantinos Chatzistavros, Theodora Pistola, Sotiris Diplaris, Konstantinos Ioannidis, Stefanos Vrochidis, Ioannis Kompatsiaris |
CBMI | 4 |
| 2022 | A survey for image based methods in construction: from images to digital twinsabstractIn the construction domain, Digital twins are mostly used for facilities management of buildings, but their applications are still very limited. The virtualization of buildings and bridges in the last 15 years in the form of Building or Bridge Information Models is clearly identified as the starting point for the DTs. The industry has erected a frame with semantically rich 3D reference models that are now heavily enriched with visual sensor data captured on construction sites. This article provides an overview of the research and current practices of computer vision methods in the construction industry and presents typical examples of their applications for 3D reconstruction, safety management and structural monitoring for quality assurance. It then highlights the dominant achievements presented in the literature and concludes with the challenges and research directions applicable to digital twins that need to be addressed and exploited in the future. Ilias Koulalis, Nikolaos I. Dourvas, Theocharis Triantafyllidis, Konstantinos Ioannidis, Stefanos Vrochidis, Ioannis Kompatsiaris |
CBMI | 4 |
| 2022 | Vehicle Color Identification Framework using Pixel-level Color Estimation from Segmentation Masks of Car PartsabstractColor comprises one of the most significant and dominant cues for various applications. As one of the most noticeable and stable attributes of vehicles, color can constitute a valuable key component in several practices of intelligent surveillance systems. In this paper, we propose a deep-learning-based framework that combines semantic segmentation masks with pixels clustering for automatic vehicle color recognition. Different from conventional methods, which usually consider only the features of the vehicle's front side, the proposed algorithm is able for view-independent color identification, which is more effective for the surveillance tasks. To the best of our knowledge, this is the first work that employs semantic segmentation masks along with color clustering for the extraction of the vehicle's color representative parts and the recognition of the dominant color, respectively. To evaluate the performance of the proposed method, we introduce a challenging multi-view dataset of 500 car-related RGB images extending the publicly available DSMLR Car Parts dataset for vehicle parts segmentation. The experiments demonstrate that the proposed approach achieves excellent performance and accurate results reaching an accuracy of 93.06% in the multi-view scenario. To facilitate further research, the evaluation dataset and the pre-trained models will be released at https://github.com/klearchos-stav/vehicle_color_recognition. Klearchos Stavrothanasopoulos, Konstantinos Gkountakos, Konstantinos Ioannidis, Theodora Tsikrika, Stefanos Vrochidis, Ioannis Kompatsiaris |
IPAS | 3 |
| 2022 | Automatic Visual Recognition of Unexploded Ordnances Using Supervised Deep LearningabstractUnexploded Ordnance (UXO) classification is a challenging task which is currently tackled using electromagnetic induction devices that are expensive and may require physical presence in potentially hazardous environments. The limited availability of open UXO data has, until now, impeded the progress of image-based UXO classification, which may offer a safe alternative at a reduced cost. In addition, the existing sporadic efforts focus mainly on small scale experiments using only a subset of common UXO categories. Our work aims to stimulate research interest in image-based UXO classification, with the curation of a novel dataset that consists of over 10000 annotated images from eight major UXO categories. Through extensive experimentation with supervised deep learning we uncover key insights into the challenging aspects of this task. Finally, we set the baseline on our novel benchmark by training state-of-the-art Convolutional Neural Networks and a Vision Transformer that are able to discriminate between highly overlapping UXO categories with 84.33% accuracy. Georgios Begkas, Panagiotis Giannakeris, Konstantinos Ioannidis, George Kalpakis, Theodora Tsikrika, Stefanos Vrochidis, Ioannis Kompatsiaris |
ICMR | 3 |
| 2021 | Crowd Violence Detection from Video FootageabstractSurveillance systems currently deploy a variety of devices that can capture visual content (such as CCTV, body-worn cameras, and smartphone cameras), thus rendering the monitoring of video footage obtained from multiple such devices a complex task. This becomes especially challenging when monitoring social events that involve large crowds, particularly when there is a risk of crowd violence. This paper presents and demonstrates a crowd violence detection system that can process, analyze, and alert potential stakeholders, when violence-related content is identified in crowd-based video footage. Based on deep neural networks, the proposed end-to-end framework utilizes a 3D Convolutional Neural Network (CNN) to deal with the (near) real-time analysis of video streams and video files for crowd violence detection. The framework is trained, evaluated, and demonstrated using the Violent Flows dataset, a dataset related to crowd violence that is widely used for research. The presented framework is provided as a standalone application for desktop environments and can analyze both video streams and video files. Konstantinos Gkountakos, Konstantinos Ioannidis, Theodora Tsikrika, Stefanos Vrochidis, Ioannis Kompatsiaris |
CBMI | 2 |
| 2021 | Spatio-Temporal Activity Detection and Recognition in Untrimmed Surveillance VideosabstractThis work presents a spatio-temporal activity detection and recognition framework for untrimmed surveillance videos consisting of a three-step pipeline: object detection, tracking, and activity recognition. The framework relies on the YOLO v4 architecture for object detection, Euclidean distance for tracking, while the activity recognizer uses a 3D Convolutional Deep learning architecture employing spatio-temporal boundaries and addressing it as multi-label classification. The evaluation experiments on the VIRAT dataset achieve accurate detections of the temporal boundaries and recognitions of activities in untrimmed videos, with better performance for the multi-label compared to the multi-class activity recognition. Konstantinos Gkountakos, Despoina Touska, Konstantinos Ioannidis, Theodora Tsikrika, Stefanos Vrochidis, Ioannis Kompatsiaris |
ICMR | 3 |
| 2021 | Fusion of Multimodal Sensor Data for Effective Human Action Recognition in the Service of Medical Platforms
Panagiotis Giannakeris, Athina Tsanousa, Thanassis Mavropoulos, Georgios Meditskos, Konstantinos Ioannidis, Stefanos Vrochidis, Ioannis Kompatsiaris |
MMM (2) | 5 |
| 2021 | Smart integration of sensors, computer vision and knowledge representation for intelligent monitoring and verbal human-computer interaction
Thanassis Mavropoulos, Spyridon Symeonidis, Athina Tsanousa, Panagiotis Giannakeris, Maria Rousi, Eleni Kamateri, Georgios Meditskos, Konstantinos Ioannidis, Stefanos Vrochidis, Ioannis Kompatsiaris |
J. Intell. Inf. Syst. | 8 |
| 2020 | Cycle-Consistent Adversarial Networks and Fast Adaptive Bi-dimensional Empirical Mode Decomposition for Style TransferabstractRecently, research endeavors have shown the potentiality of Cycle-Consistent Adversarial Networks (CycleGAN) in style transfer. In Cycle-Consistent Adversarial Networks, the consistency loss is introduced to measure the difference between the original images and the reconstructed in both directions, forward and backward. In this work, the combination of Cycle-Consistent Adversarial Networks with Fast and Adaptive Bidimensional Empirical Mode Decomposition (FABEMD) is proposed to perform style transfer on images. In the proposed approach the cycle-consistency loss is modified to include the differences between the extracted Intrinsic Mode Functions (BIMFs) images. Instead of an estimation of pixel-to-pixel difference between the produced and input images, the FABEMD is applied and the extracted BIMFs are involved in the computation of the total cycle loss. This method enriches the computation of the total loss in a content-to-content and style-to-style comparison by connecting the spatial information to the frequency components. The experimental results reveal that the proposed method is efficient and produces qualitative results comparable to state-of-the-art methods. Elissavet Batziou, Petros Alvanitopoulos, Konstantinos Ioannidis, Ioannis Patras, Stefanos Vrochidis, Ioannis Kompatsiaris |
ICPR | 3 |
| 2020 | A modified Single-Shot multibox Detector for beyond Real-Time Object DetectionabstractThis works focuses on examining the performance of the Single Shot Detector (SSD) model in resource restricted systems where maintaining the power of the full model comprises a significant prerequisite. The proposed SSD variations examine the behavior of lighter versions of SSD while propose measures to limit the unavoidable performance shortage. The outcomes of the conducted research demonstrate a remarkable trade-off between performance losses, speed improvement and the required resource reservation. Thus, the experimental results evidence the efficiency of the presented SSD alterations towards accomplishing higher frame rates and retaining the performance of the original model. Georgios Orfanidis, Konstantinos Ioannidis, Stefanos Vrochidis, Anastasios Tefas, Ioannis Kompatsiaris |
ICPR | 2 |
| 2020 | A Crowd Analysis Framework for Detecting Violence ScenesabstractThis work examines violence detection in video scenes of crowds and proposes a crowd violence detection framework based on a 3D convolutional deep learning architecture, the 3D-ResNet model with 50 layers. The proposed framework is evaluated on the Violent Flows dataset against several state-of-the-art approaches and achieves higher accuracy values in almost all cases, while also performing the violence detection activities in (near) real-time. Konstantinos Gkountakos, Konstantinos Ioannidis, Theodora Tsikrika, Stefanos Vrochidis, Ioannis Kompatsiaris |
ICMR | 2 |
| 2019 | Autonomous Swarm of Heterogeneous Robots for Surveillance Operations
Georgios Orfanidis, Savvas A. Apostolidis, Athanasios Ch. Kapoutsis, Konstantinos Ioannidis, Elias B. Kosmatopoulos, Stefanos Vrochidis, Ioannis Kompatsiaris |
ICVS | 4 |
| 2019 | Early Identification of Oil Spills in Satellite Images Using Deep CNNs
Marios Krestenitis, Georgios Orfanidis, Konstantinos Ioannidis, Konstantinos Avgerinakis, Stefanos Vrochidis, Ioannis Kompatsiaris |
MMM (1) | 3 |
| 2019 | Real-Time Active SLAM and Obstacle Avoidance for an Autonomous Robot Based on Stereo VisionabstractIn this article, the problem of real-time robot exploration and map building (active SLAM) is considered. A single stereo vision camera is exploited by a fully autonomous robot to navigate, localize itself, define its surroundings, and avoid any possible obstacle in the aim of maximizing the mapped region following the optimal route. A modified version of the so-called cognitive-based adaptive optimization algorithm is introduced for the robot to successfully complete its tasks in real time and avoid any local minima entrapment. The method’s effectiveness and performance were tested under various simulation environments as well as real unknown areas with the use of properly equipped robots. Vicky Kalogeiton, Konstantinos Ioannidis, Georgios Ch. Sirakoulis, Elias B. Kosmatopoulos |
Cybern. Syst. | 2 |
| 2018 | A Deep Neural Network for Oil Spill Semantic Segmentation in Sar ImagesabstractOil spills pose a major threat of the oceanic and coastal environments, hence, an automatic detection and a continuous monitoring system comprises an appealing option for minimizing the response time of relevant operations. Numerous efforts have been conducted towards such solutions by exploiting a variety of sensing systems such as satellite Synthetic Aperture Radar (SAR) which can identify oil spills over sea surfaces in any environmental conditions and operational time. Such approaches include the use of artificial neural networks which effectively identify the polluted areas. Considering their remarkable abilities in many applications, deep Convolutional Neural Networks (DCNN) could surpass limitations and performances of previously proposed methods. This paper describes the application of an approach that combines the merits of a DCNN with SAR imagery in order to provide a fully automated oil spill identification system. The model semantically segments the input SAR images into multiple areas of interest. The deployed DCNN was trained using multiple SAR images acquired from the Sentinel-1 satellite provided by ESA and based on EMSA records for maritime pollution events. Experiments on such challenging benchmark datasets for such an abstract problem demonstrate that the algorithm can accurately identify oil spills leading to an effective detection solution. Georgios Orfanidis, Konstantinos Ioannidis, Konstantinos Avgerinakis, Stefanos Vrochidis, Ioannis Kompatsiaris |
ICIP | 2 |