Ioannis Pitas

dblp:p/IPitas · DBLP profile ↗
← Back
429ranked-venue papers
26as first author
55since 2021 · last 2026
0009-0006-7555-8641ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 306 · 17 first-author · 34 since 2021Artificial intelligence and machine learning · 100 · 5 first-author · 15 since 2021Systems, architecture and hardware · 12 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 12 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 1 first-author · 3 since 2021Computer networks · 6 · 5 since 2021Security and privacy · 4Human-computer interaction and ubiquitous computing · 4 · 1 first-author
YearPublicationVenuePosition
2026 MEWS: Semantic image segmentation with multiclass extreme weak supervision
abstract
Unsupervised image segmentation methods typically assume zero a-priori knowledge about the data semantics. This assumption does hold in many practical scenarios where, although the training data might not be annotated, the target semantic image region classes are known. In these settings, text-driven prompting methods for semantic image segmentation offer noticeable improvement in segmentation accuracy over purely unsupervised approaches. However, such approaches are still limited by: a) inherent text-prompt semantic ambiguity, b) ineffective adaptation to target domain distributions, and c) excessive computational and architectural complexity. To address these shortcomings, we propose the Multiclass Extreme Weak Supervision (MEWS) framework for semantic image segmentation. MEWS assumes the availability of extremely few class-based pixel-level image annotations, e.g., few annotated image pixels per class in very few training images. Such pixel-based image prompts are thereby employed to form image region class prototypes. They can be used to leverage low-complexity unsupervised image segmentation architectures to be trained by our novel prototype-based triplet loss that learns discriminative image features by promoting intra-class image feature compactness while enforcing inter-class feature vector separation. Consequently, the proposed MEWS image segmentation architecture leads to increased weakly supervised training efficiency, bridging the performance gap between supervised and unsupervised image segmentation methods. Our experimental results indicate that the proposed methods compare favorably against text-based prompting image segmentation methods. It yields superior image segmentation accuracy in publicly available image segmentation datasets (e.g., Cityscapes), as well as in Natural Disaster Management (NDM) ones. • A novel Multiclass Extreme Weakly Supervised (MEWS) semantic segmentation framework is proposed that generalizes the original binary EWS DNN architecture by utilizing only sparse, per-class, few-pixel per class labelling. • A class prototype-based triplet loss function is designed that pulls same-class prototype feature vectors together, while pushing mean prototype feature vectors belonging to different classes apart. • A multiclass dynamic thresholding mechanism improves contrastive learning, without additional supervision or manual hyperparameter tuning. • MEWS segmentation excels in accuracy, scaling with annotation, ablation study on loss, validated on NDM Sardinia Wildfire dataset.
A. Apostolidis, Vasileios Mygdalis, Matthaios Dimitrios Tzimas, Ioannis Pitas
Neurocomputing4
2026 Federated unsupervised semantic segmentation
abstract
This work explores the application of Federated Learning (FL) to Unsupervised Semantic image Segmentation (USS). Recent USS methods extract pixel-level features using frozen visual foundation models and refine them through self-supervised objectives that encourage semantic grouping. These features are then grouped to semantic clusters to produce segmentation masks. Extending these ideas to federated settings requires feature representation and cluster centroid alignment across distributed clients, an inherently difficult task under heterogeneous data distributions in the absence of supervision. To address this, we propose FUSS ( F ederated U nsupervised image S emantic S egmentation) which is, to our knowledge, the first framework to enable fully decentralized, label-free semantic segmentation training. FUSS introduces novel federation strategies that promote global consistency in feature and prototype space, jointly optimizing local segmentation heads and shared semantic centroids. Experiments on both benchmark and real-world datasets, including binary and multi-class segmentation tasks, show that FUSS consistently outperforms local-only client trainings as well as extensions of classical FL algorithms under varying client data distributions. To fully support reproducibility, the source code, data partitioning scripts, and implementation details are publicly available at: https://github.com/evanchar/FUSS • Problem definition of Federated Unsupervised Semantic Segmentation (FUSS). • Various federated aggregation strategies are examined. • FedCC: Novel prototype alignment strategies for heterogeneous clients. • Experimental results show improved performance over federated baselines.
Evangelos Charalampakis, Vasileios Mygdalis, Ioannis Pitas
Neurocomputing3
2026 CCBIR: A concept-based system for geospatial image retrieval
abstract
Content Based Image Retrieval (CBIR) systems play a crucial role in efficiently organizing large image datasets for efficient image retrieval based on visual content. In this work, we focus on Natural Disaster Management (NDM) scenarios, where rapid and accurate image retrieval can assist in monitoring disaster events and supporting effective emergency response. We present a novel Concept-Based CBIR (CCBIR) pipeline designed to enhance the CBIR process by integrating semantic image concepts derived from image features. The pipeline extracts features using a task-specific DNN backbone and applies Non-Negative Matrix Factorization (NMF) to derive semantic image concepts. For efficient image retrieval based on concept based image similarity, the FAISS library is leveraged, which enables fast approximate nearest neighbor search in high-dimensional spaces. By transforming complex, high-dimensional features into a structured and interpretable concept space, the proposed CCBIR approach facilitates a more precise alignment between computational image representations and human semantic understanding. The pipeline has been evaluated on three datasets featuring different CBIR problems in natural disaster scenarios such as forest fires and floods, comparing favorably against state-of-the-art CBIR methods. This kind of performance showcases the effectiveness and adaptability of the CCBIR system, highlighting its potential to improve disaster response and recovery operations. • A Concept-Based CBIR (CCBIR) pipeline for content-based image retrieval is proposed. • Interpretable image concepts are extracted using Non-Negative Matrix Factorization. • Our method improves retrieval accuracy and interpretability compared to state-of-the-art. • The pipeline achieves superior performance in flood and wildfire retrieval tasks.
Evgenios Vlachos, Ioannis Pitas
Inf. Sci.2
2026 Discriminative feature distillation via correlation alignment and triplet constraints for image retrieval in wildfire scenarios
abstract
Accurate and efficient image retrieval is vital for wildfire monitoring and Natural Disaster Management (NDM). During an outbreak, new image data is captured and requires to be processed very fast to enable rapid localization of wildfire outbreaks, smoke plumes, and burned regions, as well as assessment of fire severity. However, existing retrieval systems rarely exploit spatial correlation information within image feature maps—an important cue for capturing fine-grained structural patterns in aerial and satellite imagery. In this work, we propose a correlation alignment knowledge distillation framework tailored for disaster image retrieval. Our approach transfers both semantic and structural knowledge from high-capacity teacher models—including large vision-language and self-supervised foundation models such as CLIP and DINO—to lightweight student networks optimized for deployment. By combining triplet-based metric learning with spatial correlation alignment, the proposed method produces discriminative and retrieval-effective features while significantly reducing DNN model complexity. We evaluate our approach on wildfire imagery and extended large-scale retrieval datasets that simulate real-world conditions, including class imbalance and high intra-class variability. Results demonstrate that our method maintains strong retrieval performance under resource-limited conditions, making it suitable for operational use in natural disaster response systems.
Ioanna Valsamara, Ioannis Pitas
Image Vis. Comput.2
2026 Extreme weakly supervised binary semantic image segmentation via one-pixel supervision
abstract
Despite recent advancements, Unsupervised Semantic Segmentation (USS) methods still exhibit a significant performance deficit compared to supervised approaches, particularly in binary semantic segmentation. This limitation arises because, without supervision, USS methods struggle to distinguish foreground from background image regions, particularly when the foreground contains small or uncommon objects. This issue is addressed by our proposed Extremely Weakly Supervised Binary Semantic Segmentation (EWS) framework. EWS expects minimal supervision, consisting only of a small set of one-pixel annotations explicitly belonging to the foreground class across the entire image dataset. Our approach leverages these one-pixel annotations and employs two contrastive losses to map visual transformer features into well-separated foreground and background feature clusters. Additionally, we propose a novel loss function to eliminate the need for hyperparameter tuning of the contrastive loss threshold, by dynamically computing it based on the similarity between the input image features. Even if we employ a single one-pixel annotation, EWS achieves competitive results in binary segmentation tasks while maintaining low computational costs, making it an efficient solution for critical segmentation applications. GitHub Repo: https://github.com/matJTzimas/EWS
Matthaios Dimitrios Tzimas, Vasileios Mygdalis, Christos Papaioannidis, Ioannis Pitas
Pattern Recognit.4
2026 RoboFireFuseNet: Robust fusion of visible and infrared wildfire imaging for real-time flame and smoke segmentation
abstract
Concurrent flame and smoke image region segmentation is a challenging task, particularly when relying on a single imaging modality. Leveraging the combination of visible (RGB) and infrared (IR) modalities in wildfire imaging significantly enhances the accuracy and robustness of fire segmentation. In particular, during dense wildfire smoke incidents, certain image features are only imaged by one modality. Therefore, the two wildfire imaging modalities are inherently complementary. This paper evaluates the effectiveness of RGB and IR image fusion for flame and smoke region segmentation. A novel intermediate image fusion architecture is proposed, built upon a real-time, state-of-the-art DNN semantic segmentation model, augmented with attention mechanisms that promote efficient image modality fusion. Furthermore, a U-Net-like decoder enables accurate spatial reconstruction of the lower-dimensional encoded features. Practical challenges, such as segmentation robustness in the absence of image registration and sensor failures, are also efficiently addressed. Based on our experiments, the proposed DNN segmentation model greatly outperforms existing multimodal DNN architectures in wildfire scenarios in terms of accuracy, while also comparing favorably to stateof-the-art semantic image region segmentation architectures in general urban datasets. Its real-time capabilities and enhanced robustness render it suitable for robotic applications in dynamic, high-stakes segmentation tasks. The code is available at https://gitfront.io/r/dfotiou/eiTd3o9UURjn/RoboFireFuseNetprivate/.
Dimitrios Fotiou, Vasileios Mygdalis, Ioannis Pitas
Pattern Recognit. Lett.3
2025 Cloud Learning-by-Education Node Community (C-LENC) Framework
Nick Tzavidas, Anestis Kaimakamidis, Ioannis Pitas
IEEE Big Data3
2025 Enhanced 3D Pipeline Installation Reconstruction and Modeling for Industrial Inspection
Ioannis Karantaidis, Alexandros Zamioudis, Ioannis Pitas
CIARP3
2025 Padnet: a Patch-Based Anomaly Detection Framework for Industrial Pipeline Damage Detection
abstract
Industrial pipeline inspection in petrochemical refineries is dangerous, expensive, time-consuming and prone to errors. Anomaly detection can play a crucial role towards its automation. Damages in this type of infrastructure are few and can be considered as anomalies (essentially outliers). This paper proposes a novel patch-based Anomaly Detection Network (PADNet), that employs deep learning for detecting insulated pipe damages. It consists of three main components: a) a pipeline segmentation module, b) an image patch proposal module, and c) an anomaly detection module. These components work sequentially first to localize insulated pipelines in the input UAV or ground camera images or video frames and then analyze image patches to detect and localize any damages. Importantly, the anomaly detection module can be trained using undamaged pipeline image data only, hence eliminating the need for costly damaged pipeline image annotation. Experimental results demonstrate the effectiveness of the proposed PADNet method in detecting pipeline damages, making it a promising solution for autonomous industrial infrastructure inspection.
Erofili Alexaki, Christos Papaioannidis, Vasileios Mygdalis, Ioannis Pitas
ICASSP4
2025 Divide-and-Summarize: Enhancing Deep Neural Video Summarization
abstract
Sequence-based neural architectures, such as Long Short-Term Memory (LSTM) networks and Transformers, have driven advances in supervised video summarization by modeling inter-frame dependencies. However, existing methods assume that long-range dependencies are essential for summary generation, which may lead to unnecessary computational overhead. To address this, we propose a Field of View (FOV) adjustment strategy, Divide-and-Summarize (DIV-SUM). By partitioning input videos into smaller fragments of predefined size, our approach explicitly models short-range inter-frame relationships, enabling a fully parallelizable end-to-end video summarization pipeline. Furthermore, most prior work formulates neural video summarization as a frame-wise score regression task. We introduce a simple yet effective target space quantization module, which discretizes the regression targets into classes, introducing a tolerance margin that improves performance. Our approach offers two key benefits: (1) we achieve state-of-the-art performance on the SumMe benchmark while remaining competitive on TVSum, and (2) we significantly reduce the computational cost of inference, improving efficiency without sacrificing quality.
Evangelos Charalampakis, Christos Papaioannidis, Ioannis Pitas
ICIP3
2025 Improve Real-Time Flood Segmentation by Encoding and Distilling Foreground Information
abstract
Flood segmentation systems play a crucial role in natural disaster management, particularly for real-time flood monitoring, thus real-time lightweight deep neural network (DNN) models represent the state-of-the-art (SOTA) solution. A neglected aspect during the design of such solutions is that flood segmentation is a computer vision problem where the variance of visual appearance between the foreground (flood) and the background is imbalanced. This paper mitigates this imbalance using Knowledge Distillation (KD), enhancing the performance of real-time SOTA DNN models for flood segmentation in complex and challenging environments. The proposed method employs a Self-KD approach, where a Teacher model, trained on augmented inputs with reduced background variance by exploiting traditional image processing techniques (e.g., blur-ring), guides a Student model operating on real-world data. By consistently processing augmented inputs, the Teacher model facilitates the Student’s ability to learn robust representations, effectively suppressing noisy background elements. Experimental results on a flood dataset demonstrate an improvement of up to 2.5% in mean Intersection over Union (mIoU) over baseline SOTA models which scored 85% mIoU, highlighting the effectiveness of the proposed method. Furthermore, our approach is model-agnostic, consistently improving the performance of various SOTA DNN architectures across different models.
Pantelis Mentesidis, Vasileios Mygdalis, Ioannis Pitas
ICIP3
2025 Blaze: A Dataset For Wildfire And Burnt Area UAV Image Classification And Segmentation
abstract
The easy and cost-effective deployment of Unmanned Aerial Vehicles (UAVs) that can fly over an area, and collect data for wildfire detection/assessment or for determining the extent of the disaster (e.g., measuring the area of burnt forest regions) using Deep Neural Network (DNN) algorithms, highlights the importance of UAVs and Artificial Intelligence (AI) algorithms for Natural Disaster Management (NDM). However, there is a limited availability of properly annotated data for the abovementioned tasks, which are typically necessary for training accurate and reliable DNN-based AI algorithms. In this direction, this paper introduces the BLAZE dataset, comprising approximately 5.4K annotated RGB images depicting both urban and non-urban areas before, during and after a wildfire. Moreover, using the proposed dataset, several baseline DNN-based algorithms have been trained and evaluated for the wildfire image classification and burnt areas segmentation tasks. Experimental results show that increased accuracy can be achieved for both tasks, thus proving the usefulness of the developed BLAZE dataset in the NDM domain. Data can be found here https://aiia.csd.auth.gr/blaze-fire-classification-segmentation-dataset/.
Michael Siavrakas, Christos Papaioannidis, Ioannis Pitas
ICIP3
2025 A Weighting Loss Approach for Transformer-Based Object Detection
abstract
This paper introduces a training loss function tailored for object detection in transformer-based architectures. Our approach addresses the imbalance in ground-truth bounding box sizes during training by implementing a coordinate-based error-weighting mechanism for the L1loss. This modification stabilizes optimization and enhances detection performance, particularly in detection problems requiring bounding boxes of varying sizes within the same image, such as fire/smoke detection applications. By integrating this method into the Real-Time Detection Transformer (RT-DETR), we conduct extensive experiments across three fire/smoke detection datasets and compare our findings against leading real-time object detection algorithms, such as YOLO models. To further validate the generalizability of the proposed loss function, we incorporate it into various DETR-based architectures. Our experiments demonstrate the superior fire detection accuracy of RT-DETR trained with our method across all three datasets while ensuring its effectiveness on more complex datasets. This study not only enhances the capabilities of transformer-based architectures for real-time detection tasks but also contributes to the development of more efficient and reliable fire detection systems.
Matthaios Dimitrios Tzimas, Vasileios Mygdalis, Ioannis Pitas
IJCNN3
2025 A Decentralized Sharding BFT Consensus Approach, for Efficient Decentralized DNN Inference Classification
abstract
The security and trustworthiness of participating DNN nodes are often overlooked during the design of modern Decentralized Deep Neural Networks (D-DNN). This paper introduces a shard-based distributed consensus protocol specifically tailored for DNN nodes operating over unreliable communication links. The proposed approach enhances D-DNN scalability, by enabling D-DNN systems having a large number of DNN nodes. This is achieved through a hierarchical consensus mechanism that partitions the D-DNN network into sub-networks (shards), leveraging Out-of-Distribution (OOD) detectors to localize and isolate the consensus process within each shard. Rather than randomly allocating DNN nodes into shards, the OOD detector can be employed to identify and group nodes with similar domain knowledge. This approach improves the overall D-DNN system robustness, by identifying and isolating malicious DNN nodes or once that have poor performance for a specific DNN task. Experimental results demonstrate improvements in the D-DNN system’s classification accuracy and reliability.
Dimitrios Papaioannou, Vasileios Mygdalis, Ioannis Pitas
ISCC3
2025 Distilling Structural Knowledge: Teaching Representations in Multi-DNN Agent Systems
abstract
Recent advancements in multi-agent systems have highlighted the potential of enabling efficient knowledge exchange among Deep Neural Network (DNN) agents to address complex tasks. This paper investigates the dynamics of DNN teacher-student interactions, with a focus on distilling specific knowledge from teacher DNNs to student DNNs to enhance retrieval performance. We propose an approach that optimizes the feature structure of the student DNN agent, enabling the distillation of detailed representation knowledge. Utilizing the concept of triplets, our method captures data correlations and transfers structural knowledge, aiming to compress the knowledge of representations and their structural data dependencies from larger to smaller DNN agents while preserving performance accuracy. Our triplet-based knowledge distillation strategy guides the student DNN agent to learn optimal representations for image retrieval in a multi-DNN agent system. Experimental results demonstrate enhancements in the student agent’s efficiency, showing improvements in performance across various DNN architectures and datasets.
Ioanna Valsamara, Christos Papaioannidis, Ioannis Pitas
ISCC3
2025 Spatio-temporal invariant descriptors for skeleton-based human action recognition
Aouaidjia Kamel, Chongsheng Zhang, Ioannis Pitas
Inf. Sci.3
2025 FCL-ViT: Task-aware attention tuning for Continual Learning
Anestis Kaimakamidis, Ioannis Pitas
Pattern Recognit. Lett.2
2025 Towards human society-inspired decentralized DNN inference
abstract
In human societies, individuals make their own decisions and they may select if and who may influence it, by e.g., consulting with people of their acquaintance or experts of a field. At a societal level, the overall knowledge is preserved and enhanced by individual person empowerment, where complicated consensus protocols have been developed over time in the form of societal mechanisms to assess, weight, combine and isolate individual people opinions. In distributed machine learning environments however, individual AI agents are merely part of a system where decisions are made in a centralized and aggregated fashion or require a fixed network topology, a practice prone to security risks and collaboration is nearly absent. For instance, Byzantine Failures may tamper both the training and inference stage of individual AI agents, leading to significantly reduced overall system performance. Inspired by societal practices, we propose a decentralized inference strategy where each individual agent is empowered to make their own decisions, by exchanging and aggregating information with other agents in their network. To this end, a “Quality of Inference” consensus protocol (QoI) is proposed, forming a single commonly accepted inference rule applied by every individual agent. The overall system knowledge and decisions on specific manners can thereby be stored by all individual agents in a decentralized fashion, employing e.g., blockchain technology. Our experiments in classification tasks indicate that the proposed approach forms a secure decentralized inference framework, that prevents adversaries at tampering the overall process and achieves comparable performance with centralized decision aggregation methods. • In human societies, individuals make their own decisions taking expert advice into account. • This process is simulated in decentralized DNN Inference settings. • A novel consensus protocol working as a single inference rule was developed. • A fault-tolerant inference architecture in which misbehaving AI agents are penalized.
Dimitrios Papaioannou, Vasileios Mygdalis, Ioannis Pitas
Signal Process. Image Commun.3
2025 LIX: Implicitly Infusing Spatial Geometric Prior Knowledge Into Visual Semantic Segmentation for Autonomous Driving
abstract
Despite the impressive performance achieved by data-fusion networks with duplex encoders for visual semantic segmentation, they become ineffective when spatial geometric data are not available. Implicitly infusing the spatial geometric prior knowledge acquired by a data-fusion teacher network into a single-modal student network is a practical, albeit less explored research avenue. This article delves into this topic and resorts to knowledge distillation approaches to address this problem. We introduce the Learning to Infuse "X" (LIX) framework, with novel contributions in both logit distillation and feature distillation aspects. We present a mathematical proof that underscores the limitation of using a single, fixed weight in decoupled knowledge distillation and introduce a logit-wise dynamic weight controller as a solution to this issue. Furthermore, we develop an adaptively-recalibrated feature distillation algorithm, including two novel techniques: feature recalibration via kernel regression and feature consistency quantification via centered kernel alignment. Extensive experiments conducted with intermediate-fusion and late-fusion networks across various public datasets provide both quantitative and qualitative evaluations, demonstrating the superior performance of our LIX framework when compared to other state-of-the-art approaches. Source code is available at https://mias.group/LIX.
Sicen Guo, Ziwei Long, Ioannis Pitas, Rui Fan 0001
IEEE Trans. Image Process.5
2025 These Maps Are Made by Propagation: Adapting Deep Stereo Networks to Road Scenarios With Decisive Disparity Diffusion
abstract
Stereo matching has emerged as a cost-effective solution for road surface 3D reconstruction, garnering significant attention towards improving both computational efficiency and accuracy. This article introduces decisive disparity diffusion (D3Stereo), marking the first exploration of dense deep feature matching that adapts pre-trained deep convolutional neural networks (DCNNs) to previously unseen road scenarios. A pyramid of cost volumes is initially created using various levels of learned representations. Subsequently, a novel recursive bilateral filtering algorithm is employed to aggregate these costs. A key innovation of D3Stereo lies in its alternating decisive disparity diffusion strategy, wherein intra-scale diffusion is employed to complete sparse disparity images, while inter-scale inheritance provides valuable prior information for higher resolutions. Extensive experiments conducted on our created UDTIRI-Stereo and Stereo-Road datasets underscore the effectiveness of D3Stereo strategy in adapting pre-trained DCNNs and its superior performance compared to all other explicit programming-based algorithms designed specifically for road surface 3D reconstruction. Additional experiments conducted on the Middlebury dataset with backbone DCNNs pre-trained on the ImageNet database further validate the versatility of D3Stereo strategy in tackling general stereo matching problems. Our source code and supplementary material are publicly available at https://mias.group/D3-Stereo.
Yikang Zhang 0001, Ioannis Pitas, Rui Fan 0001
IEEE Trans. Image Process.4
2025 Enhancing visual object tracking robustness through a lightweight denoising module
abstract
Abstract Visual object tracking is crucial for numerous applications ranging from smartphones to autonomous vehicles. However, the impact of input noise on tracking performance remains underexplored. This paper presents a lightweight neural network module designed to enhance the robustness of 2D tracking methods against various types of noise. By performing image-to-image translation, the proposed robust tracking module (RTM) standardizes the operational space of tracking algorithms, thereby improving their resilience. Experimental results on benchmark datasets demonstrate the effectiveness of RTM in mitigating performance degradation caused by noise. Additionally, we introduce an evaluation toolkit that facilitates the assessment of tracking robustness against common noise types. The source code of the proposed method is available at https://github.com/iason1907/RTM .
Iason Karakostas, Vasileios Mygdalis, Nikos Nikolaidis 0001, Ioannis Pitas
Vis. Comput.4
2024 Political Tweet Sentiment Analysis for Public Opinion Polling
abstract
Public opinion measurement through polling is a classical political analysis task, e.g. for predicting national and local election results. However, polls are expensive to run and their results may be biased primarily due to improper population sampling. In this paper, we propose two innovative methods for employing tweet sentiment analysis results for public opinion polling. Our first method utilizes merely the tweet sentiment analysis results outperforming multiple well-recognised methods. In addition, we introduce a novel hybrid way to estimate electorally results from both public opinion polls and tweets. This method enables more accurate, frequent, and inexpensive public opinion estimation and is used for estimating the result of the 2023 Greek national election. Our method demonstrated lower deviation from the actual election’s results than the conventional public opinion polls, introducing new possibilities for public opinion estimation using social media platforms.
Anestis Kaimakamidis, Ioannis Pitas
ICASSP2
2024 A Unified DNN-Based System for Industrial Pipeline Segmentation
abstract
This paper presents a unified system tailored for autonomous pipe segmentation within an industrial setting. To this end, it is designed to analyze RGB images captured by Unmanned Aerial Vehicle (UAV)-mounted cameras to predict binary pipe segmentation maps. The overall proposed system consists of three main components: a) a Convolutional Neural Network (CNN) that is used to obtain initial estimates of the pipe segmentation maps, b) a point extraction module that acts on the outputs of the CNN to propose strong pipe class representatives in the input image space, and c) a foundation segmentation model, utilized to refine the initial estimations based on the proposed pipe class representatives. The architecture of the proposed system was specifically designed to ensure increased generalization ability in different, unknown environments, offering an effective solution to a well-known limitation of typical segmentation CNNs, at least in the pipe segmentation task. The effectiveness of the proposed system in this particular setting is evaluated by utilizing two pipe segmentation datasets, originating from two different industrial sites, which were manually annotated with the corresponding pipe segmentation maps. Experimental results demonstrate that the proposed system outperforms the baseline segmentation CNNs, demonstrating its remarkable generalization capabilities.
Dimitrios Psarras, Christos Papaioannidis, Vasileios Mygdalis, Ioannis Pitas
ICASSP4
2024 Leveraging Collective Knowledge for Forest Fire Classification
abstract
This paper presents a novel Fire Classification Multi-Agent (FCMA) framework that utilizes peer-to-peer learning and distributed learning techniques to disseminate knowledge within the agent community. Furthermore, we define and introduce the architecture of a Deep Neural Network (DNN) agent, which can infinitely interact with other DNN agents and the external environment upon deployment. The FCMA framework is suitable for natural disaster management systems where multiple agents are required to run autonomously and foster the community’s knowledge. The FCMA provides two options for knowledge transfer, a peer-to-peer and a federated one. The experimental results display the effective knowledge transfer using both options and also compare the two options with each other in a forest fire classification setting.
Anestis Kaimakamidis, Ioannis Pitas
ISCC2
2024 Proof of Quality Inference (PoQI): An AI Consensus Protocol for Decentralized DNN Inference Frameworks
abstract
In the realm of machine learning systems, achieving consensus among networking nodes is a fundamental yet challenging task. This paper presents Proof of Quality Inference (PoQI), a novel consensus protocol designed to integrate deep learning inference under the basic format of the Practical Byzantine Fault Tolerant (P-BFT) algorithm. PoQI is applied to Deep Neural Networks (DNNs) to infer the quality and authenticity of produced estimations by evaluating the trustworthiness of the DNN node’s decisions. In this manner, PoQI enables DNN inference nodes to reach a consensus on a common DNN inference history in a fully decentralized fashion, rather than relying on a centralized inference decision-making process. Through PBFT adoption, our method ensures byzantine fault tolerance, permitting DNN nodes to reach an agreement on inference validity swiftly and efficiently. We demonstrate the efficacy of PoQI through theoretical analysis and empirical evaluations, highlighting its potential to forge trust among unreliable DNN nodes.
Dimitrios Papaioannou, Vasileios Mygdalis, Ioannis Pitas
ISCC3
2024 Domain Expertise Assessment for Multi-DNN Agent Systems
abstract
Recently, multi-agent systems that facilitate knowledge sharing among Deep Neural Network (DNN) agents, have gained increasing attention. This paper explores the dynamics of multi-agent systems that support Teacher-Student DNN interactions, where knowledge is distilled from Teachers to Students. Within such systems, selecting the most compatible Teacher for a given task is far from trivial and can lead to low-quality decisions. Hence, the need arises for accurate domain knowledge evaluation. In that context, we propose including an OOD detection module in each DNN agent to enable effective agent expertise evaluation and precise identification of suitable Teachers. This setup allows Student agents to distill knowledge from the most knowledgeable Teachers within a specific domain, ensuring optimal system performance. To effectively utilize OOD detection in this context, we address key challenges such as determining the minimum data cardinality required to ensure optimal performance and reliable inferences of the OOD detectors.
Ioanna Valsamara, Christos Papaioannidis, Ioannis Pitas
ISCC3
2024 Vision-based drone control for autonomous UAV cinematography
Ioannis Mademlis, Charalampos Symeonidis, Anastasios Tefas, Ioannis Pitas
Multim. Tools Appl.4
2023 Facilitating Experimental Reproducibility in Neural Network Research with a Unified Framework
abstract
In the realm of neural network research, achieving experiment reproducibility is paramount for building upon existing knowledge and advancing the field. This paper examines a multi-agent neural network framework on its ability to facilitate the reproduction of experiments. Also, we address the reproducibility problem when there are data or source code limitations. The framework offers crucial functionalities for facilitating experiment reproducibility achieved through data, layer outputs, architectures, and weights exchange among the framework's agents. Through the integration of these functionalities, this framework empowers researchers to reproduce and validate experimental results consistently, fostering a more robust and collaborative research environment in the field of neural networks. The experimental results demonstrate the framework's reproducibility abilities. Furthermore, we test the framework in terms of reproducibility in an emergency natural disaster management situation. Finally, we analyze how the privacy limitations of the original neural network affect the reproducibility results.
Anestis Kaimakamidis, Ioannis Pitas
BDCAT2
2023 Evaluating Deep Neural Network-based Fire Detection for Natural Disaster Management
abstract
Recently, climate change has led to more frequent extreme weather events, introducing new challenges for Natural Disaster Management (NDM) organizations. This fact makes the employment of modern technological tools such as Deep Neural Networks-based fire detectors a necessity, as they can assist such organizations manage these extreme events more effectively. In this work, we argue that the mean Average Precision (mAP) metric that is commonly used to evaluate typical object detection algorithms can not be trusted for the fire detection task, due to its high dependence on the employed data annotation strategy. This means that the mAP score of a fire detection algorithm may be low even when it predicts fire bounding boxes that accurately enclose the depicted fires. In this direction, a new evaluation metric for fire detection is proposed, denoted as Image-level mean Average Precision (ImAP), which reduces the dependence on the bounding box annotation strategy by rewarding/penalizing bounding box predictions on image level, rather than on bounding box level. Experiments using different object detection algorithms have shown that the proposed ImAP metric reveals the true fire detection capabilities of the tested algorithms more effectively.
Matthaios Dimitrios Tzimas, Christos Papaioannidis, Vasileios Mygdalis, Ioannis Pitas
BDCAT4
2023 Exploiting One-Class Classification Optimization Objectives for Increasing Adversarial Robustness
abstract
This work examines the problem of increasing the robustness of deep neural network-based image classification systems to adversarial attacks, without changing the neural architecture or employ adversarial examples in the learning process. We attribute their famous lack of robustness to the geometric properties of the deep neural network embedding space, derived from standard optimization options, which allow minor changes in the intermediate activation values to trigger dramatic changes to the decision values in the final layer. To counteract this effect, we explore optimization criteria that supervise the distribution of the intermediate embedding spaces, in a class-specific basis, by introducing and leveraging one-class classification objectives. The proposed learning procedure compares favorably to recently proposed training schemes for adversarial robustness in black-box adversarial attack settings.
Vasileios Mygdalis, Ioannis Pitas
ICASSP2
2023 Fast Single-Person 2D Human Pose Estimation Using Multi-Task Convolutional Neural Networks
abstract
This paper presents a novel neural module for enhancing existing fast and lightweight 2D human pose estimation CNNs, in order to increase their accuracy. A baseline stem CNN is augmented by a collateral module, which is tasked to encode global spatial and semantic information and provide it to the stem network during inference. The latter one outputs the final 2D human pose estimations. Since global information encoding is an inherent subtask of 2D human pose estimation, this particular setup allows the stem network to better focus on the local details of the input image and on precisely localizing each body joint, thus increasing overall 2D human pose estimation accuracy. Furthermore, the collateral module is designed to be lightweight, adding negligible runtime computational cost, so that the unified architecture retains the fast execution property of the stem network. Evaluation of the proposed method on public 2D human pose estimation datasets shows that it increases the accuracy of different baseline stem CNNs, while outperforming all competing fast 2D human pose estimation methods.
Christos Papaioannidis, Ioannis Mademlis, Ioannis Pitas
ICASSP3
2023 Efficient Feature Extraction for Non-Maximum Suppression in Visual Person Detection
abstract
Non-Maximum Suppression (NMS) is a post-processing step in almost every visual object detector, tasked with rapidly pruning the number of overlapping detected candidate rectangular Regions-of-Interest (RoIs) and replacing them with a single, more spatially accurate detection (in pixel coordinates). The common Greedy NMS algorithm suffers from drawbacks, due to the need for careful manual tuning. In visual person detection, most NMS methods typically suffer when analyzing crowded scenes with high levels of in-between occlusions. This paper proposes a modification on a deep neural architecture for NMS, suitable for such cases and capable of efficiently cooperating with recent neural object detectors. The method approaches the NMS problem as a rescoring task, aiming to ideally assign precisely one detection per object. The proposed modification exploits the extraction of RoI representations, semantically capturing the region’s visual appearance, from information-rich feature maps computed by the detector’s intermediate layers. Experimental evaluation on two common public person detection datasets shows improved accuracy against competing methods, with acceptable inference speed.
Charalampos Symeonidis, Ioannis Mademlis, Ioannis Pitas, Nikos Nikolaidis 0001
ICASSP3
2023 Domain Adaptation in Power Line Segmentation: A New Synthetic Dataset
abstract
Power line segmentation is a critical component of UAV intelligent inspection systems to ensure the safe and reliable operation of power grids. For challenging-to-label tasks like this, simulators can efficiently generate large amounts of labeled data. In this work, a large-scale annotated synthetic power lines dataset generated utilizing the unity game engine and the unity perception package1. To address domain shift between real and synthetic domain, input-level adaptation performed. Additionally, a new power line segmentation loss developed to mitigate the effects of unbalanced pixel distributions among power lines and background. Experiments demonstrate that our approach achieves state-of-the-art performance on power line segmentation task.
Georgios Kalitsios, Vasileios Mygdalis, Ioannis Pitas
ICIP3
2023 Deep Reinforcement Learning with semi-expert distillation for autonomous UAV cinematography
abstract
Unmanned Aerial Vehicles (UAVs, or drones) have revolutionized modern media production. Being rapidly deployable "flying cameras", they can easily capture aesthetically pleasing aerial footage of static or moving filming targets/subjects. Current approaches rely either on manual UAV/gimbal control by human experts, or on a combination of complex computer vision algorithms and hardware configurations for automating the flight+filming process. This paper explores an efficient Deep Reinforcement Learning (DRL) alternative, which implicitly merges the target detection and path planning steps into a single algorithm. To achieve this, a baseline DRL approach is augmented with a novel policy distillation component, which transfers knowledge from a suitable, semi-expert Model Predictive Control (MPC) controller into the DRL agent. Thus, the latter is able to autonomously execute a specific UAV cinematography task with purely visual input. Unlike the MPC controller, the proposed DRL agent does not need to know the 3D world position of the filming target during inference. Experiments conducted in a photorealistic simulator showcase superior performance and training speed compared to the baseline agent, while surpassing the MPC controller in terms of visual occlusion avoidance.
Andreas Sochopoulos, Ioannis Mademlis, Evangelos Charalampakis, Sotirios Papadopoulos, Ioannis Pitas
ICME5
2023 Escaping local minima in deep reinforcement learning for video summarization
abstract
State-of-the-art deep neural unsupervised video summarization methods mostly fall under the adversarial reconstruction framework. This employs a Generative Adversarial Network (GAN) structure and Long Short-Term Memory (LSTM) autoencoders during its training stage. The typical result is a selector LSTM that sequentially receives video frame representations and outputs corresponding scalar importance factors, which are then used to select key-frames. This basic approach has been augmented with an additional Deep Reinforcement Learning (DRL) agent, trained using the Discriminator’s output as a reward, which learns to optimize the selector’s outputs. However, local minima are a well-known problem in DRL. Thus, this paper presents a novel regularizer for escaping local loss minima, in order to improve unsupervised key-frame extraction. It is an additive loss term employed during a second training phase, that rewards the difference of the neural agent’s parameters from those of a previously found good solution. Thus, it encourages the training process to explore more aggressively the parameter space in order to discover a better local loss minimum. Evaluation performed on two public datasets shows considerable increases over the baseline and against the state-of-the-art.
Panagiota Alexoudi, Ioannis Mademlis, Ioannis Pitas
ICMR3
2023 Real-Time Object Geopositioning from Monocular Target Detection/Tracking for Aerial Cinematography
abstract
In recent years, the field of automated aerial cinematography has seen a significant increase in demand for real-time 3D target geopositioning for motion and shot planning. To this end, many of the existing cinematography plans require the use of complex sensors that need to be equipped on the subject or rely on external motion systems. This work addresses this problem by combining monocular visual target detection and tracking with a simple ground intersection model. Under the assumption that the targets to be filmed typically stand on the ground, 3D target localization is achieved by estimating the direction and the norm of the look-at vector. The proposed algorithm employs an error estimation model that accounts for the error in detecting the bounding box, the height estimation errors, and the uncertainties of the pitch and yaw angles. This algorithm has been fully implemented in a heavy-lifting aerial cinematography hexacopter, and its performance has been evaluated through experimental flights. Results show that typical errors are within 5 meters of absolute distance and 3 degrees of angular error for distances to the target of around 100 meters.
Daniel Aláez, Vasileios Mygdalis, Jesús E. Villadangos, Ioannis Pitas
MMSP4
2023 A multiple-UAV architecture for autonomous media production
Ioannis Mademlis, Arturo Torres-González, Jesús Capitán, Maurizio Montagnuolo, Alberto Messina, Fulvio Negro, Cédric Le Barz, Rita Cunha, Bruno J. Guerreiro, Fan Zhang 0017, Stephen Boyle, Gregoire Guerout, Anastasios Tefas, Nikos Nikolaidis 0001, David Bull 0001, Ioannis Pitas
Multim. Tools Appl.17
2023 Fast CNN-Based Single-Person 2D Human Pose Estimation for Autonomous Systems
abstract
This paper presents a novel Convolutional Neural Network (CNN) architecture for 2D human pose estimation from RGB images that balances between high 2D human pose/skeleton estimation accuracy and rapid inference. Thus, it is suitable for safety-critical embedded AI scenarios in autonomous systems, where computational resources are typically limited and fast execution is often required, but accuracy cannot be sacrificed. The architecture is composed of a shared feature extraction backbone and two parallel heads attached on top of it: one for 2D human body joint regression and one for global human body structure modelling through Image-to-Image Translation (I2I). A corresponding multitask loss function allows training of the unified network for both tasks, through combining a typical 2D body joint regression with a novel I2I term. Along with enhanced information flow between the parallel neural heads via skip synapses, this strategy is able to extract both ample semantic and rich spatial information, while using a less complex CNN; thus it permits fast execution. The proposed architecture is evaluated on public 2D human pose estimation datasets, achieving the best accuracy-speed ratio compared to the state-of-the-art. Additionally, it is evaluated on a pedestrian intention recognition task for self-driving cars, leading to increased accuracy and speed in comparison to competing approaches.
Christos Papaioannidis, Ioannis Mademlis, Ioannis Pitas
IEEE Trans. Circuits Syst. Video Technol.3
2023 Neural Attention-Driven Non-Maximum Suppression for Person Detection
abstract
Non-maximum suppression (NMS) is a post-processing step in almost every visual object detector. NMS aims to prune the number of overlapping detected candidate regions-of-interest (RoIs) on an image, in order to assign a single and spatially accurate detection to each object. The default NMS algorithm (GreedyNMS) is fairly simple and suffers from severe drawbacks, due to its need for manual tuning. A typical case of failure with high application relevance is pedestrian/person detection in the presence of occlusions, where GreedyNMS doesn't provide accurate results. This paper proposes an efficient deep neural architecture for NMS in the person detection scenario, by capturing relations of neighboring RoIs and aiming to ideally assign precisely one detection per person. The presented Seq2Seq-NMS architecture assumes a sequence-to-sequence formulation of the NMS problem, exploits the Multihead Scale-Dot Product Attention mechanism and jointly processes both geometric and visual properties of the input candidate RoIs. Thorough experimental evaluation on three public person detection datasets shows favourable results against competing methods, with acceptable inference runtime requirements.
Charalampos Symeonidis, Ioannis Mademlis, Ioannis Pitas, Nikos Nikolaidis 0001
IEEE Trans. Image Process.3
2022 OTE: Optimal Trustworthy EdgeAI solutions for smart cities
abstract
This work studies and defines the problem of providing extensive and opportunistic Edge AI-based area coverage in smart city application scenarios, by researching and determining the optimal configuration of sensing and computational resources for minimizing the environmental/technology footprint of the solution. A typical smart city computing continuum consists of statically installed multimodal sensing Internet-of-Things (IoT) nodes at various city locations, accompanied by interconnected computational Cloud/Edge/IoT nodes. This paper presents Optimal Trustworthy EdgeAI (OTE), an entirely novel research pipeline, that complements existing smart city infrastructure with intelligent drone Edge/IoT nodes (in the form of modularly equipped unmanned aerial vehicles), capable of autonomous repositioning according to individual/collective sensing and coverage criteria. Thereby, we envisage the emerging cutting-edge technologies of trustworthy sensing, perceiving, modelling technologies for predicting the behavior of moving targets (e.g., citizens/vehicles/objects), understanding natural phenomena (e.g., sea wave motion, urban flora/fauna, biodiversity) in order to anticipate events (people's bad habits, environmental changes), by exploiting novel continuous data processing services across the whole span of the enhanced Cloud-Edge-IoT computing continuum.
Vasileios Mygdalis, Lorenzo Carnevale, J. Ramiro Martinez de Dios, Dmitriy Shutin, Giovanni Aiello, Massimo Villari, Ioannis Pitas
CCGRID7
2022 Exploiting Caption Diversity for Unsupervised Video Summarization
abstract
Most unsupervised Deep Neural Networks (DNNs) for video summarization rely on adversarial learning, autoencoding and training without utilizing any ground-truth summary. In several cases, the Convolutional Neural Network (CNN)-derived video frame representations are sequentially fed to a Long Short-Term Memory (LSTM) network, which selects key-frames and, during training, attempts to reconstruct the original/full video from the summary, while confusing an adversarially optimized Discriminator. Additionally, regularizers aiming at maximizing the summary’s visual semantic diversity can be employed, such as the Determinantal Point Process (DPP) loss term. In this paper, a novel DPP-based regularizer is proposed that exploits a pretrained DNN-based image captioner in order to additionally enforce maximal key-frame diversity from the perspective of textual semantic content. Thus, the selected key-frames are encouraged to differ not only with regard to what objects they depict, but also with regard to their textual descriptions, which may additionally capture activities, scene context, etc. Empirical evaluation indicates that the proposed regularizer leads to state-of-the-art performance.
Michail Kaseris, Ioannis Mademlis, Ioannis Pitas
ICASSP3
2022 An Efficient Framework for Human Action Recognition Based on Graph Convolutional Networks
abstract
This paper presents a novel framework for skeleton-based Human Action Recognition (HAR) based on Graph Convolution Networks (GCNs). The proposed framework aims to increase human action recognition performance of GCN-based methods by incorporating a missing-joint-handling pre-processing step and a novel adjacency matrix construction method in a single human action recognition pipeline. The missing-joint-handling pre-processing step is utilized to infer missing data in the input sequence, which may occur due to imperfect skeleton extraction, based on imputation methods. The novel adjacency matrix construction method is executed offline to compute an improved weighted adjacency matrix specifically designed for HAR, which is utilized in every layer of the employed GCN. Moreover, both the pre-processing step and the adjacency construction method can be utilized along with any GCN architecture, allowing any GCN-based HAR method to be employed in the proposed framework. Experimental evaluation on two public datasets indicate favorable human action classification scores compared to the employed baseline and all competing methods both for 2D and 3D skeleton-based human action recognition, while using a GCN architecture with less learnable parameters.
Nikolaos Kilis, Christos Papaioannidis, Ioannis Mademlis, Ioannis Pitas
ICIP4
2022 Fast Semantic Image Segmentation for Autonomous Systems
abstract
Fast semantic image segmentation is crucial for autonomous systems, as it allows an autonomous system (e.g., self-driving car, drone, etc.) to interpret its environment on-the-fly and decide on necessary actions by exploiting dense semantic maps. The speed of semantic segmentation on embedded computational hardware is as important as its accuracy. Thus, this paper proposes a novel framework for semantic image segmentation that is both fast and accurate. It augments existing real-time semantic image segmentation architectures by an auxiliary, parallel neural branch that is tasked to predict semantic maps in an alternative manner by utilizing Generative Adversarial Networks (GANs). Additional attention-based neural synapses linking the two branches allow information to flow between them during both the training and the inference stage. Extensive experiments on three public datasets for autonomous driving and for aerial-perspective image analysis indicate non-negligible gains in segmentation accuracy, without compromises on inference speed.
Christos Papaioannidis, Ioannis Mademlis, Ioannis Pitas
ICIP3
2022 Auth-Persons: A Dataset for Detecting Humans in Crowds from Aerial Views
abstract
Recent advances in artificial intelligence, control and sensing technologies have facilitated the development of autonomous Unmanned Aerial Vehicles (UAVs). Detecting humans from video input captured on-the-fly from UAVs is a critical task for ensuring flight safety, mostly handled with lightweight Deep Neural Networks (DNNs). However the detection of individual people in the case of dense crowds and/or distribution shifts (i.e., significant visual differences between the training and the test sets) is still very challenging. This paper presents AUTH-Persons, a new, annotated, publicly available video dataset, that consists of both real and synthetic footage, suitable for training and evaluating aerial-view person detection algorithms. The synthetic data were collected from 8 visually distinct photorealistic outdoor environments and they mostly contain scenes with crowded areas, where heavy occlusions and high person densities pose challenges to common detectors. This dataset is employed to evaluate the generalization performance of various state-of-the-art detection frameworks, by testing them on environments that are visually distinct from those they have been trained on. Finally, given that Non-Maximum Suppression (NMS) methods at the end of person detection pipelines typically suffer in crowded scenes, the performance of various NMS algorithms is also compared in AUTH-Persons.
Charalampos Symeonidis, Ioannis Mademlis, Ioannis Pitas, Nikos Nikolaidis 0001
ICIP3
2022 Whitening Transformation inspired Self-Attention for Powerline Element Detection
abstract
Powerline inspection operations involve capturing and inspecting visual footage of powerline elements from elevated positions above and around the powerline and are currently performed with the help of helicopters and/or Unmanned Aerial Vehicles (UAVs). Current technological advances in the areas of robotics and machine learning are towards enabling fully autonomous operations. To this end, one of the tasks to be addressed is the robust, precise and fast powerline object detection problem. Recently introduced Transformer-based object detection methods demonstrate time and accuracy advances with respect to previous works. In this work, we present an enhanced Transformer-based architecture that further improves the state-of-the-art by incorporating a content-specific object query generator and by substituting the original attention operation with a whitening-inspired transformation at certain stages of the architecture. We evaluate our method in a recently captured powerline detection dataset and we show that our novel contributions offer a significant boost regarding detection accuracy.
Emmanouil Patsiouras, Vasileios Mygdalis, Ioannis Pitas
ICPR3
2022 Autonomous UAV Cinematography
abstract
The use of camera-equipped Unmanned Aerial Vehicles (UAVs, or "drones") for professional media production is already an exciting commercial reality. Currently available consumer UAVs for cinematography applications are equipped with high-end cameras and a degree of cognitive autonomy relying on artificial intelligence (AI). Current research promises to further exploit the potential of autonomous functionalities in the immediate future, resulting in portable flying robotic cameras with advanced intelligence concerning autonomous landing, subject detection/tracking, cinematic shot execution, 3D localization and environmental mapping, as well as autonomous obstacle avoidance combined with on-line motion re-planning. Disciplines driving this progress are computer vision, machine/deep learning and aerial robotics. This Tutorial emphasizes the definition and formalization of UAV cinematography aesthetic components, as well as the use of robotic planning/control methods for autonomously capturing them on footage, without the need for manual tele-operation. Additionally, it focuses on state-of-the-art Imitation Learning and Deep Reinforcement Learning approaches for automated UAV/camera control, path planning and cinematography planning, in the general context of "flying & filming".
Ioannis Pitas, Ioannis Mademlis
ACM Multimedia1
2022 Hyperspherical class prototypes for adversarial robustness
Vasileios Mygdalis, Ioannis Pitas
Pattern Recognit.2
2022 Rethinking Road Surface 3-D Reconstruction and Pothole Detection: From Perspective Transformation to Disparity Map Segmentation
abstract
Potholes are one of the most common forms of road damage, which can severely affect driving comfort, road safety, and vehicle condition. Pothole detection is typically performed by either structural engineers or certified inspectors. However, this task is not only hazardous for the personnel but also extremely time consuming. This article presents an efficient pothole detection algorithm based on road disparity map estimation and segmentation. We first incorporate the stereo rig roll angle into shifting distance calculation to generalize perspective transformation. The road disparities are then efficiently estimated using semiglobal matching. A disparity map transformation algorithm is then performed to better distinguish the damaged road areas. Subsequently, we utilize simple linear iterative clustering to group the transformed disparities into a collection of superpixels. The potholes are finally detected by finding the superpixels, whose intensities are lower than an adaptively determined threshold. The proposed algorithm is implemented on an NVIDIA RTX 2080 Ti GPU in CUDA. The experimental results demonstrate that our proposed road pothole detection algorithm achieves state-of-the-art accuracy and efficiency.
Rui Fan 0001, Umar Özgünalp, Yuan Wang 0015, Ming Liu 0001, Ioannis Pitas
IEEE Trans. Cybern.5
2021 Adversarial Optimization Scheme For Online Tracking Model Adaptation In Autonomous Systems
abstract
Online tracking model updating is typically addressed as a regression problem, involving the minimization of the dispersion between the obtained tracker model response maps in each consecutive frame and some target distribution (e.g., Gaussian), using a closed-form solution. Inspired by the recent applications of Generative Adversarial Networks (GANs), we propose to solve this problem with an adversarial optimization scheme, by employing a GeneratorDiscriminator network pair. That is, the role of the Generator is assigned to the tracking model so that it produces response maps belonging to some target distribution, while an additional discriminator network is trained to identify if the tracker response maps produced by the generator belong to this target distribution, or not. Therefore, the tracker model exploits the discriminator network as an additional information pool about the target distribution. It is shown that this simple addition improves tracking performance in standard benchmark datasets, without significantly hurting training complexity, thus rendering the proposed method suitable for embedded system application such as in autonomous cars and Unmanned Aerial Systems.
Iason Karakostas, Vasileios Mygdalis, Ioannis Pitas
ICIP3
2021 Adversarial Unsupervised Video Summarization Augmented With Dictionary Loss
abstract
Automated unsupervised video summarization by key-frame extraction consists in identifying representative video frames, best abridging a complete input sequence, and temporally ordering them to form a video summary, without relying on manually constructed ground-truth key-frame sets. State-of-the-art unsupervised deep neural approaches consider the desired summary to be a subset of the original sequence, composed of video frames that are sufficient to visually reconstruct the entire input. They typically employ a pre-trained CNN for extracting a vector representation per RGB video frame and a baseline LSTM adversarial learning framework for identifying key-frames. In this paper, to better guide the network towards properly selecting video frames that can faithfully reconstruct the original video, we augment the baseline framework with an additional LSTM autoencoder, which learns in parallel a fixed-length representation of the entire original input sequence. This is exploited during training, where a novel loss term inspired by dictionary learning is added to the network optimization objectives, further biasing key-frame selection towards video frames which are collectively able to recreate the original video. Empirical evaluation on two common public relevant datasets indicates highly favourable results.
Michail Kaseris, Ioannis Mademlis, Ioannis Pitas
ICIP3
2021 Autonomous UAV Safety by Visual Human Crowd Detection Using Multi-Task Deep Neural Networks
abstract
Camera-equipped UAVs, or drones, are increasingly employed in a wide range of applications. Thus, ensuring their safe flight in areas containing people is a top priority. In this paper, a deep neural network-based method is proposed for the task of visual human crowd detection from UAV footage, allowing a drone to rapidly extract semantic segmentation maps from captured video frames during flight. These maps can be exploited (e.g., by a path planner) to define no-fly zones over, or near human crowds and, hence, enhance UAV flight safety. To this end, a novel neural architecture for binary (crowd/non- crowd) semantic segmentation from single RGB images is proposed, based on Convolutional Neural Networks (CNNs). It consists of a semantic segmentation and an image-to-image translation (I2I) neural branch. The overall network is trained using a novel multi-task loss function that addresses both tasks by processing the output of the corresponding branch. During inference, information flows across branches through additional skip synapses to further assist the crowd detection task. In order to evaluate the proposed method, we introduce a real and a synthetic human crowd RGB image dataset. The proposed method outperforms previous aerial crowd detection methods by a large margin and without any post-processing. Moreover, it demonstrates increased generalization ability, while running at real-time and near-real-time speeds on a ground computer and on embedded AI hardware, respectively.
Christos Papaioannidis, Ioannis Mademlis, Ioannis Pitas
ICRA3
2021 Leader and breakaway detection in racing sports videos
abstract
This paper addresses the important problem of leader detection in racing sports videos (e.g., cycling, boating and car racing events), as his/her proper framing is a pivotal issue in racing sports cinematography, where the events have a linear spatial deployment. Over the last few years, as autonomous drone vision and cinematography emerged, new challenges appeared in drone vision. While, until recently, most computer vision methods typically addressed still camera AV footage, drone sports cinematography typically employs moving cameras. In this paper, we solve the problem of leader detection in a group of similarly moving targets in sports videos, e.g. the leader of a sports cyclist group and his/her breakaway during a cycling event. This is very useful in drone sports cinematography, as it is important that the drone camera automatically centers on such a leader. We demonstrate that the novel method described in this paper can effectively solve the problem of leader detection in sports videos.
Sotirios Papadopoulos, Charalampos Symeonidis, Ioannis Pitas
MMSP3
2021 Multiview vision-based human crowd localization for UAV fleet flight safety
Efstratios Kakaletsis, Ioannis Mademlis, Nikos Nikolaidis 0001, Ioannis Pitas
Signal Process. Image Commun.4
2021 Occlusion detection and drift-avoidance framework for 2D visual object tracking
abstract
This paper presents a long-term 2D tracking framework for the coverage of live outdoor (e.g., sports) events that is suitable for embedded system application (e.g. Unmanned Aerial Vehicles). This application scenario requires 2D target (e.g., athlete, ball, bicycle, boat) tracking for visually assisting the UAV pilot (or cameraman) to maintain proper target framing, or even for actual 3D target following/localization when the drone flies autonomously. In these cases, it should be expected that the target to be tracked/followed, may disappear from the UAV camera field of view, due to fast 3D target motion, illumination changes, or due to visual target occlusions by obstacles, even if the actual UAV continues following it (either autonomously, by exploiting alternative target localization sensors, or by pilot maneuvering). Therefore, the 2D tracker should be able to recover from such situations. The proposed framework solves exactly this problem. Target occlusions are detected from the 2D tracker responses. Depending on the occlusion immensity, the proposed framework decides whether to not update the tracking model, or to employ target re-detection in a broader window. As a result, the proposed framework allows continued target tracking once the target re-appears in the video stream, without tracker re-initialization.
Iason Karakostas, Vasileios Mygdalis, Anastasios Tefas, Ioannis Pitas
Signal Process. Image Commun.4
2021 Graph Attention Layer Evolves Semantic Segmentation for Road Pothole Detection: A Benchmark and Algorithms
abstract
Existing road pothole detection approaches can be classified as computer vision-based or machine learning-based. The former approaches typically employ 2D image analysis/ understanding or 3D point cloud modeling and segmentation algorithms to detect (i.e., recognize and localize) road potholes from vision sensor data, e.g., RGB images and/or depth/disparity images. The latter approaches generally address road pothole detection using convolutional neural networks (CNNs) in an end-to-end manner. However, road potholes are not necessarily ubiquitous and it is challenging to prepare a large well-annotated dataset for CNN training. In this regard, while computer vision-based methods were the mainstream research trend in the past decade, machine learning-based methods were merely discussed. Recently, we published the first stereo vision-based road pothole detection dataset and a novel disparity transformation algorithm, whereby the damaged and undamaged road areas can be highly distinguished. However, there are no benchmarks currently available for state-of-the-art (SoTA) CNNs trained using either disparity images or transformed disparity images. Therefore, in this paper, we first discuss the SoTA CNNs designed for semantic segmentation and evaluate their performance for road pothole detection with extensive experiments. Additionally, inspired by graph neural network (GNN), we propose a novel CNN layer, referred to as graph attention layer (GAL), which can be easily deployed in any existing CNN to optimize image feature representations for semantic segmentation. Our experiments compare GAL-DeepLabv3+, our best-performing implementation, with nine SoTA CNNs on three modalities of training data: RGB images, disparity images, and transformed disparity images. The experimental results suggest that our proposed GAL-DeepLabv3+ achieves the best overall pothole detection accuracy on all training data modalities. The source code, dataset, and benchmark are publicly available at mias.group/GAL-Pothole-Detection.
Rui Fan 0001, Hengli Wang, Yuan Wang 0015, Ming Liu 0001, Ioannis Pitas
IEEE Trans. Image Process.5
2020 Shot type constraints in UAV cinematography for autonomous target tracking
Iason Karakostas, Ioannis Mademlis, Nikos Nikolaidis 0001, Ioannis Pitas
Inf. Sci.4
2020 Dense convolutional feature histograms for robust visual object tracking
Paraskevi Nousi, Anastasios Tefas, Ioannis Pitas
Image Vis. Comput.3
2020 K-Anonymity inspired adversarial attack and multiple one-class classification defense
Vasileios Mygdalis, Anastasios Tefas, Ioannis Pitas
Neural Networks3
2020 Deep autoencoders for attribute preserving face de-identification
Paraskevi Nousi, Sotirios Papadopoulos, Anastasios Tefas, Ioannis Pitas
Signal Process. Image Commun.4
2020 Re-identification framework for long term visual object tracking based on object detection and classification
Paraskevi Nousi, Danai Triantafyllidou, Anastasios Tefas, Ioannis Pitas
Signal Process. Image Commun.4
2020 3D Object Pose Estimation Using Multi-Objective Quaternion Learning
abstract
In this paper, a framework is proposed for object recognition and pose estimation from color images using convolutional neural networks (CNNs). 3D object pose estimation along with object recognition has numerous applications, such as robot positioning versus a target object and robotic object grasping. Previous methods addressing this problem relied on both color and depth (RGB-D) images to learn low-dimensional viewpoint descriptors for object pose retrieval. In the proposed method, a novel quaternion-based multi-objective loss function is used, which combines manifold learning and regression to learn 3D pose descriptors and direct 3D object pose estimation, using only color (RGB) images. The 3D object pose can then be obtained either by using the learned descriptors in the nearest neighbor (NN) search or by direct neural network regression. An extensive experimental evaluation has proven that such descriptors provide greater pose estimation accuracy than the state-of-the-art methods. In addition, the learned 3D pose descriptors are almost object-independent and, thus, generalizable to unseen objects. Finally, when the object identity is not of interest, the 3D object pose can be regressed directly from the network, by overriding the NN search, thus significantly reducing the object pose inference time.
Christos Papaioannidis, Ioannis Pitas
IEEE Trans. Circuits Syst. Video Technol.2
2020 Pothole Detection Based on Disparity Transformation and Road Surface Modeling
abstract
Pothole detection is one of the most important tasks for road maintenance. Computer vision approaches are generally based on either 2D road image analysis or 3D road surface modeling. However, these two categories are always used independently. Furthermore, the pothole detection accuracy is still far from satisfactory. Therefore, in this paper, we present a robust pothole detection algorithm that is both accurate and computationally efficient. A dense disparity map is first transformed to better distinguish between damaged and undamaged road areas. To achieve greater disparity transformation efficiency, golden section search and dynamic programming are utilized to estimate the transformation parameters. Otsu's thresholding method is then used to extract potential undamaged road areas from the transformed disparity map. The disparities in the extracted areas are modeled by a quadratic surface using least squares fitting. To improve disparity map modeling robustness, the surface normal is also integrated into the surface modeling process. Furthermore, random sample consensus is utilized to reduce the effects caused by outliers. By comparing the difference between the actual and modeled disparity maps, the potholes can be detected accurately. Finally, the point clouds of the detected potholes are extracted from the reconstructed 3D road surface. The experimental results show that the successful detection accuracy of the proposed system is around 98.7% and the overall pixel-level accuracy is approximately 99.6%.
Rui Fan 0001, Umar Özgünalp, Brett Hosking, Ming Liu 0001, Ioannis Pitas
IEEE Trans. Image Process.5
2020 Corrections to "Pothole Detection Based on Disparity Transformation and Road Surface Modeling"
abstract
Unfortunately, we made two minor mistakes in the above paper. First of all, the first graph on row (c) inFig. 11was same as the third graph on row (c) inFig. 11. Secondly, “precision” and “recall” inTable IIIneed to be switched. The correct figure and table have no influence on the discussion and conclusions in the above paper, and they are given here.
Rui Fan 0001, Umar Özgünalp, Brett Hosking, Ming Liu 0001, Ioannis Pitas
IEEE Trans. Image Process.5
2020 Domain-Translated 3D Object Pose Estimation
abstract
Synthetic 3D object models have been proven crucial in object pose estimation, as they are utilized to generate a huge number of accurately annotated data. The object pose estimation problem is usually solved for images originating from the real data domain by employing synthetic images for training data enrichment, without fully exploiting the fact that synthetic and real images may have different data distributions. In this work, we argue that 3D object pose estimation problem is easier to solve for images originating from the synthetic domain, rather than the real data domain. To this end, we propose a 3D object pose estimation framework consisting of a two-step process, where a novel pose-oriented image-to-image translation step is first employed to translate noisy real images to clean synthetic ones and then, a 3D object pose estimation method is applied on the translated synthetic images to finally predict the 3D object poses. A novel pose-oriented objective function is employed for training the image-to-image translation network, which enforces that pose-related object image characteristics are preserved in the translated images. As a result, the pose estimation network does not require real data for training purposes. Experimental evaluation has shown that the proposed framework greatly improves the 3D object pose estimation performance, when compared to state-of-the-art methods.
Christos Papaioannidis, Vasileios Mygdalis, Ioannis Pitas
IEEE Trans. Image Process.3
2019 Shot Type Feasibility in Autonomous UAV Cinematography
abstract
Aerial cinematography relying on camera-equipped umanned aerial vehicles (UAVs), or drones, has revolutionized media production during the past years. Autonomous UAV function-alities are already being employed to a degree, in a manner structured mainly around visual target tracking. From a cinematographic point of view, the desired shot type (i.e., Close-Up, Long Shot, etc.) is the most important factor affecting the artistic result. Achieving a specific shot type depends on the target-to-camera distance and the camera focal length. However, the interaction between UAV/camera motion trajectory (e.g., Orbit, Chase, etc.) and the visual tracker requirements constrains the range of feasible shot types at each time instance. In this paper, which extends previous work, these constraints are explored for a number of standard UAV/camera motion types, UAV shot types are classified and rules regarding shot feasibility over time are analytically derived. The proposed rules are evaluated in a realistic UAV simulation environment and achieve high performance, indicating possible benefits from their integration into an intelligent shooting system.
Iason Karakostas, Ioannis Mademlis, Nikos Nikolaidis 0001, Ioannis Pitas
ICASSP4
2019 Deep Convolutional Feature Histograms for Visual Object Tracking
abstract
Visual Object Tracking remains an open and challenging task in the Computer Vision field, requiring tracking algorithms to achieve a feeble balance between precision and speed performance. In this work, inspired by the classic Mean Shift algorithm for object tracking using histograms, as well as the recent advances of deep Convolutional Neural Networks (CNNs), we propose a novel tracker that incorporates elements from both worlds. Our tracker uses a deep CNN as the feature extraction backbone, which is capable of extracting semantically meaningful features from the target and its background, as well as a fully learnable Bag-of-Features mechanism which extracts histograms from those features. The tracker operates in a fully-convolutional fashion, allowing for the direct and efficient evaluation of multiple possible target locations. Extensive experimental results demonstrate the efficiency and effectiveness of the proposed tracker, allowing it to run at high speeds even on systems with lower computational capacity.
Paraskevi Nousi, Anastasios Tefas, Ioannis Pitas
ICASSP3
2019 Adversarial Face De-Identification
abstract
Recently, much research has been done on how to secure personal data, notably facial images. Face de-identification is one example of privacy protection that protects person identity by fooling intelligent face recognition systems, while typically allowing face recognition by human observers. While many face de-identification methods exist, the generated de-identified facial images do not resemble the original ones. This paper proposes the usage of adversarial examples for face de-identification that introduces minimal facial image distortion, while fooling automatic face recognition systems. Specifically, it introduces P-FGVM, a novel adversarial attack method, which operates on the image spatial domain and generates adversarial de-identified facial images that resemble the original ones. A comparison between P-FGVM and other adversarial attack methods shows that P-FGVM both protects privacy and preserves visual facial image quality more efficiently.
Efstathios Chatzikyriakidis, Christos Papaioannidis, Ioannis Pitas
ICIP3
2019 Joint Lightweight Object Tracking and Detection for Unmanned Vehicles
abstract
In this paper, we address the problem of lightweight and effective visual object tracking and we present a real-time tracking system suitable for integration in embedded autonomous platforms. We propose a novel tracking framework for classification-based re-detection and tracking, with learnable management of tracking and detection results. The proposed framework includes a novel, very efficient object reidentification method, which filters the detection candidates and systematically corrects the tracking results. In our experiments, we demonstrate the effectiveness of the proposed system by comparing its performance against several other state-of-the art trackers and report the results on the UAV123 and UAV20L datasets. The results indicate that the proposed method is significantly more robust and accurate against recent state-of-the-art trackers, surpassing problems caused by real-world scenarios, while maintaining fast tracking speeds, making it suitable for use in real-time vision applications for autonomous robots, such as Unmanned Aerial Vehicles (UAVs).
Paraskevi Nousi, Danai Triantafyllidou, Anastasios Tefas, Ioannis Pitas
ICIP4
2019 Computational UAV Cinematography for Intelligent Shooting Based on Semantic Visual Analysis
abstract
Audiovisual coverage of sports events using Unmanned Aerial Vehicles (UAVs) is becoming increasingly popular. Intelligent audiovisual (A/V) shooting tools, accurately identifying the 2D region of cinematographic attention (RoCA) depicting rapidly moving target ensembles and automatically controlling the UAVs/cameras through visual content analysis, are thus needed. A novel algorithmic pipeline is proposed, implementing computational UAV cinematography for assisting sports coverage, based on semantic, human-centered visual analysis. Athlete and ball detection / tracking results as well as their spatial distribution on the image plane are the semantic features extracted from UAV video feed and exploited for RoCA extraction, based solely on present and past target detections. A PID controller visually controlling a real or virtual camera to track the RoCA and produce aesthetically pleasing shots, without exploiting 3D location-related information, is employed. The proposed method is evaluated on actual UAV footage from soccer matches and promising results are obtained.
Fotini Patrona, Ioannis Mademlis, Anastasios Tefas, Ioannis Pitas
ICIP4
2019 Semantic Map Annotation Through UAV Video Analysis Using Deep Learning Models in ROS
Efstratios Kakaletsis, Maria Tzelepi, Pantelis I. Kaplanoglou, Charalampos Symeonidis, Nikos Nikolaidis 0001, Anastasios Tefas, Ioannis Pitas
MMM (2)7
2019 Greedy Salient Dictionary Learning for Activity Video Summarization
Ioannis Mademlis, Anastasios Tefas, Ioannis Pitas
MMM (1)3
2019 Exploiting multiplex data relationships in Support Vector Machines
Vasileios Mygdalis, Anastasios Tefas, Ioannis Pitas
Pattern Recognit.3
2019 Neurons With Paraboloid Decision Boundaries for Improved Neural Network Classification Performance
abstract
In mathematical terms, an artificial neuron computes the inner product of a d-dimensional input vector x with its weight vector w, compares it with a bias value w0and fires based on the result of this comparison. Therefore, its decision boundary is given by the equation wTx + w0= 0. In this paper, we propose replacing the linear hyperplane decision boundary of a neuron with a curved, paraboloid decision boundary. Thus, the decision boundary of the proposed paraboloid neuron is given by the equation (hTx + h0)2- ||x - p||22= 0, where h and h0denote the parameters of the directrix and p denotes the coordinates of the focus. Such paraboloid neural networks are proven to have superior recognition accuracy in a number of applications.
Nikolaos Tsapanos, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas
IEEE Trans. Neural Networks Learn. Syst.4
2018 Regularized Svd-Based Video Frame Saliency for Unsupervised Activity Video Summarization
abstract
Storage, browsing and analysis of human activity videos can be significantly facilitated by automated video summarization. Unsupervised key-frame extraction remains the most widely applicable technique for summarizing activity videos. However, their specific properties make the problem difficult to solve. Typical relevant algorithms fall under the video frame clustering or the dictionary-of-representatives families, with salient dictionary learning having been recently proposed. Under this formulation, the video frames selected as key-frames are the ones which simultaneously best reconstruct the entire video and are salient compared to the rest. This paper improves upon such a method by replacing the video frame saliency estimation term with one based on Regularized SVD-based Low Rank Approximation, taking advantage of the well-established correlation between midrange matrix singular values and salient regions. Extensive empirical evaluation showcases the high performance of both the salient dictionary learning framework and the specific proposed method.
Ioannis Mademlis, Anastasios Tefas, Ioannis Pitas
ICASSP3
2018 Label Propagation on Facial Images Using Similarity and Dissimilarity Labelling Constraints
abstract
In this paper, a novel multimedia data (specifically facial images) label propagation method is presented that is based on the inclusion of labelling constraints in the objective function of the MLPP-CLP state of the art algorithm. The proposed method can incorporate pairwise facial image similarity and dissimilarity constraints into the objective function of the aforementioned method. Experiments which have been conducted on facial image labelling in three stereoscopic movies, confirm the increased labelling accuracy of the proposed method.
Efstratios Kakaletsis, Olga Zoidi, Ioannis Tsingalis, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas
ICIP6
2018 UAV Cinematography Constraints Imposed by Visual Target Tracking
abstract
Camera-equipped drones have recently revolutionized aerial cinematography, allowing easy acquisition of impressive footage. Although they are currently manually operated, autonomous functionalities based on machine learning and computer vision are becoming popular. However, the emerging area of autonomous UAV filming has to face several challenges, especially when visually tracking fast and unpredictably moving targets. In the latter case, an important issue is how to determine the shot types that are achievable without risking failure of the 2D visual tracker. This paper studies the constraints imposed to cinematography decision-making during autonomous UAV shooting. It focuses on formalizing and geometrically modelling common target-following UAV motion types, in order to analytically determine the maximum permissible camera focal length (therefore, the range of feasible shot types) for avoiding visual target tracking failure.
Iason Karakostas, Ioannis Mademlis, Nikos Nikolaidis 0001, Ioannis Pitas
ICIP4
2018 Learning Multi-Graph Regularization for SVM Classification
abstract
A classification method that emphasizes on learning the hyperplane that separates the training data with the maximum margin in a regularized space, is presented. In the proposed method, this regularized space is derived by exploiting multiple graph structures, in the SVM optimization process. Each of the employed graph structure carries some information concerning a geometric or semantic property about the training data, e.g., local neighborhood area and global geometric data relationships. The proposed method introduces information from each graph type to the standard SVM objective, as a projection of the SVM hyperplane to such a direction, where a specific property of the training data is highlighted. We show that each data property can be encoded in a regularized kernel matrix. Finally, response in the optimal classification space can be obtained by exploiting a weighted combination of multiple regularized kernel matrices. Experimental results in face recognition and object classification denote the effectiveness of the proposed method.
Vasileios Mygdalis, Anastasios Tefas, Ioannis Pitas
ICIP3
2018 Convolutional Neural Networks for Visual Information Analysis with Limited Computing Resources
abstract
Over the past decade, Deep Convolutional Neural Networks with heavy architectures and large numbers of parameters have achieved state-of-the-art results and eclipsed other methods in multiple visual analysis tasks, including object detection. However, the real-time requirements of such tasks directly conflict with the restricted computational capabilities of embedded systems, prohibiting the immediate deployment of bulky models, and necessitating their optimization for inference. Parameter pruning techniques reduce the number of parameters while reducing the input size leads to smaller internal representations, leading by extension to fewer computational operations. Furthermore, inference optimization schemes provided by Deep Learning frameworks can yield significant speed ups, for example by allowing half-precision floating point operations. We investigate the behavior of various model configurations in object detection tasks and perform a comparative study on inference optimization methods which aim to reduce the computational cost of Convolutional Neural Networks, while examining the effect of such methods on their performance, and propose architecture modifications for this purpose.
Paraskevi Nousi, Emmanouil Patsiouras, Anastasios Tefas, Ioannis Pitas
ICIP4
2018 Challenges in Autonomous UAV Cinematography: An Overview
abstract
Autonomous UAV cinematography is an active research field with exciting potential for the media industry. It bears the promise of greatly facilitating UAV shooting for various applications, while significantly reducing the costs compared to manual shooting. However, the general problem has not been clearly defined and the challenges arising from current legislation and technology restrictions have not been fully charted. A complete overview of issues related to autonomous UAV cinematography is needed, pertaining to the current situation in the field, so as to guide immediate-future research. The purpose of this paper is to lay exactly this groundwork, with the expectation of providing a global perspective to multiple domain-specific research communities. The outlined issues are partitioned into challenges deriving from ethical/legal/safety considerations and from operational/production requirements. A brief survey of current technological solutions, including their limitations, is also provided for each issue.
Ioannis Mademlis, Vasileios Mygdalis, Nikos Nikolaidis 0001, Ioannis Pitas
ICME4
2018 Efficient Camera Control using 2D Visual Information for Unmanned Aerial Vehicle-based Cinematography
abstract
Using Unmanned Aerial Vehicles (UAVs), also known as drones, for covering public sport events, such as bicycle races, is becoming increasingly popular. Even though the problem of controlling the flight path of a drone is well studied in the literature, little work has been done on controlling the shooting camera for producing professional grade video footage. In this work we propose a fast and efficient proportional-integral-derivative (PID) based control algorithm that rely solely on 2D visual information and we demonstrate that it is possible to accurately control the camera without inferring the 3D position of the target. To ensure that the proposed method will not exhibit undesired behavior, a genetic algorithm is used to tune its parameters using a properly defined fitness function. The proposed method is evaluated using two datasets that contain actual drone footage: a dataset that contains videos of a single cyclist, and a dataset that contains actually footage from a bicycle race event, the Giro D'Italia bicycle race.
Nikolaos Passalis, Anastasios Tefas, Ioannis Pitas
ISCAS3
2018 Semi-supervised subclass support vector data description for image and video classification
Vasileios Mygdalis, Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
Neurocomputing4
2018 A salient dictionary learning framework for activity video summarization via key-frame extraction
Ioannis Mademlis, Anastasios Tefas, Ioannis Pitas
Inf. Sci.3
2018 Fast constrained person identity label propagation in stereo videos using a pruned similarity matrix
Efstratios Kakaletsis, Olga Zoidi, Ioannis Tsingalis, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas
Signal Process. Image Commun.6
2018 Positive and Negative Label Propagations
abstract
This paper extends the state-of-the-art label propagation (LP) framework in the propagation of negative labels. More specifically, the state-of-the-art LP methods propagate information of the form “the sample i should be assigned the label k.” The proposed method extends the state-of-theart framework by considering additional information of the form “the sample i should not be assigned the label k.” A theoretical analysis is presented in order to include negative LP in the problem formulation. Moreover, a method for selecting the negative labels in cases when they are not inherent from the data structure is presented. Furthermore, the incorporation of negative label information in two multigraph LP methods is presented. Finally, a discussion on the proposed algorithm extension to out of sample data, as well as scalability issues, is presented. Experimental results in various scenarios showed that the incorporation of negative label information increases, in all cases, the classification accuracy of the state of the art.
Olga Zoidi, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas
IEEE Trans. Circuits Syst. Video Technol.4
2017 Summarization of human activity videos via low-rank approximation
abstract
Summarization of videos depicting human activities is a timely problem with important applications, e.g., in the domains of surveillance or film/TV production, that steadily becomes more relevant. Research on video summarization has mainly relied on global clustering or local (frame-by-frame) saliency methods to provide automated algorithmic solutions for key-frame extraction. This work presents a method based on selecting as key-frames video frames able to optimally reconstruct the entire video. The novelty lies in modelling the reconstruction algebraically as a Column Subset Selection Problem (CSSP), resulting in extracting key-frames that correspond to elementary visual building blocks. The problem is formulated under an optimization framework and approximately solved via a genetic algorithm. The proposed video summarization method is being evaluated using a publicly available annotated dataset and an objective evaluation metric. According to the quantitative results, it clearly outperforms the typical clustering approach.
Ioannis Mademlis, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas
ICASSP4
2017 Summarization of human activity videos using a salient dictionary
abstract
Video summarization has become more prominent during the last decade, due to the massive amount of available digital video content. A video summarization algorithm is typically fed an input video and expected to extract a set of important key-frames which represent the entire content, convey semantic meaning and are significantly more concise than the original input. The most wide-spread approach relies on video frame clustering and extraction of the frames closest to the cluster centroids as key-frames. Such a process, although efficient, offloads the burden of semantic scene content modelling exclusively to the employed video frame description/representation scheme, while summarization itself is approached simply as a distance-based data partitioning problem. This work focuses on videos depicting human activities (e.g., from surveillance feeds) which display an attractive property, i.e., each video frame can be seen as a linear combination of elementary visual words (i.e., basic activity components). This is exploited so as to identify the video frames containing only the elementary visual building blocks, which ideally form a set of independent basis vectors that can linearly reconstruct the entire video. In this manner, the semantic content of the scene is considered by the video summarization process itself. The above process is modulated by a traditional distance-based video frame saliency estimation, biasing towards more spread content coverage and outlier inclusion, under a joint optimization framework derived from the Column Subset Selection Problem (CSSP). The proposed algorithm results in a final key-frame set which acts as as salient dictionary for the input video. Empirical evaluation conducted on a publicly available dataset suggest that the presented method outperforms both a baseline clustering-based approach and a state-of-the-art sparse dictionary learning-based algorithm.
Ioannis Mademlis, Anastasios Tefas, Ioannis Pitas
ICIP3
2017 Approximate kernel extreme learning machine for large scale data classification
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
Neurocomputing3
2017 De-identifying facial images using singular value decomposition and projections
Panteleimon Chriskos, Olga Zoidi, Anastasios Tefas, Ioannis Pitas
Multim. Tools Appl.4
2017 Multimodal speaker clustering in full length movies
Ioannis Kapsouras, Anastasios Tefas, Nikos Nikolaidis 0001, Geoffroy Peeters, Elie-Laurent Benaroya, Ioannis Pitas
Multim. Tools Appl.6
2017 One-Class Classification Based on Extreme Learning and Geometric Class Information
Alexandros Iosifidis, Vasileios Mygdalis, Anastasios Tefas, Ioannis Pitas
Neural Process. Lett.4
2017 Big Media Data Analysis
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas, Moncef Gabbouj
Signal Process. Image Commun.3
2017 Automatic Detection of 3D Quality Defects in Stereoscopic Videos Using Binocular Disparity
abstract
The 3D video quality issues that may disturb the human visual system and negatively impact the 3D viewing experience are well known and become more relevant as the availability of 3D video content increases, primarily through 3D cinema, but also through 3D television. In this paper, we propose four algorithms that exploit available stereo disparity information, in order to detect disturbing stereoscopic effects, namely, stereoscopic window violations, bent window effects, uncomfortable fusion object objects, and depth jump cuts on stereo videos. After detecting such issues, the proposed algorithms characterize them, based on the stress they cause to the viewer's visual system. Qualitative representative examples, quantitative experimental results on a custom-made video data set, a parameter sensitivity study, and comments on the computational complexity of the algorithms are provided, in order to assess the accuracy and the performance of stereoscopic quality defect detection.
Sotirios Delis, Ioannis Mademlis, Nikos Nikolaidis 0001, Ioannis Pitas
IEEE Trans. Circuits Syst. Video Technol.4
2016 One class classification applied in facial image analysis
abstract
In this paper, we apply One-Class Classification methods in facial image analysis problems. We consider the cases where the available training data information originates from one class, or one of the available classes is of high importance. We propose a novel extension of the One-Class Extreme Learning Machines algorithm aiming at minimizing both the training error and the data dispersion and consider solutions that generate decision functions in the ELM space, as well as in ELM spaces of arbitrary dimensionality. We evaluate the performance in publicly available datasets. The proposed method compares favourably to other state-of-the-art choices.
Vasileios Mygdalis, Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
ICIP4
2016 Multi-view semantic temporal video segmentation
abstract
In this work, we propose a multi-view temporal video segmentation approach that employs a Gaussian scoring process for determining the best segmentation positions. By exploiting the semantic action information that the dense trajectories video description offers, this method can detect intra-shot actions as well, unlike shot boundary detection approaches. We compare the temporal segmentation results of the proposed method to both single-view and multi-view methods, and also compare the action recognition results obtained on ground truth video segments to the ones obtained on the proposed multi-view segments, on the IMPART multi-view action data set.
Thomas Theodoridis, Anastasios Tefas, Ioannis Pitas
ICIP3
2016 Exploiting local and global geometric data relationships in Support Vector Data Description
abstract
In this paper, we describe a one-class classification method based on Support Vector Data Description, which exploits multiple graph structures in its optimization process. We derive in a generic solution which can be employed for supervised one-class classification tasks. The devised method can produce linear or non-linear decision functions, depending on the adopted kernel function. In our experiments, we simultaneously adopted two graphs that describe local and global geometric training data relationships, respectively. We evaluated the proposed classifier in publicly available datasets, where its performance compared favorably against closely related methods.
Vasileios Mygdalis, Anastasios Tefas, Ioannis Pitas
ICPR3
2016 Robustness in blind camera identification
abstract
In this paper, we focus on studying the effects of various image operations on sensor fingerprint camera identification. It is known that artifacts in the image processing pipeline, such as pixel defects or unevenness of the responses in the CCD array as well black current noise leave telltale footprints. Nowadays, camera identification based on the analysis of these artifacts is a well established technology for linking an image to a specific camera. The sensor fingerprint is estimated from images taken from a device. A similarity measure is deployed in order to associate an image with the camera. However, when the images used in the sensor fingerprint estimation have been processed using e.g. gamma correction, contrast enhancement, histogram equalization or white balance, the properties of the detection statistic change, hence affecting fingerprint detection. In this paper we study this effect experimentally, towards quantifying the robustness of fingerprint detection in the presence of image processing operations.
Stamatis Samaras, Vasileios Mygdalis, Ioannis Pitas
ICPR3
2016 Movie shot selection preserving narrative properties
abstract
Automatic shot selection is an important aspect of movie summarization that is helpful both to producers and to audiences, e.g., for market promotion or browsing purposes. However, most of the related research has focused on shot selection based on low-level video content, which disregards semantic information, or on narrative properties extracted from text, which requires the movie script to be available. In this work, semantic shot selection based on the narrative prominence of movie characters in both the visual and the audio modalities is investigated, without the need for additional data such as a script. The output is a movie summary that only contains video frames from selected movie shots. Selection is controlled by a user-provided shot retention parameter, that removes key-frames/key-segments from the skim based on actor face appearances and speech instances. This novel process (Multimodal Shot Pruning, or MSP) is algebraically modelled as a multimodal matrix Column Subset Selection Problem, which is solved using an evolutionary computing approach.
Ioannis Mademlis, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas
MMSP4
2016 Laplacian one class extreme learning machines for human action recognition
abstract
A novel OCC method for human action recognition namely the Laplacian One Class Extreme Learning Machines is presented. The proposed method exploits local geometric data information within the OC-ELM optimization process. It is shown that emphasizing on preserving the local geometry of the data leads to a regularized solution, which models the target class more efficiently than the standard OC-ELM algorithm. The proposed method is extended to operate in feature spaces determined by the network hidden layer outputs, as well as in ELM spaces of arbitrary dimensions. Its superior performance against other OCC options is consistent among five publicly available human action recognition datasets.
Vasileios Mygdalis, Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
MMSP4
2016 Exploiting stereoscopic disparity for augmenting human activity recognition performance
Ioannis Mademlis, Alexandros Iosifidis, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas
Multim. Tools Appl.5
2016 Big Data Analysis for Media Production
abstract
A typical high-end film production generates several terabytes of data per day, either as footage from multiple cameras or as background information regarding the set (laser scans, spherical captures, etc). This paper presents solutions to improve the integration of the multiple data sources, and understand their quality and content, which are useful both to support creative decisions on-set (or near it) and enhance the postproduction process. The main cinema specific contributions, tested on a multisource production dataset made publicly available for research purposes, are the monitoring and quality assurance of multicamera set-ups, multisource registration and acceleration of 3-D reconstruction, anthropocentric visual analysis techniques for semantic content annotation, and integrated 2-D–3-D web visualization tools. We discuss as well improvements carried out in basic techniques for acceleration, clustering and visualization, which were necessary to deal with the very large multisource data, and can be applied to other big data problems in diverse application fields.
Josep Blat, Alun Evans, Hansung Kim 0001, Evren Imre, Lukás Polok, Viorela Ila, Nikos Nikolaidis 0001, Pavel Zemcík, Anastasios Tefas, Pavel Smrz, Adrian Hilton 0001, Ioannis Pitas
Proc. IEEE12
2016 Graph Embedded One-Class Classifiers for media data classification
Vasileios Mygdalis, Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
Pattern Recognit.4
2016 Graph Embedded Extreme Learning Machine
abstract
In this paper, we propose a novel extension of the extreme learning machine (ELM) algorithm for single-hidden layer feedforward neural network training that is able to incorporate subspace learning (SL) criteria on the optimization process followed for the calculation of the network's output weights. The proposed graph embedded ELM (GEELM) algorithm is able to naturally exploit both intrinsic and penalty SL criteria that have been (or will be) designed under the graph embedding framework. In addition, we extend the proposed GEELM algorithm in order to be able to exploit SL criteria in arbitrary (even infinite) dimensional ELM spaces. We evaluate the proposed approach on eight standard classification problems and nine publicly available datasets designed for three problems related to human behavior analysis, i.e., the recognition of human face, facial expression, and activity. Experimental results denote the effectiveness of the proposed approach, since it outperforms other ELM-based classification schemes in all the cases.
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
IEEE Trans. Cybern.3
2016 Multimodal Stereoscopic Movie Summarization Conforming to Narrative Characteristics
abstract
Video summarization is a timely and rapidly developing research field with broad commercial interest, due to the increasing availability of massive video data. Relevant algorithms face the challenge of needing to achieve a careful balance between summary compactness, enjoyability, and content coverage. The specific case of stereoscopic 3D theatrical films has become more important over the past years, but not received corresponding research attention. In this paper, a multi-stage, multimodal summarization process for such stereoscopic movies is proposed, that is able to extract a short, representative video skim conforming to narrative characteristics from a 3D film. At the initial stage, a novel, low-level video frame description method is introduced (frame moments descriptor) that compactly captures informative image statistics from luminance, color, optical flow, and stereoscopic disparity video data, both in a global and in a local scale. Thus, scene texture, illumination, motion, and geometry properties may succinctly be contained within a single frame feature descriptor, which can subsequently be employed as a building block in any key-frame extraction scheme, e.g., for intra-shot frame clustering. The computed key-frames are then used to construct a movie summary in the form of a video skim, which is post-processed in a manner that also considers the audio modality. The next stage of the proposed summarization pipeline essentially performs shot pruning, controlled by a user-provided shot retention parameter, that removes segments from the skim based on the narrative prominence of movie characters in both the visual and the audio modalities. This novel process (multimodal shot pruning) is algebraically modeled as a multimodal matrix column subset selection problem, which is solved using an evolutionary computing approach. Subsequently, disorienting editing effects induced by summarization are dealt with, through manipulation of the video skim. At the last step, the skim is suitably post-processed in order to reduce stereoscopic video defects that may cause visual fatigue.
Ioannis Mademlis, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas
IEEE Trans. Image Process.4
2016 Visual Voice Activity Detection in the Wild
abstract
The visual voice activity detection (V-VAD) problem in unconstrained environments is investigated in this paper. A novel method for V-VAD in the wild, exploiting local shape and motion information appearing at spatiotemporal locations of interest for facial video segment description and the bag of words model for facial video segment representation, is proposed. Facial video segment classification is subsequently performed using the state-of-the-art classification algorithms. Experimental results on one publicly available V-VAD dataset denote the effectiveness of the proposed method, since it achieves better generalization performance in unseen users, when compared to the recently proposed state-of-the-art methods. Additional results on a new unconstrained dataset provide evidence that the proposed method can be effective even in such cases in which any other existing method fails.
Fotini Patrona, Alexandros Iosifidis, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas
IEEE Trans. Multim.5
2015 Enhancing class discrimination in Kernel Discriminant Analysis
abstract
In this paper, we propose an optimization scheme aiming at optimal nonlinear data projection, in terms of Fisher ratio maximization. To this end, we formulate an iterative optimization scheme consisting of two processing steps: optimal data projection calculation and optimal class representation determination. Compared to the standard approach employing the class mean vectors for class representation, the proposed optimization scheme increases class discrimination in the reduced-dimensionality feature space. We evaluate the proposed method in standard classification problems, as well as on the classification of human actions and face, and show that it is able to achieve better generalization performance, when compared to the standard approach.
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
ICASSP3
2015 Exploiting subclass information in one-class support vector machine for video summarization
abstract
In this paper, we propose a method for video summarization based on human activity description. We formulate this problem as the one of automatic video segment selection based on a learning process that employs salient video segment paradigms. For this one-class classification problem, we introduce a novel variant of the One-Class Support Vector Machine (OC-SVM) classifier that exploits subclass information in the OC-SVM optimization problem, in order to jointly minimize the data dispersion within each subclass and determine the optimal decision function. We evaluate the proposed approach in three Hollywood movies, where the performance of the proposed SOC-SVM algorithm is compared with that of the OC-SVM. Experimental results denote that the proposed approach is able to outperform OC-SVM-based video segment selection.
Vasileios Mygdalis, Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
ICASSP4
2015 Multimodal Speaker Diarization Utilizing Face Clustering Information
Ioannis Kapsouras, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas
ICIG (2)4
2015 Merging linear discriminant analysis with Bag of Words model for human action recognition
abstract
In this paper we propose a novel method for human action recognition, that unifies discriminative Bag of Words (BoW)-based video representation and discriminant subspace learning. An iterative optimization scheme is proposed for sequential discriminant BoWs-based action representation and code-book adaptation based on action discrimination in a reduced dimensionality feature space where action classes are better discriminated. Experiments on four publicly available action recognition data sets demonstrate that the proposed unified approach increases the discriminative ability of the obtained video representation, providing enhanced action classification performance.
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
ICIP3
2015 Large-scale nonlinear facial image classification based on approximate kernel Extreme Learning Machine
abstract
In this paper, we propose a scheme that can be used in large-scale nonlinear facial image classification problems. An approximate solution of the kernel Extreme Learning Machine classifier is formulated and evaluated. Experiments on two publicly available facial image datasets using two popular facial image representations illustrate the effectiveness and efficiency of the proposed approach. The proposed Approximate Kernel Extreme Learning Machine classifier is able to scale well in both time and memory, while achieving good generalization performance. Specifically, it is shown that it outperforms the standard ELM approach for the same time and memory requirements. Compared to the original kernel ELM approach, it achieves similar (or better) performance, while scaling well in both time and memory with respect to the training set cardinality.
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
ICIP3
2015 Facial image analysis based on two-dimensional linear discriminant analysis exploiting symmetry
abstract
In this paper a novel subspace learning technique is introduced for facial image analysis. The proposed technique takes into account the symmetry nature of facial images. This information is exploited by properly incorporating a symmetry constraint into the objective function of the Two-Dimensional Linear Discriminant Analysis (2DLDA) to determine symmetric projection vectors. The performance of the proposed Symmetric Two-Dimensional Linear Discriminant Analysis was evaluated on real face recognition databases. Experimental results highlight the superiority of the proposed technique in comparison to standard approach.
Konstantinos Papachristou, Anastasios Tefas, Ioannis Pitas
ICIP3
2015 Visual voice activity detection based on spatiotemporal information and bag of words
abstract
A novel method for Visual Voice Activity Detection (V-VAD) that exploits local shape and motion information appearing at spatiotemporal locations of interest for facial region video description and the Bag of Words (BoW) model for facial region video representation is proposed in this paper. Facial region video classification is subsequently performed based on Single-hidden Layer Feedforward Neural (SLFN) network trained by applying the recently proposed kernel Extreme Learning Machine (kELM) algorithm on training facial videos depicting talking and non-talking persons. Experimental results on two publicly available V-VAD data sets, denote the effectiveness of the proposed method, since better generalization performance in unseen users is achieved, compared to recently proposed state-of-the-art methods.
Fotini Patrona, Alexandros Iosifidis, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas
ICIP5
2015 Kernel matrix trimming for improved Kernel K-means clustering
abstract
The Kernel k-Means algorithm for clustering extends the classic k-Means clustering algorithm. It uses the kernel trick to implicitly calculate distances on a higher dimensional space, thus overcoming the classic algorithm's inability to handle data that are not linearly separable. Given a set of n elements to cluster, the n × n kernel matrix is calculated, which contains the dot products in the higher dimensional space of every possible combination of two elements. This matrix is then referenced to calculate the distance between an element and a cluster center, as per classic k-Means. In this paper, we propose a novel algorithm for zeroing elements of the kernel matrix, thus trimming the matrix, which results in reduced memory complexity and improved clustering performance.
Nikolaos Tsapanos, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas
ICIP4
2015 Object motion analysis description in stereo video content
Theodoris Theodoridis, Konstantinos Papachristou, Nikos Nikolaidis 0001, Ioannis Pitas
Comput. Vis. Image Underst.4
2015 Distance-based human action recognition using optimized class representations
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
Neurocomputing3
2015 DropELM: Fast neural network regularization with Dropout and DropConnect
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
Neurocomputing3
2015 Subclass Graph Embedding and a Marginal Fisher Analysis paradigm
Anastasios Maronidis, Anastasios Tefas, Ioannis Pitas
Pattern Recognit.3
2015 A distributed framework for trimmed Kernel k-Means clustering
Nikolaos Tsapanos, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas
Pattern Recognit.4
2015 On the kernel Extreme Learning Machine classifier
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
Pattern Recognit. Lett.3
2015 Sparse extreme learning machine classifier exploiting intrinsic graphs
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
Pattern Recognit. Lett.3
2015 Facial image clustering in stereoscopic videos using double spectral analysis
Georgios Orfanidis, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas
Signal Process. Image Commun.4
2015 Class-Specific Reference Discriminant Analysis With Application in Human Behavior Analysis
abstract
In this paper, a novel nonlinear subspace learning technique for class-specific data representation is proposed. A novel data representation is obtained by applying nonlinear class-specific data projection to a discriminant feature space, where the data belonging to the class under consideration are enforced to be close to their class representation, while the data belonging to the remaining classes are enforced to be as far as possible from it. A class is represented by an optimized class vector, enhancing class discrimination in the resulting feature space. An iterative optimization scheme is proposed to this end, where both the optimal nonlinear data projection and the optimal class representation are determined in each optimization step. The proposed approach is tested on three problems relating to human behavior analysis: Face recognition, facial expression recognition, and human action recognition. Experimental results denote the effectiveness of the proposed approach, since the proposed class-specific reference discriminant analysis outperforms kernel discriminant analysis, kernel spectral regression, and class-specific kernel discriminant analysis, as well as support vector machine-based classification, in most cases.
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
IEEE Trans. Hum. Mach. Syst.3
2014 Facial image clustering in stereo videos using local binary patterns and double spectral analysis
abstract
In this work we propose the use of local binary patterns in combination with double spectral analysis for facial image clustering applied to 3D (stereoscopic) videos. Double spectral clustering involves the fusion of two well known algorithms: Normalized cuts and spectral clustering in order to improve the clustering performance. The use of local binary patterns upon selected fiducial points on the facial images proved to be a good choice for describing images. The framework is applied on 3D videos and makes use of the additional information deriving from the existence of two channels, left and right for further improving the clustering results.
Georgios Orfanidis, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas
CIDM4
2014 Stereoscopic video description for human action recognition
abstract
In this paper, a stereoscopic video description method is proposed that indirectly incorporates scene geometry information derived from stereo disparity, through the manipulation of video interest points. This approach is flexible and able to cooperate with any monocular low-level feature descriptor. The method is evaluated on the problem of recognizing complex human actions in natural settings, using a publicly available action recognition database of unconstrained stereoscopic 3D videos, coming from Hollywood movies. It is compared both against competing depth-aware approaches and a state-of-the-art monocular algorithm. Experimental results denote that the proposed approach outperforms them and achieves state-of-the-art performance.
Ioannis Mademlis, Alexandros Iosifidis, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas
CIMSIVP5
2014 Minimum Variance Extreme Learning Machine for human action recognition
abstract
In this paper we propose an algorithm for Single-hidden Layer Feedforward Neural networks training. Based on the observation that the learning process of such networks can be considered to be a non-linear mapping of the training data to a high-dimensional feature space, followed by a data projection process to a low-dimensional space where classification is performed by a linear classifier, we extend the Extreme Learning Machine (ELM) algorithm in order to exploit the training data dispersion in its optimization process. The proposed Minimum Variance Extreme Learning Machine classifier is evaluated in human action recognition, where we compare its performance with that of other ELM-based classifiers, as well as the kernel Support Vector Machine classifier.
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
ICASSP3
2014 Semi-supervised dimensionality reduction on data with multiple representations for label propagation on facial images
abstract
In this paper a novel method is introduced for semi-supervised dimensionality reduction on facial images extracted from stereo videos. It operates on image data with multiple representations and calculates a projection matrix that preserves locality information and a priori pairwise information, in the form of must-link and cannot-link constraints between the various data representations, as well as label information for a percentage of the data. The final data representation is a linear combination of the projections of all data representations. The performance of the proposed Semi-supervised Multiple Locality Preserving Projections method was evaluated in person identity label propagation on facial images extracted from stereo movies. Experimental results showed that the proposed method outperforms state of the art methods.
Olga Zoidi, Nikos Nikolaidis 0001, Ioannis Pitas
ICASSP3
2014 Human action recognition based on bag of features and multi-view neural networks
abstract
In this paper, we employ Single-hidden Layer Feedforward Neural networks in order to perform human action recognition based on multiple action representations. In order to determine both optimized network and action representation combination weights, we propose an optimization process that jointly minimizes the overall network training error and the within-class variance of the training data in the corresponding hidden layer spaces. The proposed approach has been evaluated by using the state-of-the-art Bag of Features-based action video representation on three publicly available action recognition databases, where it outperforms two commonly used video representation combination approaches, as well as the best single-descriptor classification outcome.
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
ICIP3
2014 Stereoscopic video shot clustering into semantic concepts based on visual and disparity information
abstract
In this paper, we propose a framework for clustering shots from stereoscopic videos into clusters that correspond to semantic concepts exploiting visual and disparity information. Various color, disparity and texture descriptors are applied to shot key frames for obtaining low-level representations. Self Organizing Maps are subsequently employed upon various combinations of these representations in order to determine a lattice of representative semantic concepts. Experimental results on performances and football stereoscopic videos show that the use of disparity information leads to better clustering compared to using visual information only.
Konstantinos Papachristou, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas
ICIP4
2014 Label propagation on data with multiple representations through multi-graph locality preserving projections
abstract
In this paper a novel method is introduced for propagating label information on data with multiple representations. The method performs dimensionality reduction of the data by calculating a projection matrix that preserves locality information and a priori pairwise information, in the form of must-link and cannot-link constraints between the various data representations. The final data representations are then fused, in order to perform label propagation. The performance of the proposed method was evaluated on facial images extracted from stereo movies and on the UCF11 action recognition database. Experimental results showed that the proposed method outperforms state of the art methods.
Olga Zoidi, Nikos Nikolaidis 0001, Ioannis Pitas
ICIP3
2014 Semi-supervised Classification of Human Actions Based on Neural Networks
abstract
In this paper, we propose a novel algorithm for Single-hidden Layer Feed forward Neural networks training which is able to exploit information coming from both labeled and unlabeled data for semi-supervised action classification. We extend the Extreme Learning Machine algorithm by incorporating appropriate regularization terms describing geometric properties and discrimination criteria of the training data representation in the ELM space to this end. The proposed algorithm is evaluated on human action recognition, where its performance is compared with that of other (semi-)supervised classification schemes. Experimental results on two publicly available action recognition databases denote its effectiveness.
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
ICPR3
2014 Efficient automatic detection of 3D video artifacts
abstract
This paper summarizes some common artifacts in stereo video content. These artifacts lead to poor even uncomfortable 3D viewing experience. Efficient approaches for detecting three typical artifacts, sharpness mismatch, synchronization mismatch and stereoscopic window violation, are presented in detail. Sharpness mismatch is estimated by measuring the width deviations of edge pairs in depth planes. Synchronization mismatch is detected based on the motion inconsistencies of feature points between the stereoscopic channels in a short time frame. Stereoscopic window violation is detected, using connected component analysis, when objects hit the vertical frame boundaries while being in front of the virtual screen. For experiments, test sequences were created in a professional studio environment and state-of-the-art metrics were used for evaluating the proposed approaches. The experimental results show that our algorithms have considerable robustness in detecting 3D defects.
Mohan Liu, Ioannis Mademlis, Patrick Ndjiki-Nya, Jean-Charles Le Quintrec, Nikos Nikolaidis 0001, Ioannis Pitas
MMSP6
2014 2D/3D AudioVisual content analysis & description
abstract
In this paper, we propose a way of using the Audio-Visual Description Profile (AVDP) of the MPEG-7 standard for 2D or stereo video and multichannel audio content description. Our aim is to provide means of using AVDP in such a way, that 3D video and audio content can be correctly and consistently described. Since AVDP semantics do not include ways for dealing with 3D audiovisual content, a new semantic framework within AVDP is proposed and examples of using AVDP to describe the results of analysis algorithms on stereo video and multichannel audio content are presented.
Ioannis Pitas, Konstantinos Papachristou, Nikos Nikolaidis 0001, Marco Liuni, Elie-Laurent Benaroya, Geoffroy Peeters, Axel Röbel, Antje Linnemann, Mohan Liu, Sebastian Gerke
MMSP1
2014 Shot type characterization in 2D and 3D video content
abstract
Due to the enormous increase of video and image content on the web in the last decades, automatic video annotation became a necessity. The successful annotation of video and image content facilitate a successful indexing and retrieval in search databases. In this work we study a variety of possible shot type characterizations that can be assigned in a single video frame or still image. Possible ways to propagate these characterizations to a video segment (or to an entire shot) are also discussed. A method for the detection of Over-the-Shoulder shots in 3D (stereo) video is also proposed.
Ioannis Tsingalis, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas
MMSP4
2014 Greek folk music denoising under a symmetric α-stable noise assumption
abstract
The noise in musical audio recordings is assumed to obey an α-stable distribution. A sparse linear regression framework with structured priors is elaborated. Markov Chain Monte Carlo is used to infer the clean music signal model and the α-stable noise distribution parameters. The musical audio recordings are processed both as a whole and in segments by using a sine-bell window for analysis and overlap-and-add reconstruction. Experiments on noisy Greek folk music excerpts demonstrate better denoising under the α-stable noise assumption than the Gaussian white noise one, and when processing is performed in segments rather than in full recordings.
Nikoletta Bassiou, Constantine Kotropoulos, Ioannis Pitas
QSHINE3
2014 Regularized extreme learning machine for multi-view semi-supervised action recognition
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
Neurocomputing3
2014 Kernel Reference Discriminant Analysis
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
Pattern Recognit. Lett.3
2014 Discriminant Bag of Words based representation for human action recognition
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
Pattern Recognit. Lett.3
2014 Stereo object tracking with fusion of texture, color and disparity information
Olga Zoidi, Nikos Nikolaidis 0001, Anastasios Tefas, Ioannis Pitas
Signal Process. Image Commun.4
2014 Projected Gradients for Subclass Discriminant Nonnegative Subspace Learning
abstract
Current discriminant nonnegative matrix factorization (NMF) methods either do not guarantee convergence to a stationary limit point or assume a compact data distribution inside classes, thus ignoring intra class variance in extracting discriminant data samples representations. To address both limitations, we regard that data inside each class has a multimodal distribution, forming various subclasses and perform optimization using a projected gradients framework to ensure limit point stationarity. The proposed method combines appropriate clustering-based discriminant criteria in the NMF decomposition cost function, in order to find discriminant projections that enhance class separability in the reduced dimensional projection space, thus improving classification performance. The developed algorithms have been applied to facial expression, face and object recognition, and experimental results verified that they successfully identified discriminant parts, thus enhancing recognition performance.
Symeon Nikitidis, Anastasios Tefas, Ioannis Pitas
IEEE Trans. Cybern.3
2014 Maximum Margin Projection Subspace Learning for Visual Data Analysis
abstract
Visual pattern recognition from images often involves dimensionality reduction as a key step to discover a lower dimensional image data representation and obtain a more manageable problem. Contrary to what is commonly practiced today in various recognition applications where dimensionality reduction and classification are independently treated, we propose a novel dimensionality reduction method appropriately combined with a classification algorithm. The proposed method called maximum margin projection pursuit, aims to identify a low dimensional projection subspace, where samples form classes that are better discriminated, i.e., are separated with maximum margin. The proposed method is an iterative alternate optimization algorithm that computes the maximum margin projections exploiting the separating hyperplanes obtained from training a support vector machine classifier in the identified low dimensional space. Experimental results on both artificial data, as well as, on popular databases for facial expression, face and object recognition verified the superiority of the proposed method against various state-of-the-art dimensionality reduction algorithms.
Symeon Nikitidis, Anastasios Tefas, Ioannis Pitas
IEEE Trans. Image Process.3
2014 Symmetric Subspace Learning for Image Analysis
abstract
Subspace learning (SL) is one of the most useful tools for image analysis and recognition. A large number of such techniques have been proposed utilizing a priori knowledge about the data. In this paper, new subspace learning techniques are presented that use symmetry constraints in their objective functions. The rational behind this idea is to exploit the a priori knowledge that geometrical symmetry appears in several types of data, such as images, objects, faces, and so on. Experiments on artificial, facial expression recognition, face recognition, and object categorization databases highlight the superiority and the robustness of the proposed techniques, in comparison with standard SL techniques.
Konstantinos Papachristou, Anastasios Tefas, Ioannis Pitas
IEEE Trans. Image Process.3
2014 Person Identity Label Propagation in Stereo Videos
abstract
In this paper a novel method is introduced for propagating person identity labels on facial images extracted from stereo videos. It operates on image data with multiple representations and calculates a projection matrix that preserves locality information and a priori pairwise information, in the form of must-link and cannot-link constraints between the various data representations. The final data representation is a linear combination of the projections of all data representations. Moreover, the proposed method takes into account information obtained through data clustering. This information is exploited during the data propagation step in two ways: to regulate the similarity strength between the projected data and to indicate which samples should be selected for label propagation initialization. The performance of the proposed Multiple Locality Preserving Projections with Cluster-based Label Propagation (MLPP-CLP) method was evaluated on facial images extracted from stereo movies. Experimental results showed that the proposed method outperforms state of the art methods.
Olga Zoidi, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas
IEEE Trans. Multim.4
2013 Neural Networks for Digital Media Analysis and Description
Anastasios Tefas, Alexandros Iosifidis, Ioannis Pitas
EANN (1)3
2013 Variational Bayesian inference for stereo object tracking
abstract
In this paper, we deal with object tracking in stereo video sequences. We introduce a Bayesian framework for utilizing the results of any conventional single channel object tracker, in order to accomplish the refinement of the tracking accuracy in the left/right video channel. In this Bayesian framework, a variational Bayesian algorithm is employed to this end, where a priori information about the object displacement (movement) over time is incorporated by means of a prior distribution. This a priori information is obtained in a pre-processing step, in which the object displacement over time is estimated. Experiments demonstrate the efficiency of the proposed post-processing methodology in terms of tracking accuracy.
Giannis K. Chantas, Nikos Nikolaidis 0001, Ioannis Pitas
ICASSP3
2013 Appearance based object tracking in stereo sequences
abstract
A novel algorithm is proposed, that performs tracking of rigid objects in 3D videos, without knowledge of the camera calibration parameters, by exploiting only visual information obtained from the left and right video channels, namely luminance and disparity information. The proposed algorithm exploits noisy disparity maps that have been extracted by a real-time disparity estimation algorithm. The algorithm employs two appearance-based representation methods for describing the object texture. The first one combines luminance with disparity information and the second one employs Local Steering Kernel (LSK) descriptors.
Olga Zoidi, Nikos Nikolaidis 0001, Ioannis Pitas
ICASSP3
2013 Automatic 3D defects identification in stereoscopic videos
abstract
3DTV and 3D cinema have become quite popular during the last few years. It is now well understood that certain 3D video quality issues may have a negative effect in the 3D viewing experience. In this paper, we propose two novel algorithms that exploit available disparity information, in order to detect two disturbing stereoscopic issues, namely Stereoscopic Window Violations (SWV) and bent window effects. The algorithms' performance is tested on a number of examples. The proposed algorithms can be used for assessing the overall quality of stereoscopic video content or in order to enable fixing the detected issues in a post-production stage.
Sotirios Delis, Nikos Nikolaidis 0001, Ioannis Pitas
ICIP3
2013 Active classification for human action recognition
abstract
In this paper, we propose a novel classification method involving two processing steps. Given a test sample, the training data residing to its neighborhood are determined. Classification is performed by a Single-hidden Layer Feedforward Neural network exploiting labeling information of the training data appearing in the test sample neighborhood and using the rest training data as unlabeled. By following this approach, the proposed classification method focuses the classification problem on the training data that are more similar to the test sample under consideration and exploits information concerning to the training set structure. Compared to both static classification exploiting all the available training data and dynamic classification involving data selection for classification, the proposed active classification method provides enhanced classification performance in two publicly available action recognition databases.
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
ICIP3
2013 Exploiting the SVM constraints in NMF with application in eating and drinking activity recognition
abstract
A novel method is introduced for exploiting the support vector machine constraints in nonnegative matrix factorization. The notion of the proposed method is to find the projection matrix that projects the data to a low-dimensional space so that the data projections between the two classes are separated with maximum margin. Experiments were performed for the task of eating and drinking activity classification. Experimental results showed that the proposed method achieves better classification performance than the state of the art nonnegative matrix factorization and discriminant nonnegative matrix factorization followed by support vector machines classification.
Olga Zoidi, Anastasios Tefas, Ioannis Pitas
ICIP3
2013 Learning sparse representations for view-independent human action recognition based on fuzzy distances
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
Neurocomputing3
2013 Using robust dispersion estimation in support vector machines
Nicholas Vretos, Anastasios Tefas, Ioannis Pitas
Pattern Recognit.3
2013 Dynamic action recognition based on dynemes and Extreme Learning Machine
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
Pattern Recognit. Lett.3
2013 Multi-view action recognition based on action volumes, fuzzy distances and cluster discriminant analysis
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
Signal Process.3
2013 Minimum Class Variance Extreme Learning Machine for Human Action Recognition
abstract
In this paper, we propose a novel method aiming at view-independent human action recognition. Action description is based on local shape and motion information appearing at spatiotemporal locations of interest in a video. Action representation involves fuzzy vector quantization, while action classification is performed by a feedforward neural network. A novel classification algorithm, called minimum class variance extreme learning machine, is proposed in order to enhance the action classification performance. The proposed method can successfully operate in situations that may appear in real application scenarios, since it does not set any assumption concerning the visual scene background and the camera view angle. Experimental results on five publicly available databases, aiming at different application scenarios, denote the effectiveness of both the adopted action recognition approach and the proposed minimum class variance extreme learning machine algorithm.
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
IEEE Trans. Circuits Syst. Video Technol.3
2013 Visual Object Tracking Based on Local Steering Kernels and Color Histograms
abstract
In this paper, we propose a visual object tracking framework, which employs an appearance-based representation of the target object, based on local steering kernel descriptors and color histogram information. This framework takes as input the region of the target object in the previous video frame and a stored instance of the target object, and tries to localize the object in the current frame by finding the frame region that best resembles the input. As the object view changes over time, the object model is updated, hence incorporating these changes. Color histogram similarity between the detected object and the surrounding background is employed for background subtraction. Experiments are conducted to test the performance of the proposed framework under various conditions. The proposed tracking scheme is proven to be successful in tracking objects under scale and rotation variations and partial occlusion, as well as in tracking rather slowly deformable articulated objects.
Olga Zoidi, Anastasios Tefas, Ioannis Pitas
IEEE Trans. Circuits Syst. Video Technol.3
2013 Multidimensional Sequence Classification Based on Fuzzy Distances and Discriminant Analysis
abstract
In this paper, we present a novel method aiming at multidimensional sequence classification. We propose a novel sequence representation, based on its fuzzy distances from optimal representative signal instances, called statemes. We also propose a novel modified clustering discriminant analysis algorithm minimizing the adopted criterion with respect to both the data projection matrix and the class representation, leading to the optimal discriminant sequence class representation in a low-dimensional space, respectively. Based on this representation, simple classification algorithms, such as the nearest subclass centroid, provide high classification accuracy. A three step iterative optimization procedure for choosing statemes, optimal discriminant subspace and optimal sequence class representation in the final decision space is proposed. The classification procedure is fast and accurate. The proposed method has been tested on a wide variety of multidimensional sequence classification problems, including handwritten character recognition, time series classification and human activity recognition, providing very satisfactory classification results.
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
IEEE Trans. Knowl. Data Eng.3
2013 On the Optimal Class Representation in Linear Discriminant Analysis
abstract
Linear discriminant analysis (LDA) is a widely used technique for supervised feature extraction and dimensionality reduction. LDA determines an optimal discriminant space for linear data projection based on certain assumptions, e.g., on using normal distributions for each class and employing class representation by the mean class vectors. However, there might be other vectors that can represent each class, to increase class discrimination. In this brief, we propose an optimization scheme aiming at the optimal class representation, in terms of Fisher ratio maximization, for LDA-based data projection. Compared with the standard LDA approach, the proposed optimization scheme increases class discrimination in the reduced dimensionality space and achieves higher classification rates in publicly available data sets.
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
IEEE Trans. Neural Networks Learn. Syst.3
2013 Multiplicative Update Rules for Concurrent Nonnegative Matrix Factorization and Maximum Margin Classification
abstract
The state-of-the-art classification methods which employ nonnegative matrix factorization (NMF) employ two consecutive independent steps. The first one performs data transformation (dimensionality reduction) and the second one classifies the transformed data using classification methods, such as nearest neighbor/centroid or support vector machines (SVMs). In the following, we focus on using NMF factorization followed by SVM classification. Typically, the parameters of these two steps, e.g., the NMF bases/coefficients and the support vectors, are optimized independently, thus leading to suboptimal classification performance. In this paper, we merge these two steps into one by incorporating maximum margin classification constraints into the standard NMF optimization. The notion behind the proposed framework is to perform NMF, while ensuring that the margin between the projected data of the two classes is maximal. The concurrent NMF factorization and support vector optimization are performed through a set of multiplicative update rules. In the same context, the maximum margin classification constraints are imposed on the NMF problem with additional discriminant constraints and respective multiplicative update rules are extracted. The impact of the maximum margin classification constraints on the NMF factorization problem is addressed in Section VI. Experimental results in several databases indicate that the incorporation of the maximum margin classification constraints into the NMF and discriminant NMF objective functions improves the accuracy of the classification.
Olga Zoidi, Anastasios Tefas, Ioannis Pitas
IEEE Trans. Neural Networks Learn. Syst.3
2012 Eating and drinking activity recognition based on discriminant analysis of fuzzy distances and activity volumes
abstract
Eating and drinking activity recognition can be considered a solitary research field in activity recognition area. The development of an application capable to identify human eating and drinking activity can be really useful in a smart home environment targeting to extend independent living of older persons in the early stages of dementia. In this paper a novel method aiming at eating and drinking activity recognition is presented. Activities are considered as a sequence of human body poses forming 3D volumes, in which the third dimension refers to time. Fuzzy Vector Quantization is performed to associate the 3D volume representation of an activity video with 3D volume prototypes and Linear Discriminant Analysis is used to map activity representations in a low dimensional discriminant feature space. In this space a simple Nearest Centroid classification procedure leads to very satisfactory classification results.
Alexandros Iosifidis, Ermioni Marami, Anastasios Tefas, Ioannis Pitas
ICASSP4
2012 Visual object tracking based on the object's salient features with application in automatic nutrition assistance
abstract
A novel method for object tracking in videos which can find application in eating and drinking activity recognition is proposed. The query object is detected in the first video frame, extracting a new query image. The initial query image along with the obtained query image are then compared with patches within a determined search region around the position of the detected object in the previous frame. For each image, the local steering kernels are extracted and the similarity between a query image and the patches of the video frame is measured by calculating the cosine similarity. The proposed method finds application in eating and drinking activity recognition.
Olga Zoidi, Anastasios Tefas, Ioannis Pitas
ICASSP3
2012 Variational Bayesian inference for forward-backward visual tracking in stereo sequences
abstract
In this paper we propose a Bayesian framework for accurate object tracking in stereoscopic sequences. Object detection and forward tracking are first combined according to predefined rules to get a first set of tracked regions candidates. Backward tracking is then applied to provide another set of possible object localizations. Moreover, this strategy is applied herein in stereoscopic video. We introduce a Bayesian inference algorithm which is used to merge the information of both forward and backward tracking in order to refine the tracked region localization results. Experiments, performed on face tracking, show that the proposed method provides higher tracking accuracy than a forward tracker.
Giannis K. Chantas, Nikos Nikolaidis 0001, Ioannis Pitas
ICIP3
2012 Discriminant action representation for view-invariant person identification
abstract
In this paper we propose a novel person identification method exploiting human motion information. Persons are described by using their poses during action execution. Identification process involves Fuzzy Vector Quantization and Discriminant Learning. In the case of multiple cameras used in the identification phase, single-view identification results combination is achieved by employing a Bayesian combination strategy. The proposed identification approach does not set the assumptions of known action class and number of capturing cameras in the identification phase. Experimental results on two publicly available video databases denote the effectiveness of the proposed approach.
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
ICIP3
2012 Neural representation and learning for multi-view human action recognition
abstract
In this paper we propose a novel method aiming at view-independent multi-view action recognition. Instead of combining the information provided by all the cameras forming the camera setup, for action representation and classification, we perform single-view action representation and classification to all the available videos depicting the person under consideration independently. Action representation involves a self organizing neural network training followed by fuzzy vector quantization. Action classification is performed by a feedforward neural network which is trained for view-invariant action recognition. Multiple action classification results combination based on Bayesian learning, in the recognition phase, results to high action recognition accuracy. The performance of the proposed action recognition method is evaluated on two publicly available databases, aiming at different application scenarios.
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
IJCNN3
2012 Multi-view human movement recognition based on fuzzy distances and linear discriminant analysis
Alexandros Iosifidis, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas
Comput. Vis. Image Underst.4
2012 Multiplicative update rules for incremental training of multiclass support vector machines
Symeon Nikitidis, Nikos Nikolaidis 0001, Ioannis Pitas
Pattern Recognit.3
2012 Subclass discriminant Nonnegative Matrix Factorization for facial image analysis
Symeon Nikitidis, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas
Pattern Recognit.4
2012 Shape matching using a binary search tree structure of weak classifiers
Nikolaos Tsapanos, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas
Pattern Recognit.4
2012 Video fingerprinting using Latent Dirichlet Allocation and facial images
Nicholas Vretos, Nikos Nikolaidis 0001, Ioannis Pitas
Pattern Recognit.3
2012 Activity-Based Person Identification Using Fuzzy Representation and Discriminant Learning
abstract
In this paper, a novel view invariant person identification method based on human activity information is proposed. Unlike most methods proposed in the literature, in which “walk” (i.e., gait) is assumed to be the only activity exploited for person identification, we incorporate several activities in order to identify a person. A multicamera setup is used to capture the human body from different viewing angles. Fuzzy vector quantization and linear discriminant analysis are exploited in order to provide a discriminant activity representation. Person identification, activity recognition, and viewing angle specification results are obtained for all the available cameras independently. By properly combining these results, a view-invariant activity-independent person identification method is obtained. The proposed approach has been tested in challenging problem setups, simulating real application situations. Experimental results are very promising.
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
IEEE Trans. Inf. Forensics Secur.3
2012 View-Invariant Action Recognition Based on Artificial Neural Networks
abstract
In this paper, a novel view invariant action recognition method based on neural network representation and recognition is proposed. The novel representation of action videos is based on learning spatially related human body posture prototypes using self organizing maps. Fuzzy distances from human body posture prototypes are used to produce a time invariant action representation. Multilayer perceptrons are used for action classification. The algorithm is trained using data from a multi-camera setup. An arbitrary number of cameras can be used in order to recognize actions using a Bayesian framework. The proposed method can also be applied to videos depicting interactions between humans, without any modification. The use of information captured from different viewing angles leads to high classification performance. The proposed method is the first one that has been tested in challenging experimental setups, a fact that denotes its effectiveness to deal with most of the open issues in action recognition.
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
IEEE Trans. Neural Networks Learn. Syst.3
2011 Facial expression recognition using clustering discriminant Non-negative Matrix Factorization
abstract
Non-negative Matrix Factorization (NMF) is among the most popular subspace methods widely used in a variety of image processing problems. Recently, a discriminant NMF method that incorporates Linear Discriminant Analysis criteria and achieves an efficient decomposition of the provided data to its discriminant parts has been proposed. However, this approach poses several limitations since it assumes that the underline data distribution forms compact sets which is often unrealistic. To remedy this limitation we regard that data inside each class form various number of clusters and apply a Clustering based Discriminant Analysis. The proposed method combines appropriate discriminant constraints in the NMF decomposition cost function in order to address the problem of finding discriminant projections that enhance class separability in the reduced dimensional projection space. Experimental results performed on the Cohn-Kanade database verified the effectiveness of the proposed method in the facial expression recognition task.
Symeon Nikitidis, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas
ICIP4
2011 3D facial expression recognition using Zernike moments on depth images
abstract
In this paper we propose a new method for 3D facial expression recognition. We make use of the Zernike moments, which are calculated in the depth image of a 3D facial point cloud. Combining, the Zernike moments along with the 3D point clouds and the depth images, we succeed in tackling problems arising in facial expression recognition due to affine transformations of the data, such as translation, rotation and scaling which, in other approaches are considered very harmful in the overall accuracy of a facial expression recognition algorithm. Support vector machines are used in order to classify the previously extracted features. Results are drawn in two publicly available databases for 3D facial expression recognition.
Nicholas Vretos, Nikos Nikolaidis 0001, Ioannis Pitas
ICIP3
2011 A mutual information based face clustering algorithm for movie content analysis
Nicholas Vretos, Vassilios Solachidis, Ioannis Pitas
Image Vis. Comput.3
2011 Improving subspace learning for facial expression recognition using person dependent and geometrically enriched training sets
Anastasios Maronidis, Dimitris Bolis, Anastasios Tefas, Ioannis Pitas
Neural Networks4
2010 Improving the Robustness of Subspace Learning Techniques for Facial Expression Recognition
Dimitris Bolis, Anastasios Maronidis, Anastasios Tefas, Ioannis Pitas
ICANN (1)4
2010 Frontal View Recognition Using Spectral Clustering and Subspace Learning Methods
Anastasios Maronidis, Anastasios Tefas, Ioannis Pitas
ICANN (1)3
2010 Dynamic Shape Learning and Forgetting
Nikolaos Tsapanos, Anastasios Tefas, Ioannis Pitas
ICANN (3)3
2010 Multi-view object and human body part detection utilizing 3D scene information
abstract
The aim of this paper is to present a new method for multiview object or human body (or body part) detection. The basic idea consists of using a single view detector in every view of a scene captured by multiple cameras and then combining the results using the 3D information of the scene. The method can improve the results of the single view detector, while also localizing the object/human in the 3D space. This results in a robust way for rejecting the false detections, amending the missed detections and associating the results of the single view detector across views.
Georgios Sfiris, Nikos Nikolaidis 0001, Ioannis Pitas
ICIP3
2010 Video replica detection utilizing R-trees and frame-based voting
abstract
A novel color-based two-step, coarse-to-fine video replica detection system is proposed in this paper. The first step uses an R-tree in order to perform a coarse selection of the database (original) videos that potentially match the query video. A training procedure that utilizes attacked versions of the database videos and aims at achieving robustness to attacks is being used. A frame-based voting procedure is also involved. A refinement step that processes the set of videos returned by the first step in order to select the final matching video (if any) follows. The performance of the system has been evaluated on a database of short videos with good results.
Dimitrios Zotos, Nikos Nikolaidis 0001, Ioannis Pitas
ICME3
2010 Incremental Training of Multiclass Support Vector Machines
abstract
We present a new method for the incremental training of multiclass Support Vector Machines that provides computational efficiency for training problems in the case where the training data collection is sequentially enriched and dynamic adaptation of the classifier is required. An auxiliary function that incorporates some desired characteristics in order to provide an upper bound of the objective function which summarizes the multiclass classification task has been designed and the global minimizer for the enriched dataset is found using a warm start algorithm, since faster convergence is expected when starting from the previous global minimum. Experimental evidence on two data collections verified that our method is faster than retraining the classifier from scratch, while the achieved classification accuracy is maintained at the same level.
Symeon Nikitidis, Nikos Nikolaidis 0001, Ioannis Pitas
ICPR3
2010 Movement recognition exploiting multi-view information
abstract
In this paper a novel view-invariant movement recognition method is presented. A multi-camera setup is used to capture the movement from different observation angles. Identification of the position of each camera with respect to the subject's body is achieved by a procedure based on morphological operations and the proportions of the human body. Binary body masks from frames of all cameras, consistently arranged through the previous procedure, are concatenated to produce the so-called multi-view binary mask. These masks are rescaled and vectorized to create feature vectors in the input space. Fuzzy vector quantization is performed to associate input feature vectors with movement representations and linear discriminant analysis is used to map movements in a low dimensionality discriminant feature space. Experimental results show that the method can achieve very satisfactory recognition rates.
Alexandros Iosifidis, Nikos Nikolaidis 0001, Ioannis Pitas
MMSP3
2010 Online shape learning using binary search trees
Nikolaos Tsapanos, Anastasios Tefas, Ioannis Pitas
Image Vis. Comput.3
2010 Salient feature and reliable classifier selection for facial expression classification
Marios Kyperountas, Anastasios Tefas, Ioannis Pitas
Pattern Recognit.3
2010 Image replica detection system utilizing R-trees and linear discriminant analysis
Spiros Nikolopoulos, Stefanos Zafeiriou, Nikos Nikolaidis 0001, Ioannis Pitas
Pattern Recognit.4
2010 Automatic Color Based Reassembly of Fragmented Images and Paintings
abstract
The problem of reassembling image fragments arises in many scientific fields, such as forensics and archaeology. In the field of archaeology, the pictorial excavation findings are almost always in the form of painting fragments. The manual execution of this task is very difficult, as it requires great amount of time, skill and effort. Thus, the automation of such a work is very important and can lead to faster, more efficient, painting reassembly and to a significant reduction in the human effort involved. In this paper, an integrated method for automatic color based 2-D image fragment reassembly is presented. The proposed 2-D reassembly technique is divided into four steps. Initially, the image fragments which are probably spatially adjacent, are identified utilizing techniques employed in content based image retrieval systems. The second operation is to identify the matching contour segments for every retained couple of image fragments, via a dynamic programming technique. The next step is to identify the optimal transformation in order to align the matching contour segments. Many registration techniques have been evaluated to this end. Finally, the overall image is reassembled from its properly aligned fragments. This is achieved via a novel algorithm, which exploits the alignment angles found during the previous step. In each stage, the most robust algorithms having the best performance are investigated and their results are fed to the next step. We have experimented with the proposed method using digitally scanned images of actual torn pieces of paper image prints and we produced very satisfactory reassembly results.
Efthymia Tsamoura, Ioannis Pitas
IEEE Trans. Image Process.2
2009 Human identification from human movements
abstract
In this paper a multi-modal method for human identification that exploits the discrimination power of several movement types performed from the same human is proposed. Utilizing a fuzzy vector quantization (FVQ) and linear discriminant analysis (LDA) based algorithm, an unknown movement is first classified, and, then, the person performing the movement is recognized from a movement specific person classifier. In case that the unknown person performs more than one movements, a multi-modal algorithm combines the results of the individual classifiers to yield the final decision for the id of the unknown human. Using a publicly available database, we provide promising results regarding the discrimination power of the different movements for the human identification task, as well as we indicate that the combination of the individual classifiers may increase the robustness of the human recognition algorithm.
Nikolaos Gkalelis, Anastasios Tefas, Ioannis Pitas
ICIP3
2009 A model-based facial expression recognition algorithm using Principal Components Analysis
abstract
In this paper, we propose a new method for facial expression recognition. We utilize the Candide facial grid and apply principal components analysis (PCA) to find the two eigenvectors of the model vertices. These eigenvectors along with the barycenter of the vertices are used to define a new coordinate system where vertices are mapped. Support vector machines (SVMs) are then used for the facial expression classification task. The method is invariant to in-plane translation and rotation as well as scaling of the face and achieves very satisfactory results.
Nicholas Vretos, Nikos Nikolaidis 0001, Ioannis Pitas
ICIP3
2009 View indepedent human movement recognition from multi-view video exploiting a circular invariant posture representation
abstract
In this paper a novel method for view independent human movement representation and recognition, exploiting the rich information contained in multi-view videos, is proposed. The binary masks of a multi-view posture image are first vectorized, concatenated and the view correspondence problem between train and test samples is solved using the circular shift invariance property of the discrete Fourier transform (DFT) magnitudes. Then, using fuzzy vector quantization (FVQ) and linear discriminant analysis (LDA), different movements are represented and classified. This method allows view independent movement recognition, without the use of calibrated cameras, a-priori view correspondence information or 3D model reconstruction. A multi-view video database has been constructed for the assessment of the proposed algorithm. Evaluation of this algorithm on the new database, shows that it is particularly efficient and robust, and can achieve good recognition performance.
Nikolaos Gkalelis, Nikos Nikolaidis 0001, Ioannis Pitas
ICME3
2009 Frontal view recognition in multiview video sequences
abstract
In this paper, a novel method is proposed as a solution to the problem of frontal view recognition from multiview image sequences. Our aim is to correctly identify the view that corresponds to the camera placed in front of a person, or the camera whose view is closer to a frontal one. By doing so, frontal face images of the person can be acquired, in order to be used in face or facial expression recognition techniques that require frontal faces to achieve a satisfactory result. The proposed method firstly employs the Discriminant Non-Negative Matrix Factorization (DNMF) algorithm on the input images acquired from every camera. The output of the algorithm is then used as an input to a support vector machines (SVMs) system that classifies the head poses acquired from the cameras to two classes that correspond to the frontal or non frontal pose. Experiments conducted on the IDIAP database demonstrate that the proposed method achieves an accuracy of 98.6% in frontal view recognition.
Irene Kotsia, Nikos Nikolaidis 0001, Ioannis Pitas
ICME3
2009 A perceptual hashing algorithm using latent dirichlet allocation
abstract
This paper investigates the possibility of extracting latent aspects of a video, using visual information about humans (e.g. actors' faces), in order to develop a fingerprinting (replica detection) framework. We employ a generative probabilistic model, namely Latent Dirichlet Allocation (LDA), so as to capture latent aspects of a video, using facial semantic information derived from the video. We use the bag-of-words concept, (bag-of-faces in our case) in order to ensure exchangeability of the latent variables (e.g. topics). The video topics are modeled as a mixture of distributions of faces in each video. This generative probabilistic model has already been used in the case of text modeling with good results. Experimental results provide evidence that the proposed method performs very efficiently for video fingerprinting.
Nicholas Vretos, Nikos Nikolaidis 0001, Ioannis Pitas
ICME3
2009 Pairwise facial expression classification
abstract
This paper presents a novel facial expression recognition methodology. In order to classify the expression of a test face to one of seven pre-determined facial expression classes, multiple two-class classification tasks are carried out. For each such task, a unique set of features is identified that is enhanced, in terms of its ability to help produce a proper separation between the two specific classes. The selection of these sets of features is accomplished by making use of a class separability measure that is utilized in an iterative process. Fisher's linear discriminant is employed in order to produce the separation between each pair of classes and train each two-class classifier. In order to combine the classification results from all two-class classifiers, the `voting' classifier-decision fusion process is employed. The standard JAFFE database is utilized in order to evaluate the performance of this algorithm. Experimental results show that the proposed methodology provides a good solution to the facial expression recognition problem.
Marios Kyperountas, Anastasios Tefas, Ioannis Pitas
MMSP3
2009 3D head pose estimation in monocular video sequences by sequential camera self-calibration
abstract
This paper presents a novel approach for estimating 3D head pose in single-view video sequences acquired by an uncalibrated camera. Following the initialization by a face detector, a tracking technique localizes the faces in each frame in the video sequence. Head pose estimation is performed by using a structure from motion and self-calibration technique in a sequential way. The proposed method was applied to the IDIAP database that contains head pose ground truth data. The obtained results demonstrate that the method can estimate the head pose with satisfying accuracy.
Ioannis Marras, Nikos Nikolaidis 0001, Ioannis Pitas
MMSP3
2009 Facial feature detection using distance vector fields
Stylianos Asteriadis, Nikos Nikolaidis 0001, Ioannis Pitas
Pattern Recognit.3
2009 Semantic video fingerprinting and retrieval using face information
Costas I. Cotsaces, Nikos Nikolaidis 0001, Ioannis Pitas
Signal Process. Image Commun.3
2009 3-D Head Pose Estimation in Monocular Video Sequences Using Deformable Surfaces and Radial Basis Functions
abstract
This paper presents a novel approach for estimating 3-D head pose in single-view video sequences. Following initialization by a face detector, a tracking technique that utilizes a 3-D deformable surface model to approximate the facial image intensity is used to track the face in the video sequence. Head pose estimation is performed by using a feature vector which is a byproduct of the equations that govern the deformation of the surface model used in the tracking. The afore-mentioned vector is used as input in a radial basis function interpolation network in order to estimate the 3-D head pose. The proposed method was applied to IDIAP head pose estimation database. The obtained results show that the method can estimate the head direction vector with very good accuracy.
Michail Krinidis, Nikos Nikolaidis 0001, Ioannis Pitas
IEEE Trans. Circuits Syst. Video Technol.3
2009 Visual Lip Activity Detection and Speaker Detection Using Mouth Region Intensities
abstract
In this letter, we introduce a novel approach for lip activity detection and speaker detection, using solely visual information. The main idea in this work is to apply signal detection algorithms to a simple and easily extracted feature from the mouth region. We argue that the increased average value and standard deviation of the number of pixels with low intensities that the mouth region of a speaking person demonstrates can be used as visual cues for detecting visual speech. We then proceed in deriving a statistical algorithm that utilizes this fact for the efficient characterization of visual speech and silence in video sequences. Furthermore, we employ the lip activity detection method in order to determine the active speaker(s) in a multi-person environment.
Spyridon Siatras, Nikos Nikolaidis 0001, Michail Krinidis, Ioannis Pitas
IEEE Trans. Circuits Syst. Video Technol.4
2009 Color Texture Segmentation Based on the Modal Energy of Deformable Surfaces
abstract
This paper presents a new approach for the segmentation of color textured images, which is based on a novel energy function. The proposed energy function, which expresses the local smoothness of an image area, is derived by exploiting an intermediate step of modal analysis that is utilized in order to describe and analyze the deformations of a 3-D deformable surface model. The external forces that attract the 3-D deformable surface model combine the intensity of the image pixels with the spatial information of local image regions. The proposed image segmentation algorithm has two steps. First, a color quantization scheme, which is based on the node displacements of the deformable surface model, is utilized in order to decrease the number of colors in the image. Then, the proposed energy function is used as a criterion for a region growing algorithm. The final segmentation of the image is derived by a region merge approach. The proposed method was applied to the Berkeley segmentation database. The obtained results show good segmentation robustness, when compared to other state of the art image segmentation algorithms.
Michail Krinidis, Ioannis Pitas
IEEE Trans. Image Process.2
2009 Novel Multiclass Classifiers Based on the Minimization of the Within-Class Variance
abstract
In this paper, a novel class of multiclass classifiers inspired by the optimization of Fisher discriminant ratio and the support vector machine (SVM) formulation is introduced. The optimization problem of the so-called minimum within-class variance multiclass classifiers (MWCVMC) is formulated and solved in arbitrary Hilbert spaces, defined by Mercer's kernels, in order to find multiclass decision hyperplanes/surfaces. Afterwards, MWCVMCs are solved using indefinite kernels and dissimilarity measures via pseudo-Euclidean embedding. The power of the proposed approach is first demonstrated in the facial expression recognition of the seven basic facial expressions (i.e., anger, disgust, fear, happiness, sadness, and surprise plus the neutral state) problem in the presence of partial facial occlusion by using a pseudo-Euclidean embedding of Hausdorff distances and the MWCVMC. The experiments indicated a recognition accuracy rate achieved up to 99%. The MWCVMC classifiers are also applied to face recognition and other classification problems using Mercer's kernels.
Irene Kotsia, Ioannis Pitas, Stefanos Zafeiriou
IEEE Trans. Neural Networks2
2008 Texture and Shape Information Fusion for Facial Action Unit Recognition
abstract
A novel method that fuses texture and shape information to achieve Facial Action Unit (FAU) recognition from video sequences is proposed. In order to extract the texture information, a subspace method based on Discriminant Non- negative Matrix Factorization (DNMF) is applied on the difference images of the video sequence, calculated taking under consideration the neutral and the most expressive frame, to extract the desired classification label. The shape information consists of the deformed Candide facial grid (more specifically the grid node displacements between the neutral and the most expressive facial expression frame) that corresponds to the facial expression depicted in the video sequence. The shape information is afterwards classified using a two-class Support Vector Machine (SVM) system. The fusion of texture and shape information is performed using Median Radial Basis Functions (MRBFs) Neural Networks (NNs) in order to detect the set of present FAUs. The accuracy achieved in the Cohn-Kanade database is equal to 92.1% when recognizing the 17 FAUs that are responsible for facial expression development.
Irene Kotsia, Stefanos Zafeiriou, Nikos Nikolaidis 0001, Ioannis Pitas
ACHI4
2008 Motivating class-specific nonlinear projections for single and multiple view face verification
abstract
In this paper we motivate the use of class-specific nonlinear subspace methods for face verification. The problem of face verification is considered as a two-class problem (genuine versus impostor class). The typical Fisher's linear discriminant analysis (FLDA) gives only one or two projections in a two-class problem. This is a very strict limitation to the search of discriminant dimensions. As for the FLDA for N class problems (N > 2) the transformation is not person specific. In order to remedy these limitations of FLDA, exploit the individuality of human faces and take into consideration the fact that the distribution of facial images, under different viewpoints, illumination variations and facial expression is highly complex and non-linear, novel kernel discriminant algorithms are used. The new method was tested in the face verification problem using single and multiple view datasets and found to outperform other commonly used kernel approaches.
Georgios Goudelis, Stefanos Zafeiriou, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas
ICIP5
2008 Face recognition via adaptive discriminant clustering
abstract
This paper presents a methodology that tackles the face recognition problem by accommodating multiple clustering steps. At each clustering step, the test and training faces are projected to a discriminant space and the projected training data are partitioned into clusters using the k-means algorithm. Then a subset of the training data clusters is selected, based on how similar the faces in these clusters are to the test face. In the clustering step that follows a new discriminant space is defined by processing this subset and both the test and training data are projected to this space. This process is repeated until one final cluster is selected and the most similar, to the test face, face class contained is set as the identity match. The UMIST and XM2VTS face databases have been used to evaluate the algorithm and results indicate that the proposed framework provides a promising solution to the face recognition problem.
Marios Kyperountas, Anastasios Tefas, Ioannis Pitas
ICIP3
2008 Sparse human movement representation and recognition
abstract
In this paper a novel method for human movement representation and recognition is proposed. A movement type is regarded as a unique combination of basic movement patterns, the so-called dynemes. The fuzzy c-mean (FCM) algorithm is used to identify the dynemes in the input space and allow the expression of a posture in terms of these dynemes. In the so-called dyneme space, the sparse posture representations of a movement are combined to represent the movement as a single point in that space, and linear discriminant analysis (LDA) is further employed to increase movement type discrimination and compactness of representation. This method allows for simple Mahalanobis or cosine distance comparison of movements, taking implicitly into account time shifts and internal speed variations, and, thus, aiding the design of a real-time movement recognition algorithm.
Nikolaos Gkalelis, Anastasios Tefas, Ioannis Pitas
MMSP3
2008 An analysis of facial expression recognition under partial facial image occlusion
Irene Kotsia, Ioan Buciu, Ioannis Pitas
Image Vis. Comput.3
2008 Texture and shape information fusion for facial expression and facial action unit recognition
Irene Kotsia, Stefanos Zafeiriou, Ioannis Pitas
Pattern Recognit.3
2008 Dynamic training using multistage clustering for face recognition
Marios Kyperountas, Anastasios Tefas, Ioannis Pitas
Pattern Recognit.3
2008 Face-Based Digital Signatures for Video Retrieval
abstract
The characterization of a video segment by a digital signature is a fundamental task in video processing. It is necessary for video indexing and retrieval, copyright protection, and other tasks. Semantic video signatures are those that are based on high-level content information rather than on low-level features of the video stream. The major advantage of such signatures is that they are highly invariant to nearly all types of distortion. A major semantic feature of a video is the appearance of specific persons in specific video frames. Because of the great amount of research that has been performed on the subject of face detection and recognition, the extraction of such information is generally tractable, or will be in the near future. We have developed a method that uses the pre-extracted output of face detection and recognition to perform fast semantic query-by-example retrieval of video segments. We also give the results of the experimental evaluation of our method on a database of real video. One advantage of our approach is that the evaluation of similarity is convolution-based, and is thus resistant to perturbations in the signature and independent of the exact boundaries of the query segment.
Costas I. Cotsaces, Nikos Nikolaidis 0001, Ioannis Pitas
IEEE Trans. Circuits Syst. Video Technol.3
2008 Combining Fuzzy Vector Quantization With Linear Discriminant Analysis for Continuous Human Movement Recognition
abstract
In this paper, a novel method for continuous human movement recognition based on fuzzy vector quantization (FVQ) and linear discriminant analysis (LDA) is proposed. We regard a movement as a unique combination of basic movement patterns, the so-called dynemes. The proposed algorithm combines FVQ and LDA to discover the most discriminative dynemes as well as represent and discriminate the different human movements in terms of these dynemes. This method allows for simple Mahalanobis or cosine distance comparison of not aligned human movements, taking into account implicitly time shifts and internal speed variations, and, thus, aiding the design of a real-time continuous human movement recognition algorithm. The effectiveness and robustness of this method is shown by experimental results on a standard dataset with videos captured under real conditions, and on a new video dataset created using motion capture data.
Nikolaos Gkalelis, Anastasios Tefas, Ioannis Pitas
IEEE Trans. Circuits Syst. Video Technol.3
2008 Automated Facial Pose Extraction From Video Sequences Based on Mutual Information
abstract
Estimation of the facial pose in video sequences is one of the major issues in many vision systems such as face-based biometrics, scene understanding for humans, and others. The proposed method uses a novel pose estimation algorithm based on mutual information to extract any required facial poses from video sequences. The method extracts the poses automatically and classifies them according to view angle. Experimental results on the XM2VTS video database and on a new database created for the needs of this research indicated a pose classification rate of 99.2% while it was shown that it outperforms a principal component analysis reconstruction method that was used as a benchmark.
Georgios Goudelis, Anastasios Tefas, Ioannis Pitas
IEEE Trans. Circuits Syst. Video Technol.3
2008 Audio-Assisted Movie Dialogue Detection
abstract
Abstract—An audio-assisted system is investigated that detects if a movie scene is a dialogue or not. The system is based on actor indicator functions. That is, functions which define if an actor speaks at a certain time instant. In particular, the cross-correlation and the magnitude of the corresponding the cross-power spectral density of a pair of indicator functions are input to various classifiers, such as voted perceptrons, radial basis function networks, random trees, and support vector machines for dialogue/non-dialogue detection. To boost classifier efficiency AdaBoost is also exploited. The aforementioned classifiers are trained using ground truth indicator functions determined by human annotators for 41 dialogue and another 20 non-dialogue audio instances. For testing, actual indicator functions are derived by applying audio activity detection and actor clustering to audio recordings. 23 instances are randomly chosen among the aforementioned 41 dialogue instances, 17 of which correspond to dialogue scenes and 6 to non-dialogue ones. Accuracy ranging between 0.739 and 0.826 is reported. Index Terms—Audio activity detection, cross-correlation, crosspower spectral density, dialogue detection, indicator functions, speaker clustering. I.
Margarita Kotti, Dimitrios Ververidis, Georgios Evangelopoulos, Yannis Panagakis, Constantine Kotropoulos, Petros Maragos, Ioannis Pitas
IEEE Trans. Circuits Syst. Video Technol.7
2008 Camera Motion Estimation Using a Novel Online Vector Field Model in Particle Filters
abstract
In this paper, a novel algorithm for parametric camera motion estimation is introduced. More particularly, a novel stochastic vector field model is proposed, which can handle smooth motion patterns derived from long periods of stable camera motion and can also cope with rapid camera motion changes and periods when the camera remains still. The stochastic vector field model is established from a set of noisy measurements, such as motion vectors derived, e.g., from block matching techniques, in order to provide an estimation of the subsequent camera motion in the form of a motion vector field. A set of rules for a robust and online update of the camera motion model parameters is also proposed, based on the expectation maximization algorithm. The proposed model is embedded in a particle filters framework in order to predict the future camera motion based on current and prior observations. We estimate the subsequent camera motion by finding the optimum affine transform parameters so that, when applied to the current video frame, the resulting motion vector field to approximate the one estimated by the stochastic model. Extensive experimental results verify the usefulness of the proposed scheme in camera motion pattern classification and in the accurate estimation of the 2D affine camera transform motion parameters. Moreover, the camera motion estimation has been incorporated into an object tracker in order to investigate if the new schema improves its tracking efficiency, when camera motion and tracked object motion are combined.
Symeon Nikitidis, Stefanos Zafeiriou, Ioannis Pitas
IEEE Trans. Circuits Syst. Video Technol.3
2008 Piecewise Linear Digital Curve Representation and Compression Using Graph Theory and a Line Segment Alphabet
abstract
The use of an alphabet of line segments to compose a curve is a possible approach for curve data compression. Many approaches are developed with the drawback that they can process simple curves only. Curves having more sophisticated topology with self-intersections can be handled by methods considering recursive decomposition of the canvas containing the curve. In this paper, we propose a graph theory-based algorithm for tracing the curve directly to eliminate the decomposition needs. This approach obviously improves the compression performance, as longer line segments can be used. We tune our method further by selecting optimal turns at junctions during tracing the curve. We assign a polygon approximation to the curve which consists of letters coming from an alphabet of line segments. We also discuss how other application fields can take advantage of the provided curve description scheme.
András Hajdu, Ioannis Pitas
IEEE Trans. Image Process.2
2008 Discriminant Graph Structures for Facial Expression Recognition
abstract
In this paper, a series of advances in elastic graph matching for facial expression recognition are proposed. More specifically, a new technique for the selection of the most discriminant facial landmarks for every facial expression (discriminant expression-specific graphs) is applied. Furthermore, a novel kernel-based technique for discriminant feature extraction from graphs is presented. This feature extraction technique remedies some of the limitations of the typical kernel Fisher discriminant analysis (KFDA) which provides a subspace of very limited dimensionality (i.e., one or two dimensions) in two-class problems. The proposed methods have been applied to the Cohn-Kanade database in which very good performance has been achieved in a fully automatic manner.
Stefanos Zafeiriou, Ioannis Pitas
IEEE Trans. Multim.2
2008 Nonnegative Matrix Factorization in Polynomial Feature Space
abstract
Plenty of methods have been proposed in order to discover latent variables (features) in data sets. Such approaches include the principal component analysis (PCA), independent component analysis (ICA), factor analysis (FA), etc., to mention only a few. A recently investigated approach to decompose a data set with a given dimensionality into a lower dimensional space is the so-called nonnegative matrix factorization (NMF). Its only requirement is that both decomposition factors are nonnegative. To approximate the original data, the minimization of the NMF objective function is performed in the Euclidean space, where the difference between the original data and the factors can be minimized by employing L(2)-norm. In this paper, we propose a generalization of the NMF algorithm by translating the objective function into a Hilbert space (also called feature space) under nonnegativity constraints. With the help of kernel functions, we developed an approach that allows high-order dependencies between the basis images while keeping the nonnegativity constraints on both basis images and coefficients. Two practical applications, namely, facial expression and face recognition, show the potential of the proposed approach.
Ioan Buciu, Nikos Nikolaidis 0001, Ioannis Pitas
IEEE Trans. Neural Networks3
2007 Facial Expression Recognition in Videos using a Novel Multi-Class Support Vector Machines Variant
abstract
In this paper, a novel class of support vector machines (SVM) is introduced to deal with facial expression recognition. The proposed classifier incorporates statistic information about the classes under examination into the classical SVM. The developed system performs facial expression recognition in facial videos. The grid tracking and deformation algorithm used tracks the Candide grid over time as the facial expression evolves, until the frame that corresponds to the greatest facial expression intensity. The geometrical displacement of Candide nodes is used as an input to the bank of novel SVM classifiers, that are utilized to recognize the six basic facial expressions. The experiments on the Cohn-Kanade database show a recognition accuracy of 98.2%.
Irene Kotsia, Nikos Nikolaidis 0001, Ioannis Pitas
ICASSP (2)3
2007 Discriminant Graph Structures for Face Verification
abstract
Elastic graph matching is one of the most well known techniques for frontal face recognition/verification and one of the few techniques that can be combined successfully with fully automatic face localization and alignment methods. In this paper, we propose an algorithm for finding the most discriminant features upon a person's face and a person-specific graph is placed in the spatial coordinates that correspond to these discriminant features. We illustrate the improvements in performance by applying the proposed method in frontal face verification using the XM2VTS database.
Stefanos Zafeiriou, Anastasios Tefas, Ioannis Pitas
ICASSP (1)3
2007 A Novel Kernel Discriminant Analysis for Face Verification
abstract
In this paper a novel non-linear subspace method for face verification is proposed. The problem of face verification is considered as a two-class problem (genuine versus impostor class). The typical Fisher's linear discriminant analysis (FLDA) gives only one or two projections in a two-class problem. This is a very strict limitation to the search of discriminant dimensions. As for the FLDA for N class problems (N is greater than two) the transformation is not person specific. In order to remedy these limitations of FLDA, exploit the individuality of human faces and take into consideration the fact that the distribution of facial images, under different viewpoints, illumination variations and facial expression is highly complex and non-linear, novel kernel discriminant algorithms are proposed. The new methods are tested in the face verification problem using the XM2VTS database where it is verified that they outperform other commonly used kernel approaches.
Georgios Goudelis, Stefanos Zafeiriou, Anastasios Tefas, Ioannis Pitas
ICIP (4)4
2007 Content Adaptive Heterogeneous Snakes
abstract
Active contour (snake) approaches have been proved to be efficient tools to extract object boundary precisely. One drawback of these methods is that state-of-the-art snake algorithms are usually computationally expensive. To overcome this difficulty, one approach is to reduce spatial resolution to make snake iteration faster. The disadvantage of this method is that the reduction of resolution ignores image content which may lead to losing useful image data. In this paper we propose a snake iteration method on quadtree representations derived from the image. In this way we can reduce resolution in a special and adaptive way. Once the iteration is finished on the sparse quadtree grid, final adjustments can be made by performing pixelwise snake iteration steps.
András Hajdu, Ioannis Pitas
ICIP (1)2
2007 Compression Optimized Tracing of Digital Curves using Graph Theory
abstract
The use of an alphabet of line segments to compose a curve is a possible approach for curve data compression. An existing state-of-the-art method considers a quadtree decomposition of the curve to perform the substitution of the curve parts from the alphabet of line segments. In this paper, we propose a graph theory based algorithm for tracing the curve directly to eliminate the quadtree decomposition needs. This approach obviously improves the compression efficiency, as longer line segments can be used. We tune our method further by selecting optimal turns at junctions during tracing the curve. We also discuss briefly how other application fields can take advantage of the presented approach.
András Hajdu, Ioannis Pitas
ICIP (6)2
2007 Face Verification using Locally Linear Discriminant Models
abstract
When linear discriminant analysis (LDA) is employed, the correct classification of a sample heavily depends on having an adequately large training set. This is often not possible in practical applications, such as person verification, where the lack of sufficient training samples causes improper estimation of a linear separation hyper-plane between the two classes. To overcome this shortcoming a novel algorithm that can handle the verification problem more efficiently than traditional LDA is presented. The dimensionality of the samples is reduced by breaking them down, thus creating subsets of smaller dimensionality feature vectors, and applying discriminant analysis on each subset. The resulting discriminant weight sets are themselves weighted under a normalization criterion, making the discriminant functions continuous in this sense. A series of simulations that formulate the face verification problem illustrate the cases for which our method outperforms traditional LDA and various statistical observations are made about the discriminant coefficients that are generated.
Marios Kyperountas, Anastasios Tefas, Ioannis Pitas
ICIP (4)3
2007 Using Deformable Surface Models to Derive a DCT-Like 2D Transform
abstract
This paper introduces a 2D discrete, non-separable transform for image processing, which can be regarded as a combination of the well known discrete cosine transform (DCT) with an analytically derived quantization table that includes a compression ratio selection parameter. A 3D deformable surface model is used to approximate the image intensity and the introduced discrete transform is an intermediate step of the explicit surface deformation governing equations. The proposed transform is applied to lossy image compression and the obtained results are compared to those of a DCT-based compression scheme.
Michail Krinidis, Nikos Nikolaidis 0001, Ioannis Pitas
ICME3
2007 A Face Tracker Trajectories Clustering Using Mutual Information
abstract
In this paper we propose an algorithm for face tracker's trajectories clustering. Our approach is based on the mutual information of the images and more precisely its normalized version (NMI). We make use of 2 color channels from the HSV space (hue and saturation) in order to calculate a 4D joint histogram and therefore calculate the mutual information. In this paper we also develop an algorithm where we apply robust heuristics and make use of a tracker information in order to diminish dimensionality and augment accuracy of our results. It is a supervised clustering algorithm which is therefore used (fuzzy c-means) in order to gather same trajectories and same faces together.
Nicholas Vretos, Vassilios Solachidis, Ioannis Pitas
MMSP3
2007 A neural network approach to audio-assisted movie dialogue detection
Margarita Kotti, Emmanouil Benetos, Constantine Kotropoulos, Ioannis Pitas
Neurocomputing4
2007 Combining text and link analysis for focused crawling - An application for vertical search engines
George Almpanidis, Constantine Kotropoulos, Ioannis Pitas
Inf. Syst.3
2007 Erratum to "NMF, LNMF, and DNMF modeling of neural receptive fields involved in human facial expression perception" [J. Visual Communication and Image Representation 17 (2006) 958-969]
Ioan Buciu, Ioannis Pitas
J. Vis. Commun. Image Represent.2
2007 The discriminant elastic graph matching algorithm applied to frontal face verification
Stefanos Zafeiriou, Anastasios Tefas, Ioannis Pitas
Pattern Recognit.3
2007 The discrete modal transform and its application to lossy image compression
Michail Krinidis, Nikos Nikolaidis 0001, Ioannis Pitas
Signal Process. Image Commun.3
2007 2-D Feature-Point Selection and Tracking Using 3-D Physics-Based Deformable Surfaces
abstract
This paper presents a novel approach for selecting and tracking feature points in video sequences. In this approach, the image intensity is represented by a 3-D deformable surface model. The proposed approach relies on selecting and tracking feature points by exploiting the so-called generalized displacement vector that appears in the explicit surface deformation governing equations. This vector is proven to be a combination of the output of various line- and edge-detection masks, thus leading to distinct, robust features. The proposed method was compared, in terms of tracking accuracy and robustness, with a well-known tracking algorithm, Kanade-Lucas-Tomasi (KLT), and a tracking algorithm based on scale-invariant feature transform (SIFT) features. The proposed method was experimentally shown to be more precise and robust than both KLT and SIFT tracking. Moreover, the feature-point selection scheme was tested against the SIFT and Harris feature points, and it was demonstrated to provide superior results.
Michail Krinidis, Nikos Nikolaidis 0001, Ioannis Pitas
IEEE Trans. Circuits Syst. Video Technol.3
2007 Class-Specific Kernel-Discriminant Analysis for Face Verification
abstract
In this paper, novel nonlinear subspace methods for face verification are proposed. The problem of face verification is considered as a two-class problem (genuine versus impostor class). The typical Fisher's linear discriminant analysis (FLDA) gives only one or two projections in a two-class problem. This is a very strict limitation to the search of discriminant dimensions. As for the FLDA forNclass problems (Nis greater than two), the transformation is not person specific. In order to remedy these limitations of FLDA, exploit the individuality of human faces and take into consideration the fact that the distribution of facial images, under different viewpoints, illumination variations, and facial expression is highly complex and nonlinear, novel kernel-discriminant algorithms are proposed. The new methods are tested in the face verification problem using the XM2VTS, AR, ORL, Yale, and UMIST databases where it is verified that they outperform other commonly used kernel approaches such as kernel-PCA (KPCA), kernel direct discriminant analysis (KDDA), complete kernel Fisher's discriminant analysis (CKFDA), the two-class KDDA, CKFDA, and other two-class and multiclass variants of kernel-discriminant analysis based on Fisher's criterion.
Georgios Goudelis, Stefanos Zafeiriou, Anastasios Tefas, Ioannis Pitas
IEEE Trans. Inf. Forensics Secur.4
2007 A Novel Discriminant Non-Negative Matrix Factorization Algorithm With Applications to Facial Image Characterization Problems
abstract
The methods introduced so far regarding discriminant non-negative matrix factorization (DNMF) do not guarantee convergence to a stationary limit point. In order to remedy this limitation, a novel DNMF method is presented that uses projected gradients. The proposed algorithm employs some extra modifications that make the method more suitable for classification tasks. The usefulness of the proposed technique to frontal face verification and facial expression recognition problems is demonstrated.
Irene Kotsia, Stefanos Zafeiriou, Ioannis Pitas
IEEE Trans. Inf. Forensics Secur.3
2007 Learning Discriminant Person-Specific Facial Models Using Expandable Graphs
abstract
In this paper, a novel algorithm for finding discriminant person-specific facial models is proposed and tested for frontal face verification. The most discriminant features of a person's face are found and a deformable model is placed in the spatial coordinates that correspond to these discriminant features. The discriminant deformable models, for verifying the person's identity, that are learned through this procedure are elastic graphs that are dense in the facial areas considered discriminant for a specific person and sparse in other less significant facial areas. The discriminant graphs are enhanced by a discriminant feature selection method for the graph nodes in order to find the most discriminant jet features. The proposed approach significantly enhances the performance of elastic graph matching in frontal face verification
Stefanos Zafeiriou, Anastasios Tefas, Ioannis Pitas
IEEE Trans. Inf. Forensics Secur.3
2007 Optimal Approach for Fast Object-Template Matching
abstract
This paper proposes a novel algorithm for an optimal reduction of object description for object matching purposes. Our aim is to decrease the computation needs by considering simplified objects, thus reducing the number of pixels involved in the matching process. We develop the appropriate theoretical background based on centroidal Voronoi tessellations. Its use within the chamfer matching framework is also discussed. We present experimental results regarding the performance of this approach for 2-D contour and region-like object matching. As a special case, we investigate how the snake based representation of target objects can be employed in chamfer matching. The experimental results concern the use of object part matching for recognizing humans and show how the proposed simplification leads to valid replacements of the original templates.
András Hajdu, Ioannis Pitas
IEEE Trans. Image Process.2
2007 Facial Expression Recognition in Image Sequences Using Geometric Deformation Features and Support Vector Machines
abstract
In this paper, two novel methods for facial expression recognition in facial image sequences are presented. The user has to manually place some of Candide grid nodes to face landmarks depicted at the first frame of the image sequence under examination. The grid-tracking and deformation system used, based on deformable models, tracks the grid in consecutive video frames over time, as the facial expression evolves, until the frame that corresponds to the greatest facial expression intensity. The geometrical displacement of certain selected Candide nodes, defined as the difference of the node coordinates between the first and the greatest facial expression intensity frame, is used as an input to a novel multiclass Support Vector Machine (SVM) system of classifiers that are used to recognize either the six basic facial expressions or a set of chosen Facial Action Units (FAUs). The results on the Cohn-Kanade database show a recognition accuracy of 99.7% for facial expression recognition using the proposed multiclass SVMs and 95.1% for facial expression recognition based on FAU detection.
Irene Kotsia, Ioannis Pitas
IEEE Trans. Image Process.2
2007 Minimum Class Variance Support Vector Machines
abstract
In this paper, a modified class of support vector machines (SVMs) inspired from the optimization of Fisher's discriminant ratio is presented, the so-called minimum class variance SVMs (MCVSVMs). The MCVSVMs optimization problem is solved in cases in which the training set contains less samples that the dimensionality of the training vectors using dimensionality reduction through principal component analysis (PCA). Afterward, the MCVSVMs are extended in order to find nonlinear decision surfaces by solving the optimization problem in arbitrary Hilbert spaces defined by Mercer's kernels. In that case, it is shown that, under kernel PCA, the nonlinear optimization problem is transformed into an equivalent linear MCVSVMs problem. The effectiveness of the proposed approach is demonstrated by comparing it with the standard SVMs and other classifiers, like kernel Fisher discriminant analysis in facial image characterization problems like gender determination, eyeglass, and neutral facial expression detection.
Stefanos Zafeiriou, Anastasios Tefas, Ioannis Pitas
IEEE Trans. Image Process.3
2007 Enhanced Eigen-Audioframes for Audiovisual Scene Change Detection
abstract
In this paper, a novel audio-visual scene change detection algorithm is presented and evaluated experimentally. An enhanced set of eigen-audioframes is created that is related to an audio signal subspace, where audio background changes are easily discovered. An analysis is presented that justifies why this subspace favors scene change detection. Additionally, a novel process is developed in order to detect audio scene change candidates in this subspace. Visual information is used to align audio scene change indications with neighboring video shot changes and, accordingly, to reduce the false alarm rate of the audio-only scene change detection. Moreover, video fade effects are identified and used independently in order to track scene changes. The false alarm rate is reduced further by extracting acoustic features in order to verify that the scene change indications are valid. The detection methodology was tested on newscast videos provided by the TRECVID2003 video test set. The experimental results demonstrate that the proposed method achieves an F-measure exceeding 0.85. Accordingly, it effectively tackles the scene change detection problem
Marios Kyperountas, Constantine Kotropoulos, Ioannis Pitas
IEEE Trans. Multim.3
2007 Watermarking Digital 3-D Volumes in the Discrete Fourier Transform Domain
abstract
In this paper, a robust blind watermarking method for 3-D volumes is presented. A bivalued watermark is embedded in the Fourier transform magnitude of the 3-D volume. The Fourier domain has been selected because of its scaling and rotation invariance. Furthermore, in order to decrease the detection time, a special symmetry of the watermark is exploited. The proposed method is proven to be resistant to 3-D lowpass filtering, noise addition, scaling, translation, cropping and rotation. Experimental results prove the robustness of this method against the above-mentioned attacks.
Vassilios Solachidis, Ioannis Pitas
IEEE Trans. Multim.2
2007 Weighted Piecewise LDA for Solving the Small Sample Size Problem in Face Verification
abstract
A novel algorithm that can be used to boost the performance of face-verification methods that utilize Fisher's criterion is presented and evaluated. The algorithm is applied to similarity, or matching error, data and provides a general solution for overcoming the "small sample size" (SSS) problem, where the lack of sufficient training samples causes improper estimation of a linear separation hyperplane between the classes. Two independent phases constitute the proposed method. Initially, a set of weighted piecewise discriminant hyperplanes are used in order to provide a more accurate discriminant decision than the one produced by the traditional linear discriminant analysis (LDA) methodology. The expected classification ability of this method is investigated throughout a series of simulations. The second phase defines proper combinations for person-specific similarity scores and describes an outlier removal process that further enhances the classification ability. The proposed technique has been tested on the M2VTS and XM2VTS frontal face databases. Experimental results indicate that the proposed framework greatly improves the face-verification performance.
Marios Kyperountas, Anastasios Tefas, Ioannis Pitas
IEEE Trans. Neural Networks3
2007 An Optimal Detector Structure for the Fourier Descriptors Domain Watermarking of 2D Vector Graphics
abstract
Abstract-Polygonal lines constitute a key graphical primitive in 2D vector graphics data. Thus, the ability to apply a digital watermark to such an entity would enable the watermarking of cartoons, drawings, and Geographical Information Systems (GIS) data in vector graphics format. This paper builds on and extends an existing algorithm that achieves polygonal line watermarking by modifying the Fourier descriptors magnitude in an imperceptible way. Watermarks embedded by this technique can be detected in rotated, translated, scaled, or reflected polygonal lines. The detection of such watermarks had been previously carried out through a correlator detector. In this paper, analysis of the statistics of the Fourier descriptors is exploited to devise an optimal blind detector. Furthermore, the problem of watermarking multiple lines, as well as other implementation issues are being addressed. Experimental results verify the imperceptibility and robustness of the proposed method.
Víctor Rodríguez-Doncel, Nikos Nikolaidis 0001, Ioannis Pitas
IEEE Trans. Vis. Comput. Graph.3
2006 Temporal Video Segmentation by Graph Partitioning
abstract
A novel temporal video segmentation method that, in addition to abrupt cuts, can detect with very high accuracy gradual transitions such as dissolves, fades and wipes is proposed. The method relies on evaluating mutual information between multiple pairs of frames within a certain temporal frame window. This way we create a graph where the frames are nodes and the measures of similarity correspond to the weights of the edges. By finding and disconnecting the weak connections between nodes we separate the graph to subgraphs ideally corresponding to the shots. Experiments on TRECVID2004 video test set containing different types of shot transitions and significant object and camera motion inside the shots prove that the method is very efficient.
Zuzana Cernekova, Nikos Nikolaidis 0001, Ioannis Pitas
ICASSP (2)3
2006 Video Indexing by Face Occurrence-Based Signatures
abstract
The extraction of a digital signature from a video segment in order to uniquely identify it, is often a necessary prerequisite for video indexing, copyright protection and other tasks. Semantic video signatures are those that are based on high-level content information rather than on low-level features of the video stream, their major advantage being that they are invariant to nearly all types of distortion. Since a major semantic feature of a video is the appearance of specific people in specific frames, we have developed a method that uses the pre-extracted output of face detection and recognition to perform fast semantic indexing and retrieval of video segments. We give the results of the experimental evaluation of our method on an artificial database created using a probabilistic model of the creation of video
Costas I. Cotsaces, Nikos Nikolaidis 0001, Ioannis Pitas
ICASSP (2)3
2006 Fusion of Geometrical and Texture Information for Facial Expression Recognition
abstract
A novel method based on geometrical and texture information is proposed for facial expression recognition from video sequences. The discriminant non-negative matrix factorization (DNMF) algorithm is applied at the image of the last frame of the video sequence, corresponding to the greatest intensity of the facial expression, thus extracting the texture information. A support vector machines (SVMs) system is used for the classification of the geometrical information derived from tracking the Candide grid over the video sequence. The geometrical information consists of the differences of the node coordinates between the neutral (first) and the fully expressed facial expression (last) video frame. The fusion of texture and geometrical information obtained is performed using SVMs. The accuracy achieved is 98,7% when recognizing the six basic facial expressions.
Irene Kotsia, Nikos Nikolaidis 0001, Ioannis Pitas
ICIP3
2006 Optimized Chamfer Matching for Snake-Based Image Contour Representations
abstract
In this paper we present a novel method on how to take advantage of the snake representation of target objects, when doing chamfer matching for detection/recognition purposes. In this case several time-consuming steps of classic chamfer matching approaches can be simplified. Moreover, we investigate the possibility of involving fewer pixels from both the target and template object to speed up computations. We introduce an optimization method for such an object reduction, which is valid also in the general application scheme of chamfer matching. Finally, we present our experimental results regarding human body detection
András Hajdu, Athanasios Roubies, Ioannis Pitas
ICME3
2006 Virtual Dental Patient: a System for Virtual Teeth Drilling
abstract
This paper introduces, a virtual teeth drilling system named virtual dental patient designed to aid dentists in getting acquainted with the teeth anatomy, the handling of drilling instruments and the challenges associated with the drilling procedure. The basic aim of the system is to be used for the training of dental students. The application features a 3D model of the face and the oral cavity that can be adapted to the characteristics of a specific person and animated. Drilling using a haptic device is performed on realistic teeth models (constructed from real data), within the oral cavity. Results and intermediate steps of the drilling procedure can be saved for future use
Ioannis Marras, Leontios Papaleontiou, Nikos Nikolaidis 0001, Kleoniki Lyroudia, Ioannis Pitas
ICME5
2006 Image Replica Detection using R-Trees and Linear Discriminant Analysis
abstract
In this paper a novel system for image replica detection is presented. The system uses color-based descriptors in order to extract robust features for image representation. These features are used for indexing the images in a database using an R-tree. When a query about whether a test image is a replica of an image in the database is submitted, the R-tree is traversed and a set of candidate images is retrieved. Then, in order to obtain a single result and at the same time reduce the number of decision errors the system is enhanced with linear discriminant analysis (LDA). The conducted experiments show that the proposed approach is very promising
Spiros Nikolopoulos, Stefanos Zafeiriou, Panagiotis Sidiropoulos, Nikos Nikolaidis 0001, Ioannis Pitas
ICME5
2006 A Mutual Information based Face Clustering Algorithm for Movies
abstract
In this paper a new approach for face clustering is developed. Mutual information and joint entropy are exploited in order to create a metric for the clustering process. The way the joint entropy and the mutual information are calculated gives some interesting properties to the aforementioned metric, which guarantees some robustness against standard noisy transformation such as scaling, cropping and pose changes. A slight preprocessing of the input face images is done in order to undertake problems that arise from detector's known errors
Nicholas Vretos, Vassilios Solachidis, Ioannis Pitas
ICME3
2006 On the initialization of the DNMF algorithm
abstract
A subspace supervised learning algorithm named discriminant non-negative matrix factorization (DNMF) has been recently proposed for classifying human facial expressions. It decomposes images into a set of basis images and corresponding coefficients. Usually, the algorithm starts with random basis image and coefficient initialization. Then, at each iteration, both basis images and coefficients are updated to minimize the underlying cost function. The algorithm may need several thousands of iterations to obtain cost function minimization. We provide a way to significantly improve the speed of the algorithm convergence by constructing initial basis images that meet the sparseness and orthogonality requirements and approximate the final minimization solution. To experimentally evaluate the new approach, we have applied DNMF using the random and the proposed initialization procedure to recognize six basic facial expressions. While fewer iteration steps are needed with the proposed initialization, the recognition accuracy remains within satisfactory levels.
Ioan Buciu, Nikos Nikolaidis 0001, Ioannis Pitas
ISCAS3
2006 NMF, LNMF, and DNMF modeling of neural receptive fields involved in human facial expression perception
Ioan Buciu, Ioannis Pitas
J. Vis. Commun. Image Represent.2
2006 Demonstrating the stability of support vector machines for classification
Ioan Buciu, Constantine Kotropoulos, Ioannis Pitas
Signal Process.3
2006 Information theory-based shot cut/fade detection and video summarization
abstract
New methods for detecting shot boundaries in video sequences and for extracting key frames using metrics based on information theory are proposed. The method for shot boundary detection relies on the mutual information (MI) and the joint entropy (JE) between the frames. It can detect cuts, fade-ins and fade-outs. The detection technique was tested on the TRECVID2003 video test set having different types of shots and containing significant object and camera motion inside the shots. It is demonstrated that the method detects both fades and abrupt cuts with high accuracy. The information theory measure provides us with better results because it exploits the inter-frame information in a more compact way than frame subtraction. It was also successfully compared to other methods published in literature. The method for key frame extraction uses MI as well. We show that it captures satisfactorily the visual content of the shot.
Zuzana Cernekova, Ioannis Pitas, Christophoros Nikou
IEEE Trans. Circuits Syst. Video Technol.2
2006 Digital image processing techniques for the detection and removal of cracks in digitized paintings
abstract
An integrated methodology for the detection and removal of cracks on digitized paintings is presented in this paper. The cracks are detected by thresholding the output of the morphological top-hat transform. Afterward, the thin dark brush strokes which have been misidentified as cracks are removed using either a median radial basis function neural network on hue and saturation data or a semi-automatic procedure based on region growing. Finally, crack filling using order statistics filters or controlled anisotropic diffusion is performed. The methodology has been shown to perform very well on digitized paintings suffering from cracks.
I. Giakoumis, Nikos Nikolaidis 0001, Ioannis Pitas
IEEE Trans. Image Process.3
2006 Color-Based Descriptors for Image Fingerprinting
abstract
Typically, content-based image retrieval (CBIR) systems receive an image or an image description as input and retrieve images from a database that are similar to the query image in regard to properties such as color, texture, shape, or layout. A kind of system that did not receive much attention compared to CBIR systems, is one that searches for images that are not similar but exact copies of the same image that have undergone some transformation. In this paper, we present such a system referred to as an image fingerprinting system, since it aims to extract unique and robust image descriptors (in analogy to human fingerprints). We examine the use of color-based descriptors and provide comparisons for different quantization methods, histograms calculated using color-only and/or spatial-color information with different similarity measures. The system was evaluated with receiver operating characteristic (ROC) analysis on a large database of 919 original images consisting of randomly drawn art images and similar images from specific categories, along with 30 transformed images for each original, totaling 27570 images. The transformed images were produced with attacks that typically occur during digital image distribution, including different degrees of scaling, rotation, cropping, smoothing, additive noise and compression, as well as illumination contrast changes. Results showed a sensitivity of 96% at the small false positive fraction of 4% and a reduced sensitivity of 88% when 13% of all transformations involved changing the illuminance of the images. The overall performance of the system is encouraging for the use of color, and particularly spatial chromatic descriptors for image fingerprinting
Marios A. Gavrielides, Elena Sikudová, Ioannis Pitas
IEEE Trans. Multim.3
2006 Exploiting discriminant information in nonnegative matrix factorization with application to frontal face verification
abstract
In this paper, two supervised methods for enhancing the classification accuracy of the Nonnegative Matrix Factorization (NMF) algorithm are presented. The idea is to extend the NMF algorithm in order to extract features that enforce not only the spatial locality, but also the separability between classes in a discriminant manner. The first method employs discriminant analysis in the features derived from NMF. In this way, a two-phase discriminant feature extraction procedure is implemented, namely NMF plus Linear Discriminant Analysis (LDA). The second method incorporates the discriminant constraints inside the NMF decomposition. Thus, a decomposition of a face to its discriminant parts is obtained and new update rules for both the weights and the basis images are derived. The introduced methods have been applied to the problem of frontal face verification using the well-known XM2VTS database. Both methods greatly enhance the performance of NMF for frontal face verification.
Stefanos Zafeiriou, Anastasios Tefas, Ioan Buciu, Ioannis Pitas
IEEE Trans. Neural Networks4
2005 Facial expression analysis under partial occlusion
abstract
Six basic facial expressions are investigated when the human face is partially occluded, i.e. when the eyes and eyebrows or the mouth regions are occluded. Such occlusions occur when a person wears glasses (e.g. in VR application) or a mouth mask (e.g. in medical application). More specifically, we are interested in finding the part of the face that contains sufficient information in order to correctly classify these six expressions. Two facial image databases are employed in our experiments. Each image from the database is convolved with a set of Gabor filters having various orientations and frequencies. The new feature vectors are classified by using a maximum correlation classifier and the cosine similarity measure approaches. We find that, overall, the facial expression recognition method provides robustness against partial occlusion, the classification accuracy only decreasing from 89.7% (no occlusion) to 84% (eyes region occlusion) and 83.5% (mouth region occlusion) for the first database and from 94.5% (no occlusion) to 91.5% (eyes region occlusion) and 87.2% (mouth region occlusion) for the second database, respectively.
Ioan Buciu, Irene Kotsia, Ioannis Pitas
ICASSP (5)3
2005 Methods for improving discriminant analysis for face authentication
abstract
A novel algorithm that can be used to boost the performance of face authentication methods that utilize Fisher's criterion is presented. The algorithm is applied to matching error data and provides a general solution for overcoming the "small sample size" (SSS) problem, where the lack of sufficient training samples causes improper estimation of a linear separation hyperplane between the classes. Two independent phases constitute the proposed method. Initially, a set of locally linear discriminant models is used in order to calculate discriminant weights in a more accurate way than the traditional linear discriminant analysis (LDA) methodology. Additionally, defective discriminant coefficients are identified and reestimated. The second phase defines proper combinations for person-specific matching scores and describes an outlier removal process that enhances the classification ability. Our technique was tested on the M2VTS and XM2VTS frontal face databases. Experimental results indicate that the proposed framework greatly improves the authentication algorithm's performance.
Marios Kyperountas, Anastasios Tefas, Ioannis Pitas
ICASSP (2)3
2005 Enhanced transform-domain correlation-based audio watermarking
abstract
Various watermarking techniques have been proposed so far, aiming at the copyright protection of audio signals. Little effort has been made, however, in taking under consideration the spectrum of the watermark sequence itself and exploiting its frequency properties. An enhanced audio watermarking technique, based on correlation detection, is introduced in this paper, where high-frequency chaotic watermarks are multiplicatively embedded in the low frequencies of the DFT domain. A series of experiments have been conducted to demonstrate both detection reliability and robustness against attacks.
Anastasios Tefas, Alexia Giannoula, Nikos Nikolaidis 0001, Ioannis Pitas
ICASSP (2)4
2005 Real time facial expression recognition from image sequences using support vector machines
abstract
In this paper, two novel real-time methods are proposed for facial expression recognition in image sequences. The user manually places some of the Candide grid's points to the face depicted at the first frame. The grid adaptation system tracks the entire grid as the facial expression evolves through time, thus producing a grid that corresponds to the greatest intensity of the facial expression, as shown at the last frame. Certain points that are involved into creating the facial action units (FAUs) movements are selected. Their geometrical displacement information, defined as the coordinates' difference between the last and the first frame, is extracted to be the input to a bank of support vector machine (SVM) classifiers that are used to recognize either the six basic facial expressions or eight chosen FAUs. The results show a recognition accuracy of approximately 98% and 94% for direct and FAU based facial expression recognition, respectively.
Irene Kotsia, Ioannis Pitas
ICIP (2)2
2005 Object tracking based on morphological elastic graph matching
abstract
This paper presents a novel method for real-time tracking of objects in video sequences. Tracking is performed using the so-called morphological elastic graph matching algorithm. When applied to faces, initialization of the tracking algorithm is performed by means of a novel face detection and facial feature extraction step. The obtained results show good performance in scenes with complex background. Comparison with an existing feature-based tracking method using measures based on ground truth data proves the superiority of the proposed method.
Georgios N. Stamou, Nikos Nikolaidis 0001, Ioannis Pitas
ICIP (1)3
2005 Exploiting discriminant information in elastic graph matching
abstract
In this paper, we investigate the use of discriminant techniques in the elastic graph matching (EGM) algorithm. First we use discriminant analysis in the feature vectors of the nodes in order to find the most discriminant features. The similarity measure for discriminant feature vectors and the node deformation are combined in a discriminant manner in order to form a local similarity measure between nodes. Moreover, the local similarity values at the nodes of the elastic graph, are weighted by coefficients that are also derived by some discriminant analysis in order to form a total similarity measure between faces. We illustrate the improvements in performance in frontal face verification using a modified multiscale morphological analysis.
Stefanos Zafeiriou, Anastasios Tefas, Ioannis Pitas
ICIP (3)3
2005 Watermarking digital 3D volumes in the discrete Fourier transform domain
abstract
In this paper, a robust blind watermarking method for 3D volumes is presented. A bivalued watermark is embedded in the Fourier transform magnitude of the 3D volume. The Fourier domain has been selected because of its important properties in terms of scaling and rotation invariance. Furthermore, a special symmetry of the watermark is exploited, in order to decrease the detection time. The proposed method is proven to be resistant to 3D lowpass filtering, noise addition, scaling, translation, cropping and rotation. Experimental results prove the robustness of this method against the above-mentioned attacks.
Vassilios Solachidis, Ioannis Pitas
ICME2
2005 Fast free-vibration modal analysis of 2-D physics-based deformable objects
abstract
This paper presents an accurate, very fast approach for the deformations of two-dimensional physically based shape models representing open and closed curves. The introduced models are much faster than other deformable models (e.g., finite-element methods). The approach relies on the determination of explicit deformation governing equations that involve neither eigenvalue decomposition, nor any other computationally intensive numerical operation. The approach was evaluated and compared with another fast and accurate physics-based deformable shape odel, both in terms of deformation accuracy and computation time. The conclusion is that the introduced model is completely accurate and is deformed very fast on current personal computers (Pentium III), achieving more than 380 contour deformations per second.
Stelios Krinidis, Ioannis Pitas
IEEE Trans. Image Process.2
2005 Automated Evaluation of Her-2/neu Status in Breast Tissue From Fluorescent In Situ Hybridization Images
abstract
The evaluation of fluorescent in situ hybridization (FISH) images is one of the most widely used methods to determine Her-2/neu status of breast samples, a valuable prognostic indicator. Conventional evaluation is a difficult task since it involves manual counting of dots in multiple images. In this paper, we present a multistage algorithm for the automated classification of FISH images from breast carcinomas. The algorithm focuses not only on the detection of FISH dots per image, but also on combining results from multiple images taken from a slice for overall case classification. The algorithm includes mainly two stages for nuclei and dot detection respectively. The dot segmentation consists of a top-hat filtering stage followed by template matching to separate real signals from noise. Nuclei segmentation includes a nonlinearity correction step, global thresholding to identify candidate regions, and a geometric rule to distinguish between holes within a nucleus and holes between nuclei. Finally, the marked watershed transform is used to segment cell nuclei with markers detected as regional maxima of the distance transform. Combining the two stages allows the measurement of FISH signals ratio per cell nucleus and the collective classification of cases as positive or negative. The system was evaluated with receiver operating characteristic analysis and the results were encouraging for the further development of this method.
Francesco Raimondo, Marios A. Gavrielides, Georgia Karayannopoulou, Kleoniki Lyroudia, Ioannis Pitas, Ioannis Kostopoulos
IEEE Trans. Image Process.5
2005 Blind Robust Watermarking Schemes for Copyright Protection of 3D Mesh Objects
abstract
In this paper, two novel methods suitable for blind 3D mesh object watermarking applications are proposed. The first method is robust against 3D rotation, translation, and uniform scaling. The second one is robust against both geometric and mesh simplification attacks. A pseudorandom watermarking signal is cast in the 3D mesh object by deforming its vertices geometrically, without altering the vertex topology. Prior to watermark embedding and detection, the object is rotated and translated so that its center of mass and its principal component coincide with the origin and the z-axis of the Cartesian coordinate system. This geometrical transformation ensures watermark robustness to translation and rotation. Robustness to uniform scaling is achieved by restricting the vertex deformations to occur only along the r coordinate of the corresponding (r, theta, phi) spherical coordinate system. In the first method, a set of vertices that correspond to specific angles theta is used for watermark embedding. In the second method, the samples of the watermark sequence are embedded in a set of vertices that correspond to a range of angles in the theta domain in order to achieve robustness against mesh simplifications. Experimental results indicate the ability of the proposed method to deal with the aforementioned attacks.
Stefanos Zafeiriou, Anastasios Tefas, Ioannis Pitas
IEEE Trans. Vis. Comput. Graph.3
2004 Audio PCA in a novel multimedia scheme for scene change detection
abstract
A novel scene change detection algorithm is proposed in this paper that exploits both audio and video information. Audio frames are projected to the eigenspace and their distance from a reference noise eigenframe is calculated. An analysis is presented that explains why this subspace favors scene cut detection. Video information is used to align audio scene change indications with neighboring shot changes in the visual data, and accordingly to reduce the false alarm rate. Moreover, video fade effects are identified and used independently in order to track scene changes. The detection technique was tested on newscast videos provided by the TRECVID 2003 video test set. The experimental results show that the aforementioned methods used to process the audio and video information complement each other well in tackling the scene change detection problem.
Marios Kyperountas, Zuzana Cernekova, Constantine Kotropoulos, Marios A. Gavrielides, Ioannis Pitas
ICASSP (4)5
2004 Automatic emotional speech classification
abstract
Our purpose is to design a useful tool which can be used in psychology to automatically classify utterances into five emotional states such as anger, happiness, neutral, sadness, and surprise. The major contribution of the paper is to rate the discriminating capability of a set of features for emotional speech recognition. A total of 87 features has been calculated over 500 utterances from the Danish Emotional Speech database. The sequential forward selection method (SFS) has been used in order to discover a set of 5 to 10 features which are able to classify the utterances in the best way. The criterion used in SFS is the cross-validated correct classification score of one of the following classifiers: nearest mean and Bayes classifier where class pdf are approximated via Parzen windows or modelled as Gaussians. After selecting the 5 best features, we reduce the dimensionality to two by applying principal component analysis. The result is a 51.6% /spl plusmn/ 3% correct classification rate at 95% confidence interval for the five aforementioned emotions, whereas a random classification would give a correct classification rate of 20%. Furthermore, we find out those two-class emotion recognition problems whose error rates contribute heavily to the average error and we indicate that a possible reduction of the error rates reported in this paper would be achieved by employing two-class classifiers and combining them.
Dimitrios Ververidis, Constantine Kotropoulos, Ioannis Pitas
ICASSP (1)3
2004 A blind robust watermarking scheme for copyright protection of 3d mesh models
Stefanos Zafeiriou, Anastasios Tefas, Ioannis Pitas
ICIP3
2004 MPEG-4 compliant reproduction of face animation created in Maya
abstract
This work presents a method for extracting facial information encoded in MPEG-4 format from an animated face model in Maya. The method generates the appropriate data, as specified in the MPEG-4 standard, such as the face definition parameters (FDPs), the face animation parameters (FAPs) and the facial animation table (FAT). The described procedure was implemented as a plug-in for Maya, which requires as inputs only the animated face model and the correspondence between its vertices and the FDPs. The extracted data were tested on a publicly available FAP-player in order to demonstrate the faithful reproduction of the face animation originally created with the high-end 3D graphics modeller. The capabilities of the publicly available FAP player have also been augmented in order to enable loading any proprietary face model and its corresponding FAT.
Charalampos Laftsidis, Constantine Kotropoulos, Ioannis Pitas
ICME3
2004 Anatomically-based 3D face and oral cavity model for creating virtual medical patients
abstract
This work presents a new hierarchical, modular and scalable model of a human face and oral cavity based on anatomical data taken from the Visible Human Project (National Institute of Health). The described model can further be adapted to any particular face and oral cavity of a human head by means of a finite elements method (FEM). The final aim will be to construct functional virtual medical patients.
Georgios Moschos, Nikos Nikolaidis 0001, Ioannis Pitas
ICME3
2004 Searching relevant syllable context by clustering for alignment in modern Greek speech
abstract
We present some results on the optimal design of syllable databases for modern Greek. We show that speaker gender or stress are not crucial criteria for voiced syllables, whereas distance to silence has to be taken into account. Such syllable databases are exploited in an alignment system based on dynamic time warping for modern Greek speech. The overall architecture of the alignment system is also presented.
R. Rispoli, Constantine Kotropoulos, Ioannis Pitas
ICME3
2004 Automatic detection of vocal fold paralysis and edema
abstract
In this paper we propose a combined scheme of linear prediction analysis for feature extraction along with linear projection methods for feature reduction followed by known pattern recognition methods on the purpose of discriminating between normal and pathological voice samples. Two different cases of speech under vocal fold pathology are examined: vocal fold paralysis and vocal fold edema. Three known classifiers are tested and compared in both cases, namely the Fisher linear discriminant, the -nearest neighbor classifier, and the nearest mean classifier. The performance of each classifier is evaluated in terms of the probabilities of false alarm and detection or the receiver operating characteristic. The datasets used are part of a database of disordered speech developed by Massachusetts Eye and Ear Infirmary. The experimental results indicate that vocal fold paralysis and edema can easily be detected by any of the aforementioned classifiers.
Maria Marinaki, Constantine Kotropoulos, Ioannis Pitas, Nicos Maglaveras
INTERSPEECH3
2004 Marginal median SOM for document organization and retrieval
Apostolos Georgakis, Constantine Kotropoulos, Alexandros Xafopoulos, Ioannis Pitas
Neural Networks4
2004 Language identification in web documents using discrete HMMs
Alexandros Xafopoulos, Constantine Kotropoulos, George Almpanidis, Ioannis Pitas
Pattern Recognit.4
2004 Probabilistic multiple face detection and tracking using entropy measures
abstract
A joint probabilistic face detection and tracking algorithm, combining likelihood estimation and a prior probability, is proposed. The likelihood estimation scheme is based on the statistical training of sets of automatically generated feature points and a mutual information tracking cue, while the prior probability estimation is based on a Gaussian temporal model. The likelihood estimation process is the core of a multiple face detection scheme used to initialize the tracking process. The resulting system has been tested on real image sequences and is robust to significant partial occlusion and illumination changes.
Evangelos Loutas, Ioannis Pitas, Christophoros Nikou
IEEE Trans. Circuits Syst. Video Technol.2
2004 Vector rational interpolation schemes for erroneous motion field estimation applied to MPEG-2 error concealment
abstract
A study on the use of vector rational interpolation for the estimation of erroneously received motion fields of MPEG-2 predictively coded frames is undertaken in this paper, aiming further at error concealment (EC). Various rational interpolation schemes have been investigated, some of which are applied to different interpolation directions. One scheme additionally uses the boundary matching error and another one attempts to locate the direction of minimal/maximal change in the local motion field neighborhood. Another one further adopts bilinear interpolation principles, whereas a last one additionally exploits available coding mode information. The methods present temporal EC methods for predictively coded frames or frames for which motion information pre-exists in the video bitstream. Their main advantages are their capability to adapt their behavior with respect to neighboring motion information, by switching from linear to nonlinear behavior, and their real-time implementation capabilities, enabling them for real-time decoding applications. They are easily embedded in the decoder model to achieve concealment along with decoding and avoid post-processing delays. Their performance proves to be satisfactory for packet error rates up to 2% and for video sequences with different content and motion characteristics and surpass that of other state-of-the-art temporal concealment methods that also attempt to estimate unavailable motion information and perform concealment afterwards.
Sofia Tsekeridou, Faouzi Alaya Cheikh, Moncef Gabbouj, Ioannis Pitas
IEEE Trans. Multim.4
2003 CAML - A Universal Configuration Language for Dialogue Systems
Gergely Kovásznai, Constantine Kotropoulos, Ioannis Pitas
DEXA3
2003 Video shot segmentation using singular value decomposition
abstract
A new method for detecting shot boundaries in video sequences using singular value decomposition (SVD) is proposed. The method relies on performing singular value decomposition on the matrix A created from 3D histograms of single frames. We have used SVD for its capabilities to derive a low dimensional refined feature space from a high dimensional raw feature space, where pattern similarity can easily be detected. The method can detect cuts and gradual transitions, such as dissolves and fades, which cannot be detected easily by entropy measures.
Zuzana Cernekova, Constantine Kotropoulos, Ioannis Pitas
ICASSP (3)3
2003 Watermarking of 3D models using principal component analysis
abstract
A novel method for 3D model watermarking, robust to geometric distortions such as rotation, translation and scaling, is proposed. A ternary watermark is embedded in the vertex topology of a 3D model. A transformation of the model to an invariant space is proposed prior to watermark embedding. Simulation results indicate the ability of the proposed method to deal with the aforementioned attacks giving very good results.
Andreas Kalivas, Anastasios Tefas, Ioannis Pitas
ICASSP (5)3
2003 ICA and Gabor representation for facial expression recognition
abstract
Two hybrid systems for classifying seven categories of human facial expression are proposed. The first system combines independent component analysis (ICA) and support vector machines (SVMs). The original face image database is decomposed into linear combinations of several basis images, where the corresponding coefficients of these combinations are fed up into SVMs instead of an original feature vector comprised of grayscale image pixel values. The classification accuracy of this system is compared against that of baseline techniques that combine ICA with either two-class cosine similarity classifiers or two-class maximum correlation classifiers, when we classify facial expressions into these seven classes. We found that, ICA decomposition combined with SVMs outperforms the aforementioned baseline classifiers. The second system proposed operates in two steps: first, a set of Gabor wavelets (GWs) is applied to the original face image database and, second, the new features obtained are classified by using either SVMs or cosine similarity classifiers or maximum correlation classifier. The best facial expression recognition rate is achieved when Gabor wavelets are combined with SVMs.
Ioan Buciu, Constantine Kotropoulos, Ioannis Pitas
ICIP (2)3
2003 Optimal detection for multiplicative watermarks embedded in DFT domain
abstract
This paper deals with the statistical analysis of the behavior of a blind copyright protection watermarking system based on pseudorandom signals embedded in the magnitude of the Fourier transform of the host data. The host data that the watermark is embedded into is one-dimensional and nonwhite, following a specific model. The analysis performed involves theoretical evaluation of the statistics of the Fourier coefficients and an optimum detector design for multiplicative embedding. It is proved that the widely used correlator is not the optimum detector. Finally, experimental results are presented in order to show the proposed detector's efficiency versus that of the correlator detector.
Vassilios Solachidis, Ioannis Pitas
ICIP (2)2
2003 Video shot segmentation using singular value decomposition
abstract
A new method for detecting shot boundaries in video sequences using singular value decomposition (SVD) is proposed. The method relies on performing singular value decomposition on the matrix A created from 3D histograms of single frames. We have used SVD for its capabilities to derive a low dimensional refined feature space from a high dimensional raw feature space, where pattern similarity can easily be detected. The method can detect cuts and gradual transitions, such as dissolves and fades, which cannot be detected easily by entropy measures.
Zuzana Cernekova, Constantine Kotropoulos, Ioannis Pitas
ICME3
2003 Improving the detection reliability of correlation-based watermarking techniques
abstract
The performance of watermarking schemes based on correlation detection is closely related to the frequency characteristics of the watermark sequence. In order to improve both detection reliability and robustness against attacks, embedding of watermarks with high-frequency spectrum, in the low frequencies of the DFT domain, is introduced in this paper and theoretical analysis of correlation based watermarking techniques with multiplicative embedding is performed. The proposed watermarking framework is successfully applied to audio signals, demonstrating its superiority with respect to both robustness and inaudibility. Experiments are conducted, in order to verify the validity of the theoretical analysis results.
Alexia Giannoula, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas
ICME4
2003 Watermarking of 3D models using principal component analysis
abstract
A novel method for 3D model watermarking robust to geometric distortions such as rotation, translation and scaling is proposed. A ternary watermark is embedded in the vertex topology of a 3D model. A transformation of the model to an invariant space is proposed prior to watermark embedding. Simulation results indicate the ability of the proposed method to deal with the aforementioned attacks giving very good results.
Andreas Kalivas, Anastasios Tefas, Ioannis Pitas
ICME3
2003 Segmentation of ultrasonic images using Support Vector Machines
Constantine Kotropoulos, Ioannis Pitas
Pattern Recognit. Lett.2
2003 Preface
Constantine Kotropoulos, Ioannis Pitas, Athina P. Petropulu
Pattern Recognit. Lett.2
2003 Wavelet packets-based digital watermarking for image verification and authentication
Alexandre H. Paquet, Rabab K. Ward, Ioannis Pitas
Signal Process.3
2003 Asymptotically optimal detection for additive watermarking in the DCT and DWT domains
abstract
Most of the watermarking schemes that have been proposed until now employ a correlation detector (matched filter). The current paper proposes a new detector scheme that can be applied in the case of additive watermarking in the DCT (discrete cosine transform) or DWT (discrete wavelet transform) domain. Certain properties of the probability density function of the coefficients in these domains are exploited. Thus, an asymptotically optimal detector is constructed based on well known results of the detection theory. Experimental results prove the superiority of the proposed detector over the correlation detector.
Athanasios Nikolaidis, Ioannis Pitas
IEEE Trans. Image Process.2
2003 A global energy function for the alignment of serially acquired slices
abstract
An accurate, computationally efficient, and fully automated algorithm for the alignment of two-dimensional (2-D) serially acquired sections forming a three-dimensional (3-D) volume is presented. The approach relies on the optimization of a global energy function, based on the object shape, measuring the similarity between a slice and its neighborhood in the 3-D volume. Slice similarity is computed using the distance transform measure in both directions. No particular direction is privileged in the method avoiding global offsets, biases in the estimation and error propagation. The method was evaluated on real images [medical, biological, and other computerized tomography (CT) scanned 3-D data] and the experimental results demonstrated its accuracy as reconstuction errors are less than one degree in rotation and less than one pixel in translation.
Stelios Krinidis, Christophoros Nikou, Ioannis Pitas
IEEE Trans. Inf. Technol. Biomed.3
2002 A SOM Variant Based on the Wilcoxon Test for Document Organization and Retrieval
Apostolos Georgakis, Constantine Kotropoulos, Ioannis Pitas
ICANN3
2002 Recent advances in biometric person authentication
abstract
Biometrics is an emerging topic in the field of signal processing. While technologies (e.g. audio, video) for biometrics have mostly been studied separately, ultimately, biometric technologies could find their strongest role as interwined and complementary pieces of a multi-modal authentication system. In this paper, a short overview of voice, fingerprint, and face authentication algorithms is provided.
Jean-Luc Dugelay, Jean-Claude Junqua, Constantine Kotropoulos, Roland Kuhn 0001, Florent Perronnin, Ioannis Pitas
ICASSP6
2002 3D image watermarking robust to geometric distortions
abstract
A novel blind method for 3D image watermarking robust against geometric distortions is proposed. A ternary watermark is embedded in a grayscale or a color 3D volume. Construction of watermarks having appropriate structure enables fast and robust watermark detection even after several geometric distortions of the watermarked volume. Simulation results indicate the ability of the proposed method to deal with the aforementioned attacks. The proposed method is also robust against lossy compression up to a certain compression ratio. Experiments conducted indicate the superiority of the proposed method.
Anastasios Tefas, Giorgos Louizis, Ioannis Pitas
ICASSP3
2002 On the stability of support vector machines for face detection
abstract
In this paper we study the stability of support vector machines in face detection by decomposing their average prediction error into the bias, variance, and aggregation effect terms. Such an analysis indicates whether bagging, a method for generating multiple versions of a classifier from bootstrap samples of a training set, and combining their outcomes by majority voting, is expected to improve the accuracy of the classifier. We estimate the bias, variance, and aggregation effect by using bootstrap smoothing techniques when support vector machines are applied to face detection in the AT & T face database and we demonstrate that support vector machines are stable classifiers. Accordingly, bagging is not expected to improve their face detection accuracy.
Ioan Buciu, Constantine Kotropoulos, Ioannis Pitas
ICIP (3)3
2002 Shot detection in video sequences using entropy based metrics
abstract
A new method for detecting shot boundaries in video sequences using metrics based on information theory is proposed. The method relies on the mutual information and the joint entropy between frames and can detect cuts, fade-ins and fade-outs. The detection technique was tested on TV video sequences having different types of shots and significant object and camera motion inside the shots. It was favorably compared to other recently proposed shot cut detection techniques. The method is proven to detect both fades and abrupt cuts very effectively.
Zuzana Cernekova, Christophoros Nikou, Ioannis Pitas
ICIP (3)3
2002 Application of support vector machines classifiers to visual speech recognition
abstract
In this paper we propose a visual speech recognition network based on support vector machines. Each word of the dictionary is modeled by a set of temporal sequences of visemes. Each viseme is described by a support vector machine, and the temporal character of speech is modeled by integrating the support vector machines as nodes into a Viterbi decoding lattice. Experiments conducted on a small visual speech recognition task using very simple features demonstrate a word recognition rate on the level of the best rates previously reported even without training the state transition probabilities in the Viterbi lattices. This proves the suitability of support vector machines for visual speech recognition.
Mihaela Gordan, Constantine Kotropoulos, Apostolos Georgakis, Ioannis Pitas
ICIP (3)4
2002 An information theoretic approach to joint probabilistic face detection and tracking
abstract
A joint probabilistic face detection and tracking algorithm for combining a likelihood estimation and a prior probability is proposed. Face tracking is achieved by a Bayesian framework. The likelihood estimation scheme is based on statistical training of sets of automatically generated feature points, while the prior probability estimation is based on the fusion of an information theoretic tracking cue and a Gaussian temporal model. The likelihood estimation process is the cone of a multiple face detection scheme used to initialize the tracking process. The resulting system was tested on real image sequences and is robust to significant partial occlusion and illumination changes.
Evangelos Loutas, Ioannis Pitas, Christophoros Nikou
ICIP (1)2
2002 Information theory-based analysis of partial and total occlusion in object tracking
abstract
Metrics based on mutual information and not resorting to ground truth data are proposed in this paper in order to measure tracking reliability under occlusion., The variations of the proposed metrics can be used as a quantitative estimate of changes in the tracking region, caused by occlusion, sudden movement or the deformation of the tracked object. The proposed metric was tested on an object tracking scheme using multiple feature point correspondences. Experimental results have shown that mutual information can effectively characterize object appearance and reappearance in many computer vision applications.
Evangelos Loutas, Ioannis Pitas, Christophoros Nikou
ICIP (2)2
2002 Optimal detector structure for DCT and subband domain watermarking
abstract
Most of the watermarking schemes that have been proposed until now employ a correlator in the detection stage. The current paper proposes a new detector scheme that can be applied in the case of additive watermarking in the DCT or DWT domain. Certain properties of the probability density function of the coefficients in these domains are exploited in order to construct an asymptotically optimal detector based on well known results of the detection theory. Detection is performed without the use of the original image, as in methods employing different detectors. Experimental results prove the superiority of the proposed detector over the correlator.
Athanasios Nikolaidis, Ioannis Pitas
ICIP (3)2
2002 Watermarking of sets of polygonal lines using fusion techniques
abstract
A blind watermarking method for the copyright protection of sets of polygonal lines in vector graphics images and GIS data (elevation contour maps) is presented. The paper focuses mainly on the use of simple fusion rules for combining the detector outputs from each polygonal line in order to come up with a global detection result. Experimental comparison of the various fusion methods using both synthetic and real data (elevation maps) is provided.
Alexia Giannoula, Nikos Nikolaidis 0001, Ioannis Pitas
ICME (2)3
2002 3D physics-based reconstruction of serially acquired slices
abstract
This paper presents an accurate, computationally efficient, fast and fully-automated algorithm for the alignment of 2D serially acquired sections forming a 3D volume. The method accounts for the main shortcomings of 3D image alignment: corrupted data (cuts and tears), dissimilarities or discontinuities between slices and missing slices. The approach relies on the determination of inter-slice correspondences. The features used for correspondence are extracted by a 2D physics-based deformable model parameterizing the object shape. Correspondence affinities and global constraints render the method efficient and reliable. The method has been evaluated on real images and the experimental results demonstrate its accuracy, as reconstruction errors are smaller than 1 degree in rotation and smaller than 1 pixel in translation.
Stelios Krinidis, Christophoros Nikou, Ioannis Pitas
ICME (1)3
2002 Copyright protection of 3D images using watermarks of specific spatial structure
abstract
A novel blind method for 3D image watermarking, robust against geometric distortions, is proposed. A ternary watermark is embedded in a grayscale or a color 3D volume. Construction of watermarks having appropriate structure enables fast and robust watermark detection even after several geometric distortions of the watermarked volume. Simulation results indicate the ability of the proposed method to deal with the aforementioned attacks. The proposed method is also robust against lossy compression up to a certain compression ratio.
Giorgos Louizis, Anastasios Tefas, Ioannis Pitas
ICME (2)3
2002 Watermark detection: benchmarking perspectives
abstract
Benchmarking of watermarking algorithms is a complicated task that requires examination of a set of mutually dependent performance factors (algorithm complexity, decoding/detection performance, and perceptual quality). This paper will focus on detection/decoding performance evaluation and try to summarize its basic principles. A methodology for deriving the corresponding performance metrics will also be provided.
Nikos Nikolaidis 0001, Vassilios Solachidis, Anastasios Tefas, Ioannis Pitas
ICME (2)4
2002 Face verification using elastic graph matching based on morphological signal decomposition
Anastasios Tefas, Constantine Kotropoulos, Ioannis Pitas
Signal Process.3
2002 Binary Morphological Shape-Based Interpolation Applied to 3D Tooth Reconstruction
abstract
In this paper, we propose an interpolation algorithm using a mathematical morphology morphing approach. The aim of this algorithm is to reconstruct the n-dimensional object from a group of (n - 1)-dimensional sets representing sections of that object. The morphing transformation modifies pairs of consecutive sets such that they approach in shape and size. The interpolated set is achieved when the two consecutive sets are made idempotent by the morphing transformation. We prove the convergence of the morphological morphing. The entire object is modeled by successively interpolating a certain number of intermediary sets between each two consecutive given sets. We apply the interpolation algorithm for three-dimensional tooth reconstruction.
Adrian G. Bors, Lefteris Kechagias, Ioannis Pitas
IEEE Trans. Medical Imaging3
2001 Robust spatial image watermarking using progressive detection
abstract
A novel method for image watermarking robust to geometric distortions is proposed. A binary watermark is embedded in a grayscale or a color host image. The ability of progressive watermark detection enables fast and robust watermark detection even after several geometric distortions of the watermarked image. Simulation results indicate the ability of the proposed method to deal with the aforementioned attacks. Experiments conducted using the Stirmark benchmarking tests, indicate the superiority of the proposed method.
Anastasios Tefas, Ioannis Pitas
ICASSP2
2001 Bernoulli shift generated chaotic watermarks: theoretic investigation
abstract
The paper statistically analyzes the behaviour of chaotic watermark signals generated by n-way Bernoulli shift maps. For this purpose, a simple blind copyright protection watermarking system is considered. The analysis involves theoretical evaluation of the system detection reliability, when a correlator detector is used. The aim of the paper is twofold: (i) to introduce the n-way Bernoulli shift generated chaotic watermarks and theoretically contemplate their properties with respect to detection reliability and (ii) to establish theoretically their potential superiority against the widely used pseudorandom watermarks. Experimental verification of the theoretical analysis results is also performed.
Sofia Tsekeridou, Vassilios Solachidis, Nikos Nikolaidis 0001, Athanasios Nikolaidis, Anastasios Tefas, Ioannis Pitas
ICASSP6
2001 Frontal face detection using support vector machines and back-propagation neural networks
abstract
Face detection is a key problem in building systems that perform face recognition/verification and model-based image coding. Two algorithms for face detection that employ either support vector machines or backpropagation feedforward neural networks are described, and their performance is tested on the same frontal face database using the false acceptance and false rejection rates as quantitative figures of merit. The aforementioned algorithms can replace the explicitly-defined knowledge for facial regions and facial features in mosaic-based face detection algorithms.
Nikoletta Bassiou, Constantine Kotropoulos, T. Kosmidis, Ioannis Pitas
ICIP (1)4
2001 Shape-based interpolation using morphological morphing
abstract
We propose an interpolation algorithm using a morphological morphing approach. The aim of this algorithm is to reconstruct an n-dimensional object from a group of (n-1)-dimensional sets representing object sections. The morphing transformation modifies consecutive sets so that they approach in shape and size. When the two morphed sets become idempotent we generate a new set. The entire object is modeled by successively interpolating a certain number of intermediary sets between each two consecutive initial sets. The interpolation algorithm is used for 3D tooth reconstruction.
Adrian G. Bors, Lefteris Kechagias, Ioannis Pitas
ICIP (2)3
2001 Combining support vector machines for accurate face detection
abstract
The paper proposes the application of majority voting on the output of several support vector machines in order to select the most suitable learning machine for frontal face detection. The first experimental results indicate a significant reduction of the rate of false positive patterns.
Ioan Buciu, Constantine Kotropoulos, Ioannis Pitas
ICIP (1)3
2001 Occlusion resistant object tracking
abstract
Object tracking with occlusion prediction using multiple feature correspondences is proposed. The tracking region is defined by a set of point features, tracked using Kanade-Lucas-Tomasi (1991) algorithm. During total occlusion the region position is estimated using motion prediction based on a Kalman filtering scheme applied to the motion model prior to occlusion. During partial occlusion the displacements of the occluded features are predicted based on the motion of the bounding box of the moving object. Experimental results on real and artificial images have shown that the algorithm behaves well under total and partial occlusion.
Evangelos Loutas, Ioannis Pitas, Konstantinos I. Diamantaras
ICIP (2)2
2001 Digital image processing in painting restoration and archiving
abstract
Digital image processing and analysis can be an important tool for the restoration of works of art. This paper presents three applications of image processing in this field: a method for digital crack restoration of paintings, a technique for color restoration of old paintings and a method for mosaicing of partial images of works of art painted on curved surfaces. A digital archiving system for works of arts is also described.
Nikos Nikolaidis 0001, Ioannis Pitas
ICIP (1)2
2001 A benchmarking protocol for watermarking methods
abstract
A benchmarking system for watermarking algorithms is described. The proposed benchmarking system can be used to evaluate the performance of watermarking methods used for copyright protection, authentication, fingerprinting, etc. Although the system described is used for image watermarking, the general framework can be used, by introducing a different set of attacks, for benchmarking of video and audio data.
Nikos Nikolaidis 0001, Sofia Tsekeridou, Anastasios Tefas, Vassilios Solachidis, Athanasios Nikolaidis, Ioannis Pitas
ICIP (3)6
2001 Digital watermarking in the fractional Fourier transformation domain
Igor Djurovic, Srdjan Stankovic, Ioannis Pitas
J. Netw. Comput. Appl.3
2001 Using Support Vector Machines to Enhance the Performance of Elastic Graph Matching for Frontal Face Authentication
abstract
A novel method for enhancing the performance of elastic graph matching in frontal face authentication is proposed. The starting point is to weigh the local similarity values at the nodes of an elastic graph according to their discriminatory power. Powerful and well-established optimization techniques are used to derive the weights of the linear combination. More specifically, we propose a novel approach that reformulates Fisher's discriminant ratio to a quadratic optimization problem subject to a set of inequality constraints by combining statistical pattern recognition and support vector machines (SVM). Both linear and nonlinear SVM are then constructed to yield the optimal separating hyperplanes and the optimal polynomial decision surfaces, respectively. The method has been applied to frontal face authentication on the M2VTS database. Experimental results indicate that the performance of morphological elastic graph matching is highly improved by using the proposed weighting technique.
Anastasios Tefas, Constantine Kotropoulos, Ioannis Pitas
IEEE Trans. Pattern Anal. Mach. Intell.3
2001 Image authentication techniques for surveillance applications
abstract
In automatic video surveillance (VS) systems, the issue of authenticating the video content is of primary importance. Given the ease with which digital images and videos can be manipulated, practically they do not have any value as legal proof, if the possibility of authenticating their content is not provided. In this paper, the problem of authenticating video surveillance image sequences is considered. After an introduction motivating the need for a watermarking-based authentication of VS sequences, a brief survey of the main watermarking-based authentication techniques is presented and the requirements that an authentication algorithm should satisfy for VS applications, are discussed. A novel algorithm which is suitable for VS visual data authentication is also presented and the results obtained by applying it to test data are discussed.
Franco Bartolini, Anastasios Tefas, Mauro Barni, Ioannis Pitas
Proc. IEEE4
2001 Projection distortion analysis for flattened image mosaicing from straight uniform generalized cylinders
William Puech, Adrian G. Bors, Ioannis Pitas, Jean-Marc Chassery
Pattern Recognit.3
2001 Statistical analysis of a watermarking system based on Bernoulli chaotic sequences
Sofia Tsekeridou, Vassilios Solachidis, Nikos Nikolaidis 0001, Athanasios Nikolaidis, Anastasios Tefas, Ioannis Pitas
Signal Process.6
2001 Content-based video parsing and indexing based on audio-visual interaction
abstract
A content-based video parsing and indexing method is presented in this paper, which analyzes both information sources (auditory and visual) and accounts for their inter-relations and synergy to extract high-level semantic information. Both frame- and object-based access to the visual information is employed. The aim of the method is to extract semantically meaningful video scenes and assign semantic label(s) to them. Due to the temporal nature of video, time has to be accounted for. Thus, time-constrained video representations and indices are generated. The current approach searches for specific types of content information relevant to the presence or absence of speakers or persons. Audio-source parsing and indexing leads to the extraction of a speaker label mapping of the source over time. Video-source parsing and indexing results in the extraction of a talking-face shot mapping over time. Integration of the audio and visual mappings constrained by interaction rules leads to higher levels of video abstraction and even partial detection of its context.
Sofia Tsekeridou, Ioannis Pitas
IEEE Trans. Circuits Syst. Video Technol.2
2001 Region-based image watermarking
abstract
We introduce a novel method for embedding and detecting a chaotic watermark in the digital spatial image domain, based on segmenting the image and locating regions that are robust to several image manipulations. The robustness of the method is confirmed by experimental results that display the immunity of the embedded watermark to several kinds of attacks, such as compression, filtering, scaling, cropping, and rotation.
Athanasios Nikolaidis, Ioannis Pitas
IEEE Trans. Image Process.2
2001 Circularly symmetric watermark embedding in 2-D DFT domain
abstract
In this paper, a method for digital image watermarking is described that is resistant to geometric transformations. A private key, which allows a very large number of watermarks, determines the watermark, which is embedded on a ring in the DFT domain. The watermark possesses circular symmetry. Correlation is used for watermark detection. The original image is not required in detection. The proposed method is resistant to JPEG compression, filtering, noise addition, scaling, translation, cropping, rotation, printing and rescanning. Experimental results prove the robustness of this method against the aforementioned attacks.
Vassilios Solachidis, Ioannis Pitas
IEEE Trans. Image Process.2
2001 Watermarking in the space/spatial-frequency domain using two-dimensional Radon-Wigner distribution
abstract
A two-dimensional (2-D) signal with a variable spatial frequency is proposed as a watermark in the spatial domain. This watermark is characterized by a linear frequency change. It can be efficiently detected by using 2-D space/spatial-frequency distributions. The projections of the 2-D Wigner distribution--the 2-D Radon-Wigner distribution, are used in order to emphasize the watermark detection process. The watermark robustness with respect to some very important image processing attacks, such as for example, the translation, rotation, cropping, JPEG compression, and filtering, is demonstrated and tested by using Stirmark 3.1.
Srdjan Stankovic, Igor Djurovic, Ioannis Pitas
IEEE Trans. Image Process.3
2001 Robust audio watermarking in the time domain
abstract
The audio watermarking method proposed in this paper offers copyright protection to an audio signal by time domain processing. The strength of audio signal modifications is limited by the necessity to produce an output signal that is perceptually similar to the original one. The watermarking method presented here does not require the use of the original signal for watermark detection. The watermark signal is generated using a key, i.e., a single number known only to the copyright owner. Watermark embedding depends on the audio signal amplitude and frequency in a way that minimizes the audibility of the watermark signal. The embedded watermark is robust to common audio signal manipulations like MPEG audio coding, cropping, time shifting, filtering, resampling, and requantization.
P. Bassia, Ioannis Pitas, Nikos Nikolaidis 0001
IEEE Trans. Multim.2
2000 Watermarking polygonal lines using Fourier descriptors
abstract
A method for watermarking of polygonal lines is proposed. The watermark is embedded in the Fourier descriptors of the polygonal line causing minor distortions to the coordinates of the vertices of the polygonal line. Watermarks generated by this technique can be successfully detected even after rotation, translation, scaling and reflection of the host polygonal line.
Vassilios Solachidis, Nikos Nikolaidis 0001, Ioannis Pitas
ICASSP3
2000 Face authentication by using elastic graph matching and support vector machines
abstract
A novel method for enhancing the performance of elastic graph matching in face authentication is proposed. The starting point is to weigh the local matching errors at the nodes of an elastic graph according to their discriminatory power. We propose a novel approach to discriminant analysis that re-formulates Fisher's linear discriminant ratio to a quadratic optimization problem subject to inequality constraints by combining statistical pattern recognition and support vector machines. The method is applied to frontal face authentication on the M2VTS database.
Anastasios Tefas, Constantine Kotropoulos, Ioannis Pitas
ICASSP3
2000 Embedding self-similar watermarks in the wavelet domain
abstract
A wavelet-based watermarking method for still images is presented. Watermarks are composed of scaled versions of circularly-symmetric pseudo-random 2D signals in order to attain a spatial self-similar structure with respect to a cartesian grid. Embedding and detection are performed in the wavelet domain thus allowing multi-level detection. The novelty of the method lies in the use of self-similar watermarks (quasi scale-invariant), expected to be robust against geometric transformations, especially scaling. The approach proves rather efficient against many kinds of distortions, such as linear and nonlinear filtering, JPEG and wavelet compression, scaling, cropping and rotation.
Sofia Tsekeridou, Ioannis Pitas
ICASSP2
2000 Fourier Descriptors Watermarking of Vector Graphics Images
abstract
A blind method for watermarking of vector graphics images through polygonal line modification is proposed. The watermark is embedded in the Fourier descriptors of the polygonal lines, causing invisible distortions to the vertices coordinates. The properties of the Fourier descriptors ensure that watermarks generated by this technique withstand rotation, translation, scaling, reflection, change of traversal starting point/direction and smoothing.
Vassilios Solachidis, Nikos Nikolaidis 0001, Ioannis Pitas
ICIP3
2000 Using Support Vector Machines for Face Authentication Based on Elastic Graph Matching
abstract
A novel method for enhancing the performance of elastic graph matching in face authentication is proposed. Our objective is to weigh the local matching errors at the nodes of an elastic graph according to their discriminatory power. We propose a novel approach to discriminant analysis that re-formulates Fisher's linear discriminant ratio to a quadratic optimization problem subject to inequality constraints by combining statistical pattern recognition and support vector machines. The method is applied to frontal face authentication on the M2VTS database.
Anastasios Tefas, Constantine Kotropoulos, Ioannis Pitas
ICIP3
2000 Multi-Bit Image Watermarking Robust to Geometric Distortions
abstract
A novel method for multi-bit image watermarking robust to geometric distortions is proposed. A binary watermark is embedded in a grayscale or a color host image. The ability of progressive watermark detection enables fast and robust watermark detection even after several geometric distortions of the watermarked image. The embedding of a multi-bit message and its robust decoding after the watermark detection is also discussed. Simulation results indicate the ability of the proposed method to deal with the aforementioned attacks.
Anastasios Tefas, Ioannis Pitas
ICIP2
2000 Copyright Protection of Still Images Using Self-Similar Chaotic Watermarks
abstract
A spatial domain watermarking algorithm for still images is proposed in this paper. The watermarks are generated by combining scaled versions of 2-D chaotic signals, in order to attain a self-similar structure. In addition, the 2-D chaotic signal generation procedure contains operations which ensure that lowpass watermarks are obtained. Such characteristics (lowpass spectrum, self-similarity) enable robustness against compression, lowpass filtering or, generally, distortions of lowpass nature, as well as, cropping and scaling. An additional advantage of the proposed algorithm is that the detection procedure does not require the original image.
Sofia Tsekeridou, Nikos Nikolaidis 0001, Nicholas D. Sidiropoulos, Ioannis Pitas
ICIP4
2000 Comparison of Face Verification Results on the XM2VTS Database
abstract
Presents results of the face verification contest that was organized in conjunction with International Conference on Pattern Recognition 2000. Participants had to use identical data sets from a large, publicly available multimodal database XM2VTSDB. Training and evaluation was carried out according to an a priori known protocol. Verification results of all tested algorithms have been collected and made public on the XM2VTSDB website, facilitating large scale experiments on classifier combination and fusion. Tested methods included, among others, representatives of the most common approaches to face verification -elastic graph matching, Fisher's linear discriminant and support vector machines.
Jiri Matas, Miroslav Hamouz, Kenneth Jonsson, Josef Kittler, Yongping Li, Constantine Kotropoulos, Anastasios Tefas, Ioannis Pitas, Teewoon Tan, Hong Yan 0001, Fabrizio Smeraldi, N. Capdevielle, Wulfram Gerstner, Yousri Abdeljaoued, Josef Bigün, Souheil Ben Yacoub, Eddy Mayoraz
ICPR8
2000 Morphological Segmentation of Histology Cell Images
abstract
Two algorithms for segmentation of cell images are proposed. They have a unique part that contains computation of morphological gradient to extract object borders and thinning the obtained borders to get a line of one-pixel thickness. For this task, we propose the fast gray-scale thinning algorithm that is based on the idea of the analysis of binary image layers. Then, the obtained one-pixel lines are used to extract cells and compute their characteristics. The algorithms based on morphological and split/merge segmentation are developed and used for this task.
Alexandr Nedzved, Sergey Ablameyko 0001, Ioannis Pitas
ICPR3
2000 Comparison of different chaotic maps with application to image watermarking
abstract
The current paper presents a method for image watermarking based on chaotic sequences produced by different functions. The effect of embedding a chaotic watermark on an image is examined. Extensive experimental results for JPEG and filtering attacks on images watermarked by the different functions are presented together with a comparative study of them. The results show that efficient watermarking methods based on chaos can be designed under certain conditions.
Athanasios Nikolaidis, Ioannis Pitas
ISCAS2
2000 Image authentication using chaotic mixing systems
abstract
A novel method for image authentication is proposed. A watermark signal is embedded in a grayscale or a color host image. The watermark key controls a set of parameters of a chaotic system used for the watermark generation. The use of chaotic mixing increases the security of the proposed method and provides the additional feature of imperceptible encryption of the image owner logo in the host image. The method succeeds in detecting any alteration made in a watermarked image. The proposed method is robust in high quality lossy image compression. It provides the user not only with a measure for the authenticity of the test image but also with an image map that highlights the unaltered image regions when selective tampering has been made.
Anastasios Tefas, Ioannis Pitas
ISCAS2
2000 Wavelet-based self-similar watermarking for still images
abstract
This paper presents a wavelet-based watermarking method for copyright protection of still images. Watermarks are structured in such a way to attain spatial self-similarity with respect to a cartesian grid. Embedding and detection are performed in the wavelet domain thus allowing multiresolution detection. The novelty of the current approach is the use of self-similar watermarks (quasi scale-invariant), which are expected to be robust against geometric transformations, especially scaling. The approach proves rather efficient against many kinds of distortions, such as linear and nonlinear filtering, JPEG compression, scaling, cropping and rotation.
Sofia Tsekeridou, Ioannis Pitas
ISCAS2
2000 Morphological elastic graph matching applied to frontal face authentication under well-controlled and real conditions
Constantine Kotropoulos, Anastasios Tefas, Ioannis Pitas
Pattern Recognit.3
2000 Facial feature extraction and pose determination
Athanasios Nikolaidis, Ioannis Pitas
Pattern Recognit.2
2000 Analog implementation of erosion/dilation, median and order statistics filters
Spiridon Vlassis, Kostantinos Doris, Stilianos Siskos, Ioannis Pitas
Pattern Recognit.4
2000 MPEG-2 error concealment based on block-matching principles
abstract
The MPEG-2 compression algorithm is very sensitive to channel disturbances due to the use of variable-length coding. A single bit error during transmission leads to noticeable degradation of the decoded sequence quality, in that part or an entire slice information is lost until the next resynchronization point is reached. Error concealment (EC) methods, implemented at the decoder side, present one way of dealing with this problem. An error-concealment scheme that is based on block-matching principles and spatio-temporal video redundancy is presented in this paper. Spatial information (for the first frame of the sequence or the next scene) or temporal information (for the other frames) is used to reconstruct the corrupted regions. The concealment strategy is embedded in the MPEG-2 decoder model in such a way that error concealment is applied after entire frame decoding. Its performance proves to be satisfactory for packet error rates (PER) ranging from 1% to 10% and for video sequences with different content and motion and surpasses that of other EC methods under study.
Sofia Tsekeridou, Ioannis Pitas
IEEE Trans. Circuits Syst. Video Technol.2
2000 Prediction and tracking of moving objects in image sequences
abstract
We employ a prediction model for moving object velocity and location estimation derived from Bayesian theory. The optical flow of a certain moving object depends on the history of its previous values. A joint optical flow estimation and moving object segmentation algorithm is used for the initialization of the tracking algorithm. The segmentation of the moving objects is determined by appropriately classifying the unlabeled and the occluding regions. Segmentation and optical flow tracking is used for predicting future frames.
Adrian G. Bors, Ioannis Pitas
IEEE Trans. Image Process.2
2000 A generalized fuzzy mathematical morphology and its application in robust 2-D and 3-D object representation
abstract
In this paper, the generalized fuzzy mathematical morphology (GFMM) is proposed, based on a novel definition of the fuzzy inclusion indicator (FII). FII is a fuzzy set used as a measure of the inclusion of a fuzzy set into another, that is proposed to be a fuzzy set. It is proven that the FII obeys a set of axioms, which are proposed to be extensions of the known axioms that any inclusion indicator should obey, and which correspond to the desirable properties of any mathematical morphology operation. The GFMM provides a very powerful and flexible tool for morphological operations. The binary and grayscale mathematical morphologies can be considered as special cases of the proposed GFMM. An application for robust skeletonization and shape decomposition of two-dimensional (2-D) and three-dimensional (3-D) objects is presented. Simulation examples show that the object reconstruction from their skeletal subsets that can be achieved by using the GFMM is better than by using the binary mathematical morphology in most cases. Furthermore, the use of the GFMM for skeletonization and shape decomposition preserves the shape and the location of the skeletal subsets and spines.
Vassilios Chatzis, Ioannis Pitas
IEEE Trans. Image Process.2
2000 Frontal face authentication using morphological elastic graph matching
abstract
A novel dynamic link architecture based on multiscale morphological dilation-erosion is proposed for frontal face authentication. Instead of a set of Gabor filters tuned to different orientations and scales, multiscale morphological operations are employed to yield a feature vector at each node of the reference grid. Linear projection algorithms for feature selection and automatic weighting of the nodes according to their discriminatory power succeed to increase the authentication capability of the method. The performance of the morphological dynamic link architecture is evaluated in terms of the receiver operating characteristic in the M2VTS face image database. The comparison with other frontal face authentication algorithms indicates that the morphological dynamic link architecture with discriminatory power coefficients is the best algorithm with respect to the equal error rate achieved.
Constantine Kotropoulos, Anastasios Tefas, Ioannis Pitas
IEEE Trans. Image Process.3
2000 A fast implementation of 3-D binary morphological transformations
abstract
This paper proposes a fast algorithm for implementing the basic operation of Minkowski addition for the special case of binary three-dimensional (3-D) images, using 3-D structuring elements of arbitrary size and shape. The application of the proposed algorithm for all the other morphological transformations is straightforward, as they can all be expressed in terms of Minkowski addition. The efficiency of the algorithm is analyzed and some experimental results of its application are presented. As shown, the efficiency of the algorithm increases with the size of the structuring element.
Nikos Nikopoulos, Ioannis Pitas
IEEE Trans. Image Process.2
2000 Digital color restoration of old paintings
abstract
Physical and chemical changes can degrade the visual color appearance of old paintings. Five digital color restoration techniques, which can be used to simulate the original appearance of paintings, are presented. Although a small number of color samples is employed in the restoration procedure, simulation results indicate that good restoration quality can be attained.
Michail Pappas, Ioannis Pitas
IEEE Trans. Image Process.2
2000 Memory efficient propagation-based watershed and influence zone algorithms for large images
abstract
Propagation front or grassfire methods are very popular in image processing because of their efficiency and because of their inherent geodesic nature. However, because of their random-access nature, they are inefficient in large images that cannot fit in available random access memory. We explore ways to increase the memory efficiency of two algorithms that use propagation fronts: the skeletonization by influence zones and the watershed transform. Two algorithms are presented for the skeletonization by influence zones. The first computes the skeletonization on surfaces without storing the enclosing volume. The second performs the skeletonization without any region reference, by using only the propagation fronts. The watershed transform algorithm that was developed keeps in memory the propagation fronts and only one greylevel of the image. All three algorithms use much less memory than the ones presented in the literature so far. Several techniques have been developed in this work in order to minimize the effect of these set operations. These include fast search methods, double propagation fronts, directional propagation, and others.
Ioannis Pitas, Costas I. Cotsaces
IEEE Trans. Image Process.1
2000 Interpolation of 3-D Binary Images based on Morphological Skeletonization
abstract
In this paper, the morphological skeleton interpolation (MSI) algorithm is presented. It is an efficient, shape-based interpolation method used for interpolating slices in a three-dimensional (3-D) binary object. It is based on morphological skeletonization, which is used for two-dimensional (2-D) slice representation. The proposed morphological skeleton matching process provides translation, rotation, and scaling information at the same time. The interpolated slices preserve the shape of the original object slices, when the slices have similar shapes. It can also modify the shape of an object when the successive slices do not have similar shapes. Applications on artificial and real data are also presented.
Vassilios Chatzis, Ioannis Pitas
IEEE Trans. Medical Imaging2
2000 Frontal Face Authentication Using Discriminating Grids with Morphological Feature Vectors
abstract
A novel elastic graph matching procedure based on multiscale morphological operations, the so called morphological dynamic link architecture, is developed for frontal face authentication. Fast algorithms for implementing mathematical morphology operations are presented. Feature selection by employing linear projection algorithms is proposed. Discriminatory power coefficients that weigh the matching error at each grid node are derived. The performance of morphological dynamic link architecture in frontal face authentication is evaluated in terms of the receiver operating characteristic on the M2VTS face image database. Preliminary results for face recognition using the proposed technique are also presented.
Constantine Kotropoulos, Anastasios Tefas, Ioannis Pitas
IEEE Trans. Multim.3
2000 Robust Watermarking of Facial Images Based on Salient Geometric Pattern Matching
abstract
We introduce a novel method for embedding and detecting a chaotic watermark in the digital spatial domain of color facial images, based on localizing salient facial features. These features define a certain area on which the watermark is embedded and detected. An assessment of the watermarking robustness is done experimentally, by testing resistance to several attacks, such as compression, filtering, noise addition, scaling, cropping and rotation.
Athanasios Nikolaidis, Ioannis Pitas
IEEE Trans. Multim.2
1999 Circularly symmetric watermark embedding in 2-D DFT domain
abstract
This paper presents an algorithm for rotation and scale invariant watermarking of digital images. An invisible mark is embedded in magnitude of the DFT domain. It is robust to compression, filtering, cropping, translation and rotation. The watermark introduces image changes that are invisible to the human eye. The detection algorithm does not require the original image.
Vassilios Solachidis, Ioannis Pitas
ICASSP2
1999 Compensating for variable recording conditions in frontal face authentication algorithms
abstract
This paper addresses the problem of compensating for variable recording conditions such as changes in illumination, scale differences, and varying face position. It is well known that the performance of any face authentication/recognition algorithm deteriorates significantly in the presence of the aforementioned conditions as well as the expression variations. The use of simple and powerful pre-processing techniques aiming at compensating for variable recording conditions prior to the application of any authentication algorithm is proposed. It is shown that such an approach overcomes indeed the image variations and guarantees an almost stable performance for the Morphological Dynamic Link Architecture developed within the European research project M2VTS.
Anastasios Tefas, Yann Menguy, Constantine Kotropoulos, Gaël Richard, Ioannis Pitas, Philip Lockwood
ICASSP5
1999 Motion field estimation by vector rational interpolation for error concealment purposes
abstract
A study on the use of vector rational interpolation for the estimation of erroneously received motion fields of an MPEG-2 coded video bitstream has been performed. Four different motion vector interpolation schemes have been examined using motion information from available top and bottom adjacent blocks since left or right neighbours are usually lost. The presented interpolation schemes are capable of adapting their behaviour according to neighbouring motion information. Simulation results prove the satisfactory performance of the novel nonlinear interpolation schemes and the success of their application to the concealment of predictively coded frames. The motion vector rational interpolation concealment method proves to be a fast method, thus adequate for real-time applications.
Sofia Tsekeridou, Faouzi Alaya Cheikh, Moncef Gabbouj, Ioannis Pitas
ICASSP4
1999 The use of watermarks in the protection of digital multimedia products
abstract
The watermarking of digital images, audio, video, and multimedia products in general has been proposed for resolving copyright ownership and verifying originality of content. This paper studies the contribution of watermarking for developing protection schemes. A general watermarking framework (GWF) is studied and the fundamental demands are listed. The watermarking algorithms, namely watermark generation, embedding, and detection, are analyzed and necessary conditions for a reliable and efficient protection are stated. Although the GWF satisfies the majority of requirements for copyright protection and content verification, there are unsolved problems inside a pure watermarking framework. Particular solutions, based on product registration and related network services, are suggested to overcome such problems.
George Voyatzis, Ioannis Pitas
Proc. IEEE2
1999 Object classification in 3-D images using alpha-trimmed mean radial basis function network
abstract
We propose a pattern classification based approach for simultaneous three-dimensional (3-D) object modeling and segmentation in image volumes. The 3-D objects are described as a set of overlapping ellipsoids. The segmentation relies on the geometrical model and graylevel statistics. The characteristic parameters of the ellipsoids and of the graylevel statistics are embedded in a radial basis function (RBF) network and they are found by means of unsupervised training. A new robust training algorithm for RBF networks based on alpha-trimmed mean statistics is employed in this study. The extension of the Hough transform algorithm in the 3-D space by employing a spherical coordinate system is used for ellipsoidal center estimation. We study the performance of the proposed algorithm and we present results when segmenting a stack of microscopy images.
Adrian G. Bors, Ioannis Pitas
IEEE Trans. Image Process.2
1999 Fuzzy scalar and vector median filters based on fuzzy distances
abstract
In this paper, the fuzzy scalar median (FSM) is proposed, defined by using ordering of fuzzy numbers based on fuzzy minimum and maximum operations defined by using the extension principle. Alternatively, the FSM is defined from the minimization of a fuzzy distance measure, and the equivalence of the two definitions is proven. Then, the fuzzy vector median (FVM) is proposed as an extension of vector median, based on a novel distance definition of fuzzy vectors, which satisfy the property of angle decomposition. By defining properly the fuzziness of a value, the combination of the basic properties of the classical scalar and vector median (VM) filter with other desirable characteristics can be succeeded.
Vassilios Chatzis, Ioannis Pitas
IEEE Trans. Image Process.2
1999 Morphological iterative closest point algorithm
abstract
This work presents a method for the registration of three-dimensional (3-D) shapes. The method is based on the iterative closest point (ICP) algorithm and improves it through the use of a 3-D volume containing the shapes to be registered. The Voronoi diagram of the "model" shape points is first constructed in the volume. Then this is used for the calculation of the closest point operator. This way a dramatic decrease of the computational cost is achieved.
Christos A. Kapoutsis, C. P. Vavoulidis, Ioannis Pitas
IEEE Trans. Image Process.3
1999 Multimodal decision-level fusion for person authentication
abstract
The use of clustering algorithms for decision-level data fusion is proposed. Person authentication results coming from several modalities (e.g., still image, speech), are combined by using fuzzy k-means (FKM) and fuzzy vector quantization (FVQ) algorithms, and a median radial basis function (MRBF) network. The quality measure of the modalities data is used for fuzzification. Two modifications of the FKM and FVQ algorithms, based on a fuzzy vector distance definition, are proposed to handle the fuzzy data and utilize the quality measure. Simulations show that fuzzy clustering algorithms have better performance compared to the classical clustering algorithms and other known fusion algorithms. MRBF has better performance especially when two modalities are combined. Moreover, the use of the quality via the proposed modified algorithms increases the performance of the fusion system.
Vassilios Chatzis, Adrian G. Bors, Ioannis Pitas
IEEE Trans. Syst. Man Cybern. Part A3
1998 Variants of Dynamic Link Architecture Based on Mathematical Morphology for Frontal Face Authentication
abstract
Two novel variants of dynamic link architecture that are based on mathematical morphology and incorporate coefficients which weigh the contribution of each node in elastic graph matching according to its discriminatory power are developed. They are the so called Morphological Dynamic Link Architecture and the Morphological Signal Decomposition-Dynamic Lint Architecture. The proposed variants are tested for face authentication in a cooperative scenario where the candidates claim an identity to be checked. Their performance is evaluated in terms of their receiver operating characteristic and the equal error rate achieved in M2VTS database. An equal error rate in the range 3.7-6.8% is reported.
Anastasios Tefas, Constantine Kotropoulos, Ioannis Pitas
CVPR3
1998 Face Verification based on Morphological Shape Decomposition
Anastasios Tefas, Constantine Kotropoulos, Ioannis Pitas
FG3
1998 Face authentication using variants of elastic graph matching based on mathematical morphology that incorporate local discriminant coefficients
abstract
Two novel variants of dynamic link architecture that are based on mathematical morphology and incorporate local coefficients which weigh the contribution of each node according to its discriminatory power in elastic graph matching are proposed, namely, the morphological dynamic link architecture and the morphological signal decomposition-dynamic link architecture. They are tested for face authentication in a cooperative scenario where the candidates claim an identity to be checked. Their performance is evaluated in terms of their receiver operating characteristics and the equal error rate achieved in the M2VTS database. An equal error rate of 6.6%-6.8% is reported.
Constantine Kotropoulos, Anastasios Tefas, Ioannis Pitas
ICASSP3
1998 Motion and Segmentation Prediction in Image Sequences based on Moving Object Tracking
Adrian G. Bors, Ioannis Pitas
ICIP (3)2
1998 Morphological Techniques in the Iterative Closest Point Algorithm
abstract
This paper describes a method for the accurate and computationally efficient registration of 3-D shapes. The method is based on the iterative closest point (ICP) algorithm and improves it by dramatically decreasing the computational cost of the algorithm's most inefficient step, namely the implementation of the closest point operator. The decrease is achieved with the help of a 3-D volume containing the points to be registered. Prior to the implementation of the ICP algorithm, the Voronoi diagram of the "model" points is constructed in the volume, by means of the morphological Voronoi tesselation method with respect to the Euclidean distance metric. The use of the tesselated volume renders the calculation of the closest point operator extremely fast and speeds up the ICP algorithm tremendously.
Christos A. Kapoutsis, C. P. Vavoulidis, Ioannis Pitas
ICIP (1)3
1998 Frontal Face Authentication using Variants of Dynamic Link Matching Based on Mathematical Morphology
abstract
Two variants of dynamic link matching based on mathematical morphology are developed and tested for frontal face authentication, namely, the morphological dynamic link architecture and the morphological signal decomposition-dynamic link architecture. Local coefficients which weigh the contribution of each node in elastic graph matching according to its discriminatory power are derived. The performance of the proposed algorithms is evaluated in terms of their receiver operating characteristic and the equal error rate (EER) achieved in the M2VTS database. The comparison with other frontal face authentication algorithms developed within M2VTS project indicates that morphological dynamic link architecture with discriminatory power coefficients is ranked as the best algorithm in terms of the EER.
Constantine Kotropoulos, Anastasios Tefas, Ioannis Pitas
ICIP (1)3
1998 Speaker Dependent Video Indexing based on Audio-Visual Interaction
abstract
A content-based video indexing method is presented that aims at temporally indexing a video sequence according to the actual speaker. This is achieved by the integration of audio and visual information. Audio analysis leads to the extraction of a speaker identity label versus time diagram. Visual analysis includes scene cut detection, face shot determination, mouth region extraction and tracking and finally talking face shot determination. Results from both sources are combined to improve speaker dependent video indexing. Such a task enables flexible video retrieval or browsing in cases where queries according to speaker identities are imposed. Speaker recognition errors are reduced to 2%.
Sofia Tsekeridou, Ioannis Pitas
ICIP (1)2
1998 Chaotic Watermarks for Embedding in the Spatial Digital Image Domain
abstract
Binary (bi-valued) digital signals, which are suitable watermarks for digital images are presented. They are generated by using chaotic dynamical systems that provide sufficient watermark complexity and controlled lowpass characteristics. Watermark detection is performed without resorting to the original image and its reliability is studied. Efficient robustness under lossy compression, lowpass filtering and other image processing can be achieved. Possible watermark detection after geometrical transformations is discussed.
George Voyatzis, Ioannis Pitas
ICIP (2)2
1998 Digital image watermarking using mixing systems
George Voyatzis, Ioannis Pitas
Comput. Graph.2
1998 Robust image watermarking in the spatial domain
Nikos Nikolaidis 0001, Ioannis Pitas
Signal Process.2
1998 A novel method for automatic face segmentation, facial feature extraction and tracking
Karin Sobottka, Ioannis Pitas
Signal Process. Image Commun.2
1998 A method for watermark casting on digital image
abstract
Watermark casting on digital images is an important problem since it affects many aspects of the information market. We propose a method for casting digital watermarks on images, and we analyze its effectiveness. The satisfaction of some basic demands in this area is examined, and a method for producing digital watermarks is proposed. Moreover, issues like immunity to subsampling and image-dependent watermarks are examined, and simulation results are provided for the verification of the above-mentioned topics.
Ioannis Pitas
IEEE Trans. Circuits Syst. Video Technol.1
1998 Optical flow estimation and moving object segmentation based on median radial basis function network
abstract
Various approaches have been proposed for simultaneous optical flow estimation and segmentation in image sequences. In this study, the moving scene is decomposed into different regions with respect to their motion, by means of a pattern recognition scheme. The inputs of the proposed scheme are the feature vectors representing still image and motion information. Each class corresponds to a moving object. The classifier employed is the median radial basis function (MRBF) neural network. An error criterion function derived from the probability estimation theory and expressed as a function of the moving scene model is used as the cost function. Each basis function is activated by a certain image region. Marginal median and median of the absolute deviations from the median (MAD) estimators are employed for estimating the basis function parameters. The image regions associated with the basis functions are merged by the output units in order to identify moving objects.
Adrian G. Bors, Ioannis Pitas
IEEE Trans. Image Process.2
1997 Mosaicing of Flattened Images from Straight Homogeneous Generalized Cylinders
Adrian G. Bors, William Puech, Ioannis Pitas, Jean-Marc Chassery
CAIP3
1997 Morphological Iterative Closest Point Algorithm
C. P. Vavoulidis, Ioannis Pitas
CAIP2
1997 Perspective distortion analysis for mosaicing images painted on cylindrical surfaces
abstract
A set of monocular images of a curved painting is taken from different viewpoints around its curved surface. After deriving the surface localization in the camera coordinate system we backproject the image on the curved surface and we flatten it. We analyze the perspective distortions of the scene in the case when it is mapped on a cylindrical surface. Based on the result of this analysis we derive the necessary number of views in order to represent the entire scene depicted on a cylindrical surface. We employ a matching-based mosaicing method for reconstructing the scene from the curved surface. The proposed method is appropriate to be used for painting reconstruction.
Adrian G. Bors, William Puech, Ioannis Pitas, Jean-Marc Chassery
ICASSP3
1997 Rule-based face detection in frontal views
abstract
Face detection is a key problem in building automated systems that perform face recognition. A very attractive approach for face detection is based on multiresolution images (also known as mosaic images). Motivated by the simplicity of this approach, a rule-based face detection algorithm in frontal views is developed that extends the work of G. Yang and T.S. Huang (see Pattern Recognition, vol.27, no.1, p.53-63, 1994). The proposed algorithm has been applied to frontal views extracted from the European ACTS M2VTS database that contains the videosequences of 37 different persons. It has been found that the algorithm provides a correct facial candidate in all cases. However, the success rate of the detected facial features (e.g. eyebrows/eyes, nostrils/nose, and mouth) that validate the choice of a facial candidate is found to be 86.5% under the most strict evaluation conditions.
Constantine Kotropoulos, Ioannis Pitas
ICASSP2
1997 Face Authentication Based on Morphological Grid Matching
abstract
A novel dynamic link architecture based on multiscale morphological dilation-erosion is proposed for face verification in a cooperative scenario where the candidates claim an identity that is to be checked. The performance of the morphological dynamic link architecture (MDLA) is evaluated in terms of the receiver operating characteristic (ROC) for several threshold selections on the matching error in the M2VTS database. The experimental results indicate that the proposed method outperforms the dynamic link matching with Gabor based feature vectors.
Constantine Kotropoulos, Ioannis Pitas
ICIP (1)2
1997 Fuzzy cell Hough transform for curve detection
Vassilios Chatzis, Ioannis Pitas
Pattern Recognit.2
1997 Cylindrical surface localization in monocular vision
William Puech, Jean-Marc Chassery, Ioannis Pitas
Pattern Recognit. Lett.3
1996 Segmentation and Tracking of Faces in Color Images
abstract
The authors present a new approach for automatically segmentation and tracking of faces in color images. Segmentation of faces is performed by evaluating color and shape information. First, skin-like regions are determined based on the color attributes hue and saturation. Then regions with elliptical shape are selected as face hypotheses. They are verified by searching for facial features in their interior. After a face is reliably detected it is tracked over time. Tracking is realized by using an active contour model. The exterior forces of the snake are defined based on color features. They push or pull snaxels perpendicular to the snake. Results for tracking are shown for an image sequence consisting of 150 frames.
Karin Sobottka, Ioannis Pitas
FG2
1996 Copyright protection of images using robust digital signatures
abstract
A method for copyright protection of digital images is presented. Copyright protection is achieved by embedding an "invisible" signal, known as digital signature, in the digital image. Signature casting is performed in the spatial domain by slightly modifying the intensity level of randomly selected image pixels. Signature detection is done by comparing the mean intensity value of the marked pixels against that of the not marked pixels. Statistical hypothesis testing is used for this purpose. The signature can be designed in such a way that it is resistant to JPEG compression and lowpass filtering. This is done by minimizing the energy content of the signature signal in higher frequencies. Experiments on real image data verify the effectiveness of the method.
Nikos Nikolaidis 0001, Ioannis Pitas
ICASSP2
1996 Image watermarking using DCT domain constraints
abstract
Watermarking algorithms are used for image copyright protection. The algorithms proposed select certain blocks in the image based on a Gaussian network classifier. The pixel values of the selected blocks are modified such that their discrete cosine transform (DCT) coefficients fulfil a constraint imposed by the watermark code. Two different constraints are considered. The first approach consists of embedding a linear constraint among selected DCT coefficients and the second one defines circular detection regions in the DCT domain. A rule for generating the DCT parameters of distinct watermarks is provided. The watermarks embedded by the proposed algorithms are resistant to JPEG compression.
Adrian G. Bors, Ioannis Pitas
ICIP (3)2
1996 Multichannel adaptive L-filters in color image filtering
abstract
Three novel adaptive multichannel L-filters based on marginal data ordering are proposed. They rely on well-known algorithms for the unconstrained minimization of the mean squared error (MSE), namely, the least mean squares (LMS), the normalized LMS (NLMS) and the LMS-Newton (LMSN) algorithm. Performance comparisons in color image filtering have been made both in RGB and U/sup */V/sup */W/sup */ color spaces. The proposed adaptive multichannel L-filters outperform the other candidates in noise suppression for color images corrupted by mixed impulsive and additive white contaminated Gaussian noise.
Constantine Kotropoulos, Ioannis Pitas, Maria Gabrani
ICIP (1)2
1996 A method for signature casting on digital images
abstract
Signature (watermark) casting on digital images is an important problem, since it affects many aspects of the information market. We propose a method for casting digital watermarks on images and we analyze its effectiveness. The satisfaction of some basic demands in this area is examined and a method for producing digital watermarks is proposed. Moreover, immunity to subsampling is examined and simulation results are provided for the verification of the above mentioned topics.
Ioannis Pitas
ICIP (3)1
1996 Face localization and facial feature extraction based on shape and color information
abstract
Recognition of human faces out of still images or image sequences is a research field of fast increasing interest. At first, facial regions and facial features like eyes and mouth have to be extracted. In the present paper we propose an approach that copes with problems of these first two steps. We perform face localization based on the observation that human faces are characterized by their oval shape and skin-color, also in the case of varying light conditions. For that we segment faces by evaluating shape and color (HSV) information. Then face hypotheses are verified by searching for facial features inside of the face-like regions. This is done by applying morphological operations and minima localization to intensity images.
Karin Sobottka, Ioannis Pitas
ICIP (3)2
1996 Applications of toral automorphisms in image watermarking
abstract
Digital watermarking methods have been proposed for various purposes and especially for copyright protection of multimedia data. The digital watermark is embedded in a digital signal or an image and must be unrecognizable by unauthorized persons and detectable only by the legal copyright owner. We use toral automorphisms as chaotic 2-D integer vector generators in order to manipulate digital image watermarking. We propose also an embedding algorithm which provides robustness under filtering and compression.
George Voyatzis, Ioannis Pitas
ICIP (2)2
1996 Introducing the select and split fuzzy cell Hough transform
abstract
In this paper a new variation of Hough transform is proposed. The parameter space of Hough transform is recursively roughly split into fuzzy cells which are defined as fuzzy numbers. This fuzzy partition of the parameter space provides the advantage to use the uncertainty of the contour points location, which is increased when noisy images have to be used. It can be used to detect contours in an image, with better accuracy, especially in noisy images. Moreover, the regions of the parameter space having no contours are rejected during the iterations. As a result the computation time is significantly decreased.
Vassilios Chatzis, Ioannis Pitas
ICPR2
1996 Extraction of facial regions and features using color and shape information
abstract
There are many applications for systems coping with the problem of face localization and recognition, e.g. model-based video coding, security systems and mug shot matching. Due to variations in illumination, back-ground, visual angle and facial expressions, the problem of machine face recognition is complex. In this paper we present a robust approach for the extraction of facial regions and features out of color images. First, face candidates are located based on the color and shape information. Then the topographic grey-level relief of facial regions is evaluated to determine the position of facial features as eyes and month. Results are shown for two example scenes.
Karin Sobottka, Ioannis Pitas
ICPR2
1996 Mosaicing of paintings on curved surfaces
abstract
The paper presents an approach for reconstructing images painted on curved surfaces. A set of monocular images is taken from different viewpoints in order to mosaic and represent the entire scene. By using a priori knowledge about the support surface of the picture, we derive the surface localization in the camera coordinate system. An automatic mosaicing method is applied on the patterned images in order to obtain the complete scene. The mosaiced scene is visualized on a new synthetic surface by a mapping procedure.
William Puech, Jean-Marc Chassery, Adrian G. Bors, Ioannis Pitas
WACV4
1996 Nonlinear adaptive filters for speckle suppression in ultrasonic images
Eleftherios Kofidis, Sergios Theodoridis, Constantine Kotropoulos, Ioannis Pitas
Signal Process.4
1996 Multichannel L filters based on reduced ordering
abstract
Nonlinear multichannel signal processing is an emerging research topic with numerous applications. In this paper we use the so-called reduced ordering (R-ordering) principle to introduce a new family of L filters for vector-valued observations. The coefficients of the proposed filters can be deduced so that the filters are optimal with respect to the output mean squared error. Expressions for the unconstrained, unbiased and location invariant optimal filter coefficients are derived. The calculation of moments of the R-ordered vectors that are involved in these expressions is also discussed. Experiments with noisy two-channel vector fields and noisy color images are presented in order to demonstrate the superiority of the proposed filters over other multichannel filters.
Nikos Nikolaidis 0001, Ioannis Pitas
IEEE Trans. Circuits Syst. Video Technol.2
1996 Guest Editorial Introduction to the Special Issue on Nonlinear Image Processing
Gonzalo R. Arce, Petros Maragos, Yrjö Neuvo, Ioannis Pitas
IEEE Trans. Image Process.4
1996 Morphological residual representations of signals
abstract
A general approach for residual representation of one-dimensional (1-D) and two-dimensional (2-D) signals is defined. Signals are reconstructed as a sum of components that are recursively determined by using a constructive transform. Several morphological constructive transforms are proposed and their corresponding representations are discussed. The use of residual representations in signal and image compression is investigated with promising results.
Dinu Coltuc, Ioannis Pitas
IEEE Trans. Image Process.2
1996 Adaptive LMS L-filters for noise suppression in images
abstract
Several adaptive least mean squares (LMS) L-filters, both constrained and unconstrained ones, are developed for noise suppression in images and compared in this paper. First, the location-invariant LMS L-filter for a nonconstant signal corrupted by zero-mean additive white noise is derived. It is demonstrated that the location-invariant LMS L-filter can be described in terms of the generalized linearly constrained adaptive processing structure proposed by Griffiths and Jim (1982). Subsequently, the normalized and the signed error LMS L-filters are studied. A modified LMS L-filter with nonhomogeneous step-sizes is also proposed in order to accelerate the rate of convergence of the adaptive L-filter. Finally, a signal-dependent adaptive filter structure is developed to allow a separate treatment of the pixels that are close to the edges from the pixels that belong to homogeneous image regions.
Constantine Kotropoulos, Ioannis Pitas
IEEE Trans. Image Process.2
1996 Multichannel techniques in color image enhancement and modeling
abstract
We present novel multichannel methods in two target research areas. The first area is color image modeling. Multichannel AR models have been developed and applied to color texture segmentation and synthesis. The second area is color image equalization, which is performed on the three RGB channels simultaneously, using the joint PDF. Alternatively, equalization at the HSI domain is performed in order to avoid changes in digital image hue. A parallel algorithm is proposed for color image histogram calculation and equalization.
Ioannis Pitas, P. Kiniklis
IEEE Trans. Image Process.1
1996 Multichannel transforms for signal/image processing
abstract
This paper presents a novel approach to the Fourier analysis of multichannel time series. Orthogonal matrix functions are introduced and are used in the definition of multichannel Fourier series of continuous-time periodic multichannel functions. Orthogonal transforms are proposed for discrete-time multichannel signals as well. It is proven that the orthogonal matrix functions are related to unitary transforms (e.g., discrete Hartley transform (DHT), Walsh-Hadamard transform), which are used for single-channel signal transformations. The discrete-time one-dimensional multichannel transforms proposed in this paper are related to two-dimensional single-channel transforms, notably to the discrete Fourier transform (DFT) and to the DHT. Therefore, fast algorithms for their computation can be easily constructed. Simulations on the use of discrete multichannel transforms on color image compression have also been performed.
Ioannis Pitas, Anestis Karasaridis
IEEE Trans. Image Process.1
1996 Order statistics learning vector quantizer
abstract
We propose a novel class of learning vector quantizers (LVQs) based on multivariate data ordering principles. A special case of the novel LVQ class is the median LVQ, which uses either the marginal median or the vector median as a multivariate estimator of location. The performance of the proposed marginal median LVQ in color image quantization is demonstrated by experiments.
Ioannis Pitas, Constantine Kotropoulos, Nikos Nikolaidis 0001, Ruikang Yang, Moncef Gabbouj
IEEE Trans. Image Process.1
1996 Median radial basis function neural network
abstract
Radial basis functions (RBFs) consist of a two-layer neural network, where each hidden unit implements a kernel function. Each kernel is associated with an activation region from the input space and its output is fed to an output unit. In order to find the parameters of a neural network which embeds this structure we take into consideration two different statistical approaches. The first approach uses classical estimation in the learning stage and it is based on the learning vector quantization algorithm and its second-order statistics extension. After the presentation of this approach, we introduce the median radial basis function (MRBF) algorithm based on robust estimation of the hidden unit parameters. The proposed algorithm employs the marginal median for kernel location estimation and the median of the absolute deviations for the scale parameter estimation. A histogram-based fast implementation is provided for the MRBF algorithm. The theoretical performance of the two training algorithms is comparatively evaluated when estimating the network weights. The network is applied in pattern classification problems and in optical flow segmentation.
Adrian G. Bors, Ioannis Pitas
IEEE Trans. Neural Networks2
1995 Segmentation and Estimation of the Optical Flow
Adrian G. Bors, Ioannis Pitas
CAIP2
1995 Edge detection operators for angular data
abstract
Physical quantities referring to angles, like vector direction, color hue etc., are periodic in nature. Due to this periodicity edge detectors proposed for data on the line cannot be used to detect edges in angle-related signals. In this paper we use estimators of circular dispersion to introduce edge detectors for angular signals and discuss their application in edge detection on hue images. Extensions of the notion of quasi-range to circular data are also proposed. These "circular" quasi-ranges have good and user-controlled properties as edge detectors on noisy angular signals. The performance of the proposed edge operators is evaluated on angular edges, using both qualitative and quantitative criteria.
Nikos Nikolaidis 0001, Ioannis Pitas
ICIP2
1995 Parallel Digital Signal Filtering on Barrel Shifter Computers
abstract
Parallel algorithms on barrel shifter computers for a broad class of 1-D and 2-D signal operators are presented in this paper. Running max/min selection filter, moving average filter and sorting, which can be used in order statistics filtering, are examined. The proposed algorithms require a significantly smaller number of comparisons/computations per output point than the conventional ones. They are suitable to be implemented on barrel shifter computer topologies.
Georgios Angelopoulos, Ioannis Pitas
ISCAS2
1995 Jordan Decomposition Filters
abstract
In this paper we propose a new class of filters called Jordan decomposition filters. They are upper and lower bounds of the well-known max and min filters. Our approach is based on a decomposition scheme by representing bounded variation signals as a difference of two increasing functions. Jordan filters are computed as the difference between the local extrema values of the increasing functions. They are faster than max/min filters. The number of operations per sample is constant, regardless the size of the window filter. The computational complexity of the proposed filters decreases with the number of filtering steps needed in an application. Examples are provided and the application in contrast enhancement is investigated.
Dinu Coltuc, Ioannis Pitas
ISCAS2
1994 Multichannel L-filter design based on marginal data ordering
abstract
We address the design of multichannel L-filters based on marginal data ordering using the mean-squared-error as fidelity criterion. Design procedures subject to the constraints of unbiased or location-invariant estimation or without imposing any constraint are discussed. It is shown by simulations that the proposed multichannel L-filters perform better than other multichannel nonlinear filters such as the vector median, the marginal alpha-trimmed mean, the marginal median, the multichannel modified trimmed mean and the multichannel double-window trimmed mean, the multivariate ranked-order estimators as well as their single-channel counterparts.>
Constantine Kotropoulos, Ioannis Pitas
ICASSP (3)2
1994 Application of directional statistics in vector direction estimation
abstract
Vector direction estimation can be very important in applications like hue color component filtering or motion direction estimation from noisy motion vector fields. Vector representation and manipulation in polar coordinates greatly facilitates the accomplishment of the previous task. Based on angular estimators of location, the authors introduce a number of "circular" filters, i.e. filters for angular input data. These filters include circular mean, median and a-trimmed mean filters. Emphasis is given to a circular median filter for which the approximate output pdf as well as other interesting properties are derived. The effectiveness of angular estimators of location in noise filtering is studied in simulations involving color images. Simulation results clearly indicate that circular filters can be used effectively to remove noise when the estimation of color hue is of primary importance.>
Nikos Nikolaidis 0001, Ioannis Pitas
ICASSP (5)2
1994 Discrete multichannel orthogonal transforms
abstract
This paper presents a novel approach to the Fourier analysis of multichannel time series. Orthogonal matrix functions are introduced and are used in the definition of multichannel Fourier series of continuous-time periodic multichannel functions. Orthogonal transforms are proposed for discrete-time multichannel signals as well. The discrete-time one-dimensional multichannel transforms proposed in this paper are related to two-dimensional single-channel transforms, notably to the discrete Fourier transform and to the discrete Hartley transform. Therefore, fast algorithms for their computation can be easily constructed. Simulations on the use of discrete multichannel transforms on color image compression have also been performed.>
Ioannis Pitas, Anestis Karasaridis
ICASSP (5)1
1994 Two-Dimensional Vector Median Filters on Mesh Connected Computers
abstract
This paper presents a parallel algorithm for the calculation of vector median filters on mesh connected computers. A straightforward algorithm to find the vector median requires the computation of the sum of the distances (norms) of each vector to all other vectors in the filter window. By taking into account the fact that adjacent filter windows share common input samples, the computation of the common norms of adjacent windows could be done only once. In the proposed algorithm, a 2-D signal is assumed to be stored in a SIMD mesh connected computer of the same size. According to this scheme, each processing element (PE) computes some norms and the rest ones, that are required for the selection of the vector median, are received from its neighbouring PEs via the local communication channels.>
Georgios Angelopoulos, Ioannis Pitas
ICIP (3)2
1994 Multichannel Distance Filters
abstract
Two nonlinear digital filters for multichannel signal processing are presented in this paper. A method for the approximate calculation of output mean and variance for one of these filters is proposed. The results of simulations, showing great advantages of the described filters, are also presented.>
Andrzej Buchowicz, Ioannis Pitas
ICIP (2)2
1994 Cellular LMS L-filters for Noise Suppression in Still Images and Image Sequences
abstract
A novel class of nonlinear adaptive L-filters based on cellular neural networks topology is presented. Like cellular neural systems and cellular automata as well, processing nodes, called cells, communicate with each other directly only through its nearest neighbors exchanging information. Each cell is an adaptive LMS L-filter. The proposed filters share the best features of both adaptive filters and cellular neural network topologies; their adaptive structure tracks image nonstationarities and their local interconnection feature makes it suitable for VLSI implementation. Cellular adaptive LMS L-filters are suited for high-speed parallel adaptive image filtering. Some interesting applications to image and image sequence filtering are demonstrated.>
Maria Gabrani, Constantine Kotropoulos, Ioannis Pitas
ICIP (1)3
1994 A fast implementation of two-dimensional weighted median filters
abstract
This paper deals with the implementation of a fast algorithm for two-dimensional weighted median filtering. Because of the vast amount of data that must be handled, the development of fast algorithms is very important. A fast running algorithm for weighted median filtering, which is based on using a histogram and updating it, is proposed. Experimental results prove the superiority of the proposed algorithm over that of finding the weighted median either by sorting with the Quick Sort or by selecting the r-th order statistic.
Georgios Angelopoulos, Ioannis Pitas
ICPR (3)2
1994 Adaptive LMS L-filters for smoothing noisy images
abstract
Several adaptive LMS L-filters, both constrained and unconstrained ones, are developed for noise suppression in images and being compared in this paper. First, the location-invariant LMS L-filter for a nonconstant signal corrupted by zero-mean additive white noise is derived. Subsequently, the normalized and the sign LMS L-filters are studied. It is shown that both these filters turn to be identical for a certain choice of the adaptation step-size. A modified LMS L-filter with nonhomogeneous step-sizes is also proposed in order to accelerate the rate of convergence of the adaptive L-filter. Finally, a signal-dependent adaptive filter structure is developed to allow a separate treatment of the pixels that are close to the edges from the pixels that belong to homogeneous image regions.
Constantine Kotropoulos, Ioannis Pitas
ICPR (3)2