Anastasios Tefas

dblp:22/6949 · DBLP profile ↗
← Back
280ranked-venue papers
13as first author
63since 2021 · last 2026
0000-0003-1288-3667ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 143 · 4 first-author · 45 since 2021Graphics, computer vision, multimedia, augmented reality and games · 143 · 10 first-author · 19 since 2021Systems, architecture and hardware · 9 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 6 · 2 since 2021Security and privacy · 3Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 2
YearPublicationVenuePosition
2026 Optics-Informed Long Short Term Memory Cells
Loukia Avramelou, Manos Kirtas, Nikolaos Passalis, Nikos Pleros, Anastasios Tefas
ICPR (14)5
2026 An Isomorphic Transformation for Hardware Implemented Non-negative Neural Networks
Manos Kirtas, Nikolaos Passalis, Nikos Pleros, Anastasios Tefas
ISCAS4
2026 Single-pass Anchor-based Uncertainty Quantification in Deep Neural Networks
Dimitrios Spanos, Nikolaos Passalis, Anastasios Tefas
ISCAS3
2026 Non-negative isomorphic neural networks for efficient accelerators
abstract
High speed and efficient accelerators such as neuromorphic photonics are becoming increasingly popular, as they can significantly improve computation speed and energy efficiency of matrix-based calculations, leading to femtojoule per MAC efficiency. However, deploying existing DL models on such platforms is not trivial, as a wide range of photonic neural network (PNN) architectures relies on incoherent setups and power-adding operational schemes that cannot natively represent negative quantities. This results in additional hardware complexity that increases cost and reduces energy efficiency. To overcome this, we can train non-negative neural networks and potentially exploit the full range of incoherent neuromorphic photonic capabilities. However, the existing non-negative training approaches cannot achieve the same level of accuracy as their regular counterparts. To this end, we introduce a methodology to obtain the non-negative isomorphic equivalents of regular artificial neural networks (ANNs) that meet the requirements of neuromorphic hardware as it allows one to perform inference with non-negative only parameters. Furthermore, we also introduce a sign-preserving optimization approach that enables the training of such isomorphic networks in a non-negative manner. The source code implementation of the proposed numerical framework is publicly available at: https://github.com/eakirtas/nn_isomorphics_24 .
Manos Kirtas, Nikolaos Passalis, Nikos Pleros, Anastasios Tefas
Neurocomputing4
2026 Detecting and mitigating training anomalies in deep neural networks
Vasilios Moustakidis, Nikolaos Passalis, Anastasios Tefas
Neurocomputing3
2026 Improving uncertainty estimation in deep neural networks by modeling class and instance uncertainty
abstract
• A deep learning method that models data uncertainty for uncertainty estimation is proposed. • A method that generates soft labels in an adaptive manner is used for the optimization. • The uncertainty-aware learning process improves state-of-the-art methodologies in out-of-distribution detection. In the pursuit of enhancing the trustworthiness of deep learning models, there has been a growing interest in improving their uncertainty estimation for more reliable decision-making. Recent methodologies have concentrated on refining uncertainty estimation particularly in the context of identifying out-of-distribution input samples. This work introduces and explores two distinct categories of uncertainty: class uncertainty and instance uncertainty . The former highlights uncertainty stemming from inter-class resemblances, while the latter addresses analogous uncertainty within individual instances. We propose a novel framework, Adaptive Similarity Labeling (ASL), that captures both class-level and instance-level uncertainty by adaptively assigning soft labels based on semantic similarity and instance difficulty. ASL mitigates feature collapse and improves uncertainty estimation without adding inference-time overhead, and can be used regardless of the uncertainty metric used. The proposed approach is architecture-agnostic and can be combined with recent state-of-the-art architectures for uncertainty estimation. Experiments conducted on several datasets demonstrate the effectiveness of ASL in improving uncertainty estimation and out-of-distribution detection across diverse tasks.
Dimitrios Spanos, Nikolaos Passalis, Anastasios Tefas
Pattern Recognit.3
2026 Probabilistic image-based joint embedding predictive architecture
abstract
In the field of self-supervised learning, the Image-based Joint-Embedding Predictive Architecture (I-JEPA) has shown impressive results in capturing semantic information from images without the need for strong augmentations between views. I-JEPA learns to predict the representations of multiple masked target blocks within the same image, using the representation of a single unmasked context block. To achieve this, I-JEPA reconstructs the masked patches in the representation space rather than at the pixel level, as done in previous works. However, accurately reconstructing abstract representations of patches is a challenging task, as pointwise matching focuses on exact reconstruction of individual target representations rather than the preservation of their local geometric structure. In this work, we present a probabilistic matching approach that can be applied to I-JEPA’s learning process to encourage the output of semantically richer representations. Specifically, we propose a probabilistic formulation of I-JEPA, which involves efficiently modeling and matching the probability distributions of the missing patches between the target block representations and the network’s predictions based solely on the context block. Through rigorous evaluations, we demonstrate that the proposed methodology provides consistent improvements over I-JEPA in downstream classification and transfer learning tasks.
Lazaros Gogos, Dimitrios Katsikas, Nikolaos Passalis, Anastasios Tefas
Pattern Recognit. Lett.4
2025 Single Pass Uncertainty Estimation in Q-Learning via Dropout Distillation
abstract
While Reinforcement Learning (RL) has demonstrated success in numerous applications, ensuring the consistency and safety of learned policies remains a significant challenge, particularly in safety-critical domains. To enable the reliable and safe operation of RL agents, it is essential to quantify the agent’s uncertainty regarding its action selection, commonly referred to as Policy Uncertainty. In this work, we propose a novel approach for estimating policy uncertainty in Deep Q-learning (DQN) agents by leveraging Dropout Distillation Networks. Our methodology consists of three key steps: (1) training a DQN agent with dropout applied to its network layers, (2) using the trained agent to interact with the RL environment and collect a distillation dataset, and (3) training a student network (DDN) using the collected dataset. Furthermore, we identify a gap in the literature, where existing works primarily introduce uncertainty estimation algorithms without assessing their estimation accuracy. To address this, we propose a novel RL policy uncertainty evaluation algorithm capable of assessing the calibration of any agent that provides policy uncertainty estimates, utilizing a simulator of the RL environment. We compare our proposed policy uncertainty estimation approach against Monte Carlo Dropout, a widely used technique in the literature. Experimental results indicate that our method consistently achieves better calibration while requiring only a single forward pass, a crucial advantage for real-time RL applications.
Nikolaos Kaparinos, Nikolaos Passalis, Anastasios Tefas
IJCNN3
2025 Multiplicative Stochastic Gradient Descent for fast and robust deep learning training
abstract
Even recent Deep Learning (DL) architectures are highly sensitive to training hyperparameters, initial weights, and data distributions, making the development of fast and stable optimization methods a challenging task. To this end, generations of researchers have pursued the development of robust methods to train DL architectures. Multiplicative updates naturally take into account the magnitude of parameters and despite their significant contributions on the early development of Machine Learning (ML) and the strong theoretical claims, they have been merely studied in the context of DL. To this end, in this work, we propose two alternative Stochastic Gradient Descent (SGD) optimization methods that are based on a novel multiplicative update term, which proportionally scales the parameters using the normalized gradients, decoupling in this way the magnitude of gradients from the update. More specifically, we formulate two SGD alternatives, one multiplicative and another hybrid multiplicative-additive, which overcomes the limitation of purely multiplicative updates that constraints the change of parameter’s sign. The proposed methods accelerate training while leading to more robust models in contrast to traditionally used SGD, and we experimentally demonstrate its effectiveness in a wide range of tasks. Such tasks range from convex and non-convex optimization to difficult image classification benchmarks applying a wide range of Artificial Neural Network (ANN) architectures, providing quantitative and qualitative experimental results.
Manos Kirtas, Nikolaos Passalis, Anastasios Tefas
IJCNN3
2025 Leveraging active perception for real-time high-resolution pose estimation
Theodoros Manousis, Emmanouil Eleftheriadis, Nikolaos Passalis, Anastasios Tefas
Expert Syst. Appl.4
2025 Leveraging subclass learning for improving uncertainty estimation in deep Learning
abstract
Machine learning is becoming increasingly popular across various applications and has led to state-of-the-art results, but it faces challenges related to its trustworthiness. One aspect of making deep learning models more trustworthy is improving their ability to estimate the uncertainty of whether a sample is from the in-domain (ID) data distribution or not. Especially, neural networks have a tendency to make overly confident extrapolations and struggle to convey their uncertainty, which can limit their trustworthiness. Recent approaches have employed Radial Basis Function (RBF)-based models with great success in improving uncertainty estimation in Deep Learning. However, such models assume a unimodal distribution of the data for each class, which we show is critical for out-of-distribution sample detection, but can be limiting in many real world cases. To overcome these limitations, in this paper, we propose a method for training a deep model utilizing the inherent different modalities that naturally arise in a class in real data, which we call subclasses , leading to improved uncertainty quantification. The proposed method leverages a variance-preserving reconstruction-based representation learning approach that prevents feature collapse and enables robust discovery of subclasses, further improving the effectiveness of the proposed approach. The improvement of the approach is demonstrated using extensive experiments on several datasets.
Dimitrios Spanos, Nikolaos Passalis, Anastasios Tefas
Neurocomputing3
2025 Trustworthy knowledge distillation via anchor-guided distribution learning
abstract
Existing knowledge distillation methods typically do not adequately account for the representativeness of sampled training data or the adverse effects of missing classes within mini-batches, both of which can lead to suboptimal knowledge transfer. In this paper, we introduce A nchor-based K nowledge D istillation (AKD), a method that leverages the most informative and representative samples, referred to as anchors , during the transfer process, and matches the representation distribution of the data in the feature space rather than their actual representations. Anchors are strategically chosen based on their ability to encapsulate critical features of the data distribution, ensuring that the student model focuses on the most informative aspects of the teacher’s knowledge. The proposed method enables the introduction of information from a wider variety of classes at each mini-batch iteration, ensuring a more balanced data distribution and thereby improving the knowledge transfer process. Furthermore, by leveraging the static nature of anchors, we enhance the student model’s representation learning through attention maps, improving both convergence and generalization. The proposed approach is extensively evaluated across various datasets and tasks, demonstrating that the incorporation of anchors into the knowledge distillation process improving the accuracy and trustworthiness of the resulting models.
Dimitrios Spanos, Nikolaos Passalis, Anastasios Tefas
Knowl. Based Syst.3
2025 Sign potential-driven multiplicative optimization for robust deep reinforcement learning
Loukia Avramelou, Manos Kirtas, Nikolaos Passalis, Anastasios Tefas
Neural Networks4
2025 Explicit Bandwidth Learning for FOREX Trading Using Deep Reinforcement Learning
abstract
Financial time series are sequences of price observations related to financial assets collected over time. Deep Learning (DL) is currently standing as the predominant approach for addressing various time series tasks, including problems in finance, such as the development of trading agents using Deep Reinforcement Learning (DRL). However, the noisy and temporal nature of such data as well as their non-stationarity pose substantial challenges to current methodologies. DL models suffer from overfitting noise, frequently arising from the absence of strong priors. In this paper, we address the instability of trading DRL agents due to noise by proposing an end-to-end hybrid trainable filtering and feature extraction approach. The proposed method employs Gaussian filters as priors and can be attached at the beginning of any DL architecture forming a hybrid model-based and data-driven model that can directly process the raw input data. The bandwidth of the filters is determined through the learning process, ultimately allowing the agent to autonomously determine the optimal bandwidth for the task and data at hand, without requiring any additional supervision. Moreover, the proposed method leverages high-order derivatives to address the non-stationarity of financial data and provides multiple views of the input signal efficiently utilized by the subsequent model. We conduct experiments with a plethora of financial assets from the Foreign Exchange Market (FOREX) and demonstrate the method's efficiency when compared to alternative processing pipelines.
Angelos Nalmpantis, Nikolaos Passalis, Anastasios Tefas
IEEE Signal Process. Lett.3
2025 Inducing Neural Collapse via Anticlasses and One-Cold Cross-Entropy Loss
abstract
While softmax cross-entropy (CE) loss is the standard objective for supervised classification, it primarily focuses on the ground-truth classes, ignoring the relationships between the nontarget, complementary classes. This leaves valuable information unexploited during optimization. In this work, we propose a novel loss function, one-cold CE (OCCE) loss, which addresses this limitation by structuring the activations of these complementary classes. Specifically, for each class, we define an anticlass, which consists of everything that is not part of the target class-this includes all complementary classes as well as out-of-distribution (OOD) samples, noise, or in general any instance that does not belong to the true class. By setting a uniform one-cold encoded distribution over the complementary classes as a target for each anticlass, we encourage the model to equally distribute activations across all nontarget classes. This approach promotes a symmetric geometric structure of classes in the final feature space, increases the degree of neural collapse (NC) during training, addresses the independence deficit problem of neural networks, and improves generalization. Our extensive evaluation shows that incorporating OCCE loss in the optimization objective consistently enhances performance across multiple settings, including classification, open-set recognition, and OOD detection.
Dimitrios Katsikas, Nikolaos Passalis, Anastasios Tefas
IEEE Trans. Neural Networks Learn. Syst.3
2024 Large Models in Dialogue for Active Perception and Anomaly Detection
Tzoulio Chamiti, Nikolaos Passalis, Anastasios Tefas
ICPR (17)3
2024 Plasticity Driven Knowledge Transfer for Continual Deep Reinforcement Learning in Financial Trading
Dimitrios Katsikas, Nikolaos Passalis, Anastasios Tefas
ICPR (9)3
2024 Multiplicative RMSprop Using Gradient Normalization for Learning Acceleration
Manos Kirtas, Nikolaos Passalis, Anastasios Tefas
ICPR (29)3
2024 Deep reinforcement learning for financial trading using multi-modal features
Loukia Avramelou, Paraskevi Nousi, Nikolaos Passalis, Anastasios Tefas
Expert Syst. Appl.4
2024 Multiplicative update rules for accelerating deep learning training and increasing robustness
Manos Kirtas, Nikolaos Passalis, Anastasios Tefas
Neurocomputing3
2024 Vision-based drone control for autonomous UAV cinematography
Ioannis Mademlis, Charalampos Symeonidis, Anastasios Tefas, Ioannis Pitas
Multim. Tools Appl.3
2024 Semi-supervised learning for on-street parking violation prediction using graph convolutional networks
Nikolaos Karantaglis, Nikolaos Passalis, Anastasios Tefas
Neural Comput. Appl.3
2024 Vpit: real-time embedded single object 3D tracking using voxel pseudo images
abstract
Abstract In this paper, we propose a novel voxel-based 3D single object tracking (3D SOT) method called Voxel Pseudo Image Tracking (VPIT). VPIT is the first method that uses voxel pseudo images for 3D SOT. The input point cloud is structured by pillar-based voxelization, and the resulting pseudo image is used as an input to a 2D-like Siamese SOT method. The pseudo image is created in the Bird’s-eye View (BEV) coordinates; and therefore, the objects in it have constant size. Thus, only the object rotation can change in the new coordinate system and not the object scale. For this reason, we replace multi-scale search with a multi-rotation search, where differently rotated search regions are compared against a single target representation to predict both position and rotation of the object. Experiments on KITTI [1] Tracking dataset show that VPIT is the fastest 3D SOT method and maintains competitive Success and Precision values. Application of a SOT method in a real-world scenario meets with limitations such as lower computational capabilities of embedded devices and a latency-unforgiving environment, where the method is forced to skip certain data frames if the inference speed is not high enough. We implement a real-time evaluation protocol and show that other methods lose most of their performance on embedded devices; while, VPIT maintains its ability to track the object.
Illia Oleksiienko, Paraskevi Nousi, Nikolaos Passalis, Anastasios Tefas, Alexandros Iosifidis
Neural Comput. Appl.4
2024 AnIO: anchored input-output learning for time-series forecasting
Ourania Stentoumi, Paraskevi Nousi, Maria Tzelepi, Anastasios Tefas
Neural Comput. Appl.4
2024 Online probabilistic knowledge distillation on cryptocurrency trading using Deep Reinforcement Learning
Vasileios Moustakidis, Nikolaos Passalis, Anastasios Tefas
Pattern Recognit. Lett.3
2023 Residual Error Learning for Electricity Demand Forecasting
Achilleas Andronikos, Maria Tzelepi, Anastasios Tefas
EANN3
2023 Improving Electric Load Demand Forecasting with Anchor-Based Forecasting Method
abstract
In this paper we deal with the problem of Electric Load Demand Forecasting (ELDF) considering the Greek Energy Market. Motivated by the anchored-based object detection methods, we argue that considering the ELDF task we can define an anchor and transform the problem into predicting the offset instead of predicting the actual load values. The experimental evaluation considering the one-day-ahead forecasting task, validated the effectiveness of the proposed Anchor-based FOREcasting (AFORE) method. The AFORE method achieved significant improvements in terms of mean absolute percentage error under various setups, using different loss functions and model architectures.
Maria Tzelepi, Paraskevi Nousi, Anastasios Tefas
ICASSP3
2023 Enabling High-Resolution Pose Estimation in Real Time Using Active Perception
abstract
Deep Learning (DL) models have enabled very accurate pose estimation. However, most of the existing approaches require images of relatively high resolution, since locating body parts and joints accurately is challenging, which increases the computational cost of these approaches. To overcome this limitation in this paper we propose an active perception method for high-resolution pose estimation that enables efficiently selecting the most appropriate image region for analysis and then employing a bottom-up pose estimator on the corresponding region. This allows for significantly improving the efficiency of pose estimation by selectively analyzing in high resolution only the parts of the image that contain humans. To ensure the computational efficiency of the proposed method we propose using low-resolution heat maps extracted using the same pose estimation model in order to guide the active perception process. The proposed method is model agnostic since it can be combined with any bottom-up pose estimation model in order to enable high-resolution analysis. We have experimentally evaluated the proposed method using a well-known pose estimation model, Lightweight OpenPose, demonstrating its effectiveness on three high-resolution variants of the COCO2017 dataset.
Theodoros Manousis, Nikolaos Passalis, Anastasios Tefas
ICIP3
2023 Deep Learning for Active Robotic Perception
Nikolaos Passalis, Pavlos Tosidis, Theodoros Manousis, Anastasios Tefas
IJCCI4
2023 Anchored Input-Output Learning for Electrical Load Demand Forecasting
abstract
In this work we deal with the one-day-ahead electric load demand forecasting problem, namely the task of predicting the electricity demand a day ahead. We draw inspiration from objection detection methods, where the use of anchors facilitates the task by providing the network with some bias resembling the expected output. Based on the same principles, we propose the use of anchors and encode the groundtruth load demand. Furthermore, we propose the use of anchor-encoded input features to match the encoded output. We perform experiments on a dataset collected from the Greek Energy Market and show that the proposed method consistently outperforms the baseline methods. Our best model has a 17.5% relative improvement in terms of Mean Absolute Percentage Error.
Paraskevi Nousi, Maria Tzelepi, Anastasios Tefas
ISCAS3
2023 Tweaking EfficientDet for frugal training
abstract
Object detection appears to be omnipresent nowadays with detectors being available for every problem available, covering solutions from extra-light to ultra resource demanding models. Yet, the vast majority of these approaches are based on large datasets to provide the required feature diversity. This work focuses on object detection solutions which do not rely heavily on abundant training datasets but rather on medium-sized data collections. It uses Efficientdet object detector as base for the application of novel modifications which achieve better performance both in efficiency as well in effectiveness. The focus on medium-sized datasets aim at representing more commonplace datasets which can be accumulated and compiled with relative ease.
Georgios Orfanidis, Konstantinos Ioannidis, Anastasios Tefas, Stefanos Vrochidis, Ioannis Kompatsiaris
ICMR3
2023 Retrieval-Based Methodology for Few-Sample Logo Recognition
abstract
Logo recognition describes the challenging task of detecting and classifying logos in digital images and videos. Former works approach logo recognition as a closed-set problem. However, this approach is accompanied by several shortcomings linked with its incapability of recognising new classes. In this paper, we propose an open-set logo recognition method, named REtrieval-based methodology For FEw-sample LOgo Recognition (REFELOR). REFELOR is composed by a generic logo detector and a feature extractor, allowing the generalization on unseen classes, using only a few samples per logo. That is, a single-stage generic logo detector is trained to detect logos in an input image. Then, feature representations for the detected logos are extracted, using the feature extractor, while the feature representations of a database containing only a few samples per class are also extracted. Finally, the detected logo representations are classified to the corresponding class based a similarity search in the representations of the aforementioned database. In addition, a regularization technique is applied to the feature extractor, providing further improvements. The experimental evaluation validates the effectiveness of the proposed method, outperforming current state-of-the-art logo recognition methods.
Dimosthenis Moralis, Maria Tzelepi, Anastasios Tefas
MMSP3
2023 Leveraging Active and Continual Learning for Improving Deep Face Recognition in-the-Wild
abstract
Face recognition systems play a vital role in various applications by providing identification and verification based on facial features. However, these systems face challenges in large-scale in-the-wild applications, where the current static pipelines are usually unable to cope with the high velocity, variety, and volume of the data, negatively affecting their accuracy and reliability. To overcome these challenges, in this paper, we propose departing from traditional static face recognition pipelines and moving towards dynamic and adaptable approaches. To this end, we propose a combination of an active and continual learning approach that can automatically augment the model with informative samples gathered during the inference, provided that the model is already confident enough, significantly improving its recognition accuracy. Furthermore, the proposed pipeline also natively incorporates active learning, allowing for using human feedback, when available, to further improve its performance. The effectiveness of the proposed method is validated on a challenging setup using two large-scale face recognition datasets.
Pavlos Tosidis, Nikolaos Passalis, Anastasios Tefas
MMSP3
2023 A multiple-UAV architecture for autonomous media production
Ioannis Mademlis, Arturo Torres-González, Jesús Capitán, Maurizio Montagnuolo, Alberto Messina, Fulvio Negro, Cédric Le Barz, Rita Cunha, Bruno J. Guerreiro, Fan Zhang 0017, Stephen Boyle, Gregoire Guerout, Anastasios Tefas, Nikos Nikolaidis 0001, David Bull 0001, Ioannis Pitas
Multim. Tools Appl.14
2023 Mixed-precision quantization-aware training for photonic neural networks
abstract
Abstract The energy demanding nature of deep learning (DL) has fueled the immense attention for neuromorphic architectures due to their ability to operate in a very high frequencies in a very low energy consumption. To this end, neuromorphic photonics are among the most promising research directions, since they are able to achieve femtojoule per MAC efficiency. Although electrooptical substances provide a fast and efficient platform for DL, they also introduce various noise sources that impact the effective bit resolution, introducing new challenges to DL quantization. In this work, we propose a quantization-aware training method that gradually performs bit reduction to layers in a mixed-precision manner, enabling us to operate lower-precision networks during deployment and further increase the computational rate of the developed accelerators while keeping the energy consumption low. Exploiting the observation that intermediate layers have lower-precision requirements, we propose to gradually reduce layers’ bit resolutions, by normally distributing the reduction probability of each layer. We experimentally demonstrate the advantages of mixed-precision quantization in both performance and inference time. Furthermore, we experimentally evaluate the proposed method in different tasks, architectures, and photonic configurations, highlighting its immense capabilities to reduce the average bit resolution of DL models while significantly outperforming the evaluated baselines.
Manos Kirtas, Nikolaos Passalis, Athina Oikonomou, Miltiadis Moralis-Pegios, George Giamougiannis, Apostolos Tsakyridis, George Mourgias-Alexandris, Nikos Pleros, Anastasios Tefas
Neural Comput. Appl.9
2023 Modeling limit order trading with a continuous action policy for deep reinforcement learning
Avraam Tsantekidis, Nikolaos Passalis, Anastasios Tefas
Neural Networks3
2023 Mutual Information-Based Neural Network Distillation for Improving Photonic Neural Network Training
Alexandros Chariton, Nikolaos Passalis, Nikos Pleros, Anastasios Tefas
Neural Process. Lett.4
2022 A Robust, Quantization-Aware Training Method for Photonic Neural Networks
Athina Oikonomou, Manos Kirtas, Nikolaos Passalis, George Mourgias-Alexandris, Miltiadis Moralis-Pegios, Nikos Pleros, Anastasios Tefas
EANN7
2022 Improving Binary Semantic Scene Segmentation for Robotics Applications
Maria Tzelepi, Nikolaos Tragkas, Anastasios Tefas
EANN3
2022 Sentiment-Aware Distillation for Bitcoin Trend Forecasting Under Partial Observability
abstract
Deep Learning (DL) models are increasingly used for financial forecasting problems, such as price or trend prediction of a financial asset. However, most methods either rely solely on price information or require difficult to implement data harvesting pipelines, e.g., from social media, to deploy them. The main contribution of this paper is a method that exploits sentiment information as a source of additional supervision during the training process, allowing for improving the profitability of the developed strategies compared to baseline agents, while also allowing for operating the agent under partial observability, i.e., without requiring sentiment information as input during inference. As demonstrated in the conducted experiments on the Bitcoin-USD currency pair, this approach can indeed lead to significant improvements in the performance of DL agents, as well as help reduce the overfitting phenomena that often occur when training such agents.
Georgios Panagiotatos, Nikolaos Passalis, Avraam Tsantekidis, Anastasios Tefas
ICASSP4
2022 Bag-Of-Features-Based Knowledge Distillation For Lightweight Convolutional Neural Networks
abstract
Knowledge distillation enables us to transfer the knowledge from a large and complex neural network into a smaller and faster one. This allows for improving the accuracy of the smaller network. However, directly transferring the knowledge between enormous feature maps, as they are extracted from convolutional layers, is not straightforward. In this work, we propose an efficient mutual information-based approach for transferring the knowledge between feature maps extracted from different networks. The proposed method employs an efficient Neural Bag-of-Features formulation to estimate the joint and marginal probabilities and then optimizes the whole pipeline in an end-to-end manner. The effectiveness of the proposed method is demonstrated using a lightweight, fully convolutional neural network architecture, which aims toward high-resolution analysis and targets photonic neural network accelerators.
Alexandros Chariton, Nikolaos Passalis, Anastasios Tefas
ICIP3
2022 Multilayer Online Self-Acquired Knowledge Distillation
abstract
Online knowledge distillation has been proposed as an auspicious approach for circumventing the flaws of the conventional offline distillation (i.e., complex, and computationally and memory demanding process). In this work, a novel online self-distillation method, named Multilayer Online Self-Acquired Knowledge Distillation (MOSAKD), is proposed, aiming to develop fast-to-execute and effective models that can comply with applications with memory and computational restrictions, e.g., robotics applications. The MOSAKD method is able to mine additional knowledge both from the intermediate and the output layers of a deep neural model in an online fashion. To achieve this goal, k-nn non-parametric density estimation for estimating the unknown probability distributions of the data samples in the feature space generated by any neural layer is used. This enables us to compute the soft labels that explicitly express the similarities of the data with the classes, by directly estimating the posterior class probabilities of the data samples. The experimental evaluation on four datasets, including a dataset of synthetic images, indicates the effectiveness of the MOSAKD method and the superiority over existing online distillation methods.
Maria Tzelepi, Charalampos Symeonidis, Nikos Nikolaidis 0001, Anastasios Tefas
ICPR4
2022 Online Knowledge Distillation for Financial Timeseries Forecasting
abstract
Recent advances in Deep Neural Networks (DNNs) led to enormous progress in many different fields, covering a wide range of applications, including financial time-series analysis. Many different financial time-series forecasting tasks have been successfully tackled using such approaches, including predicting the next day’s return of FOREX currency pairs. However, using DNNs for such tasks is not always straightforward due to training stability issues that often arise. Indeed, the noisy nature of the data can often cause considerable different behaviors between DL models, despite following the same training process, model architecture, and hyper-parameters. At the same time, the methods proposed for generating the training labels can sometimes further reinforce such issues. All these phenomena can reduce the reliability of training DL models for financial forecasting tasks, while also making the training process especially time-consuming, requiring several validation and back-testing runs. To overcome these limitations, we propose an ensemble-based online distillation method that can significantly reduce this behavior. The proposed method is efficient since, in contrast to offline distillation approaches, it works in a single step, while it also allows for reducing the number of hyper-parameters to be tuned, e.g., the number of epochs for training the teachers. As demonstrated in the conducted experiments, the soft labels extracted through the proposed approach can mitigate the effect of noisy annotation that often exists in FOREX data, leading to significant performance improvements.
Pavlos Floratos, Avraam Tsantekidis, Nikolaos Passalis, Anastasios Tefas
INISTA4
2022 OpenDR: An Open Toolkit for Enabling High Performance, Low Footprint Deep Learning for Robotics
abstract
Existing Deep Learning (DL) frameworks typically do not provide ready-to-use solutions for robotics, where very specific learning, reasoning, and embodiment problems exist. Their relatively steep learning curve and the different methodologies employed by DL compared to traditional approaches, along with the high complexity of DL models, which often leads to the need of employing specialized hardware accelerators, further increase the effort and cost needed to employ DL models in robotics. Also, most of the existing DL methods follow a static inference paradigm, as inherited by the traditional computer vision pipelines, ignoring active perception, which can be employed to actively interact with the environment in order to increase perception accuracy. In this paper, we present the Open Deep Learning Toolkit for Robotics (OpenDR). OpenDR aims at developing an open, non-proprietary, efficient, and modular toolkit that can be easily used by robotics companies and research institutions to efficiently develop and deploy AI and cognition technologies to robotics applications, providing a solid step towards addressing the aforementioned challenges. We also detail the design choices, along with an abstract interface that was created to overcome these challenges. This interface can describe various robotic tasks, spanning beyond traditional DL cognition and inference, as known by existing frameworks, incorporating openness, homogeneity and robotics-oriented perception e.g., through active perception, as its core design principles.
Nikolaos Passalis, S. Pedrazzi, Robert Babuska, Wolfram Burgard, D. Dias, F. Ferro, Moncef Gabbouj, Ole Green, Alexandros Iosifidis, Erdal Kayacan, Jens Kober, O. Michel, Nikos Nikolaidis 0001, Paraskevi Nousi, Roel Pieters, Maria Tzelepi, Abhinav Valada, Anastasios Tefas
IROS18
2022 Autoencoder-driven spiral representation learning for gravitational wave surrogate modelling
Paraskevi Nousi, Styliani-Christina Fragkouli, Nikolaos Passalis, Panagiotis Iosif, Theocharis Apostolatos, George Pappas, Nikolaos Stergioulas, Anastasios Tefas
Neurocomputing8
2022 Probabilistic online self-distillation
Maria Tzelepi, Nikolaos Passalis, Anastasios Tefas
Neurocomputing3
2022 Multisource financial sentiment analysis for detecting Bitcoin price change indications using deep learning
Nikolaos Passalis, Loukia Avramelou, Solon Seficha, Avraam Tsantekidis, Stavros Doropoulos, Giorgos Makris, Anastasios Tefas
Neural Comput. Appl.7
2022 Quantization-aware training for low precision photonic neural networks
Manos Kirtas, Athina Oikonomou, Nikolaos Passalis, George Mourgias-Alexandris, Miltiadis Moralis-Pegios, Nikos Pleros, Anastasios Tefas
Neural Networks7
2022 Predicting on-street parking violation rate using deep residual neural networks
Nikolaos Karantaglis, Nikolaos Passalis, Anastasios Tefas
Pattern Recognit. Lett.3
2022 Editorial paper for Pattern Recognition Letters VSI on cross model understanding for visual question answering
Shaohua Wan 0001, Zan Gao 0001, Hanwang Zhang, Xiaojun Chang, Chen Chen 0001, Anastasios Tefas
Pattern Recognit. Lett.6
2021 Efficient Realistic Data Generation Framework Leveraging Deep Learning-Based Human Digitization
Charalampos Symeonidis, Paraskevi Nousi, Pavlos Tosidis, Konstantinos Tsampazis, Nikolaos Passalis, Anastasios Tefas, Nikos Nikolaidis 0001
EANN6
2021 Pseudo-Active Vision For Improving Deep Visual Perception Through Neural Sensory Refinement
abstract
Active vision approaches hold the credentials for improving the accuracy of Deep Learning (DL) models for many challenging visual analysis tasks and varying environmental conditions. However, active vision approaches are typically closely tied to the underlying hardware, slowing down their adoption, while they typically increase the latency of perception systems, since sensory data must be recaptured. In this work, we propose a pseudo-active data refinement method that works by appropriately refining the sensory input, without having to reacquire the sensor data through traditional camera control approaches. The proposed method is fully differentiable and can be trained for the task at hand in an end-to-end fashion, while it can be directly deployed in a wide variety of systems, tasks and conditions. The effectiveness and robustness of the proposed method is demonstrated across a variety of tasks using two challenging datasets.
Nikolaos Passalis, Anastasios Tefas
ICIP2
2021 Efficient Training of Lightweight Neural Networks Using Online Self-Acquired Knowledge Distillation
abstract
Knowledge Distillation has been established as a highly promising approach for training compact and faster models by transferring knowledge from heavyweight and powerful models. However, KD in its conventional version constitutes an enduring, computationally and memory demanding process. In this paper, Online Self-Acquired Knowledge Distillation (OSAKD) is proposed, aiming to improve the performance of any deep neural model in an online manner. We utilize k-nn non-parametric density estimation technique for estimating the unknown probability distributions of the data samples in the output feature space. This allows us for directly estimating the posterior class probabilities of the data samples, and we use them as soft labels that encode explicit information about the similarities of the data with the classes, negligibly affecting the computational cost. The experimental evaluation on four datasets validates the effectiveness of proposed method.
Maria Tzelepi, Anastasios Tefas
ICME2
2021 Online Subclass Knowledge Distillation
Maria Tzelepi, Nikolaos Passalis, Anastasios Tefas
Expert Syst. Appl.3
2021 Transferring trading strategy knowledge to deep learning models
Avraam Tsantekidis, Anastasios Tefas
Knowl. Inf. Syst.2
2021 Diversity-driven knowledge distillation for financial trading using Deep Reinforcement Learning
Avraam Tsantekidis, Nikolaos Passalis, Anastasios Tefas
Neural Networks3
2021 Deep adaptive group-based input normalization for financial trading
Angelos Nalmpantis, Nikolaos Passalis, Avraam Tsantekidis, Anastasios Tefas
Pattern Recognit. Lett.4
2021 Improving knowledge distillation using unified ensembles of specialized teachers
Adamantios Zaras, Nikolaos Passalis, Anastasios Tefas
Pattern Recognit. Lett.3
2021 Occlusion detection and drift-avoidance framework for 2D visual object tracking
abstract
This paper presents a long-term 2D tracking framework for the coverage of live outdoor (e.g., sports) events that is suitable for embedded system application (e.g. Unmanned Aerial Vehicles). This application scenario requires 2D target (e.g., athlete, ball, bicycle, boat) tracking for visually assisting the UAV pilot (or cameraman) to maintain proper target framing, or even for actual 3D target following/localization when the drone flies autonomously. In these cases, it should be expected that the target to be tracked/followed, may disappear from the UAV camera field of view, due to fast 3D target motion, illumination changes, or due to visual target occlusions by obstacles, even if the actual UAV continues following it (either autonomously, by exploiting alternative target localization sensors, or by pilot maneuvering). Therefore, the 2D tracker should be able to recover from such situations. The proposed framework solves exactly this problem. Target occlusions are detected from the 2D tracker responses. Depending on the occlusion immensity, the proposed framework decides whether to not update the tracking model, or to employ target re-detection in a broader window. As a result, the proposed framework allows continued target tracking once the target re-appears in the video stream, without tracker re-initialization.
Iason Karakostas, Vasileios Mygdalis, Anastasios Tefas, Ioannis Pitas
Signal Process. Image Commun.3
2021 Deep supervised hashing using quadratic spherical mutual information for efficient image retrieval
Nikolaos Passalis, Anastasios Tefas
Signal Process. Image Commun.2
2021 Hypersphere-Based Weight Imprinting for Few-Shot Learning on Embedded Devices
abstract
Weight imprinting (WI) was recently introduced as a way to perform gradient descent-free few-shot learning. Due to this, WI was almost immediately adapted for performing few-shot learning on embedded neural network accelerators that do not support back-propagation, e.g., edge tensor processing units. However, WI suffers from many limitations, e.g., it cannot handle novel categories with multimodal distributions and special care should be given to avoid overfitting the learned embeddings on the training classes since this can have a devastating effect on classification accuracy (for the novel categories). In this article, we propose a novel hypersphere-based WI approach that is capable of training neural networks in a regularized, imprinting-aware way effectively overcoming the aforementioned limitations. The effectiveness of the proposed method is demonstrated using extensive experiments on three image data sets.
Nikolaos Passalis, Alexandros Iosifidis, Moncef Gabbouj, Anastasios Tefas
IEEE Trans. Neural Networks Learn. Syst.4
2021 Probabilistic Knowledge Transfer for Lightweight Deep Representation Learning
abstract
Knowledge-transfer (KT) methods allow for transferring the knowledge contained in a large deep learning model into a more lightweight and faster model. However, the vast majority of existing KT approaches are designed to handle mainly classification and detection tasks. This limits their performance on other tasks, such as representation/metric learning. To overcome this limitation, a novel probabilistic KT (PKT) method is proposed in this article. PKT is capable of transferring the knowledge into a smaller student model by keeping as much information as possible, as expressed through the teacher model. The ability of the proposed method to use different kernels for estimating the probability distribution of the teacher and student models, along with the different divergence metrics that can be used for transferring the knowledge, allows for easily adapting the proposed method to different applications. PKT outperforms several existing state-of-the-art KT techniques, while it is capable of providing new insights into KT by enabling several novel applications, as it is demonstrated through extensive experiments on several challenging data sets.
Nikolaos Passalis, Maria Tzelepi, Anastasios Tefas
IEEE Trans. Neural Networks Learn. Syst.3
2021 Price Trailing for Financial Trading Using Deep Reinforcement Learning
abstract
Machine learning methods have recently seen a growing number of applications in financial trading. Being able to automatically extract patterns from past price data and consistently apply them in the future has been the focus of many quantitative trading applications. However, developing machine learning-based methods for financial trading is not straightforward, requiring carefully designed targets/rewards, hyperparameter fine-tuning, and so on. Furthermore, most of the existing methods are unable to effectively exploit the information available across various financial instruments. In this article, we propose a deep reinforcement learning-based approach, which ensures that consistent rewards are provided to the trading agent, mitigating the noisy nature of profit-and-loss rewards that are usually used. To this end, we employ a novel price trailing-based reward shaping approach, significantly improving the performance of the agent in terms of profit, Sharpe ratio, and maximum drawdown. Furthermore, we carefully designed a data preprocessing method that allows for training the agent on different FOREX currency pairs, providing a way for developing market-wide RL agents and allowing, at the same time, to exploit more powerful recurrent deep learning models without the risk of overfitting. The ability of the proposed methods to improve various performance metrics is demonstrated using a challenging large-scale data set, containing 28 instruments, provided by Speedlab AG.
Avraam Tsantekidis, Nikolaos Passalis, Anastasia-Sotiria Toufa, Konstantinos Saitas Zarkias, Stergios Chairistanidis, Anastasios Tefas
IEEE Trans. Neural Networks Learn. Syst.6
2020 Heterogeneous Knowledge Distillation Using Information Flow Modeling
abstract
Knowledge Distillation (KD) methods are capable of transferring the knowledge encoded in a large and complex teacher into a smaller and faster student. Early methods were usually limited to transferring the knowledge only between the last layers of the networks, while latter approaches were capable of performing multi-layer KD, further increasing the accuracy of the student. However, despite their improved performance, these methods still suffer from several limitations that restrict both their efficiency and flexibility. First, existing KD methods typically ignore that neural networks undergo through different learning phases during the training process, which often requires different types of supervision for each one. Furthermore, existing multi-layer KD methods are usually unable to effectively handle networks with significantly different architectures (heterogeneous KD). In this paper we propose a novel KD method that works by modeling the information flow through the various layers of the teacher model and then train a student model to mimic this information flow. The proposed method is capable of overcoming the aforementioned limitations by using an appropriate supervision scheme during the different phases of the training process, as well as by designing and training an appropriate auxiliary teacher model that acts as a proxy model capable of “explaining” the way the teacher works to the student. The effectiveness of the proposed method is demonstrated using four image datasets and several different evaluation setups.
Nikolaos Passalis, Maria Tzelepi, Anastasios Tefas
CVPR3
2020 Adaptive Normalization for Forecasting Limit Order Book Data Using Convolutional Neural Networks
abstract
Deep learning models are capable of achieving state-of-the-art performance on a wide range of time series analysis tasks. However, their performance crucially depends on the employed normalization scheme, while they are usually unable to efficiently handle non-stationary features without first appropriately pre-processing them. These limitations impact the performance of deep learning models, especially when used for forecasting financial time series, due to their non-stationary and multimodal nature. In this paper we propose a data-driven adaptive normalization layer which is capable of learning the most appropriate normalization scheme that should be applied on the data. To this end, the proposed method first identifies the distribution from which the data were generated and then it dynamically shifts and scales them in order to facilitate the task at hand. The proposed nor-malization scheme is fully differentiable and it is trained in an end-to-end fashion along with the rest of the parameters of the model. The proposed method leads to significant performance improvements over several competitive normalization approaches, as demonstrated using a large-scale limit order book dataset.
Nikolaos Passalis, Anastasios Tefas, Juho Kanniainen, Moncef Gabbouj, Alexandros Iosifidis
ICASSP2
2020 Leveraging Deep Reinforcement Learning For Active Shooting Under Open-World Setting
abstract
Recent advances in Deep Reinforcement Learning (DRL) led to the development of powerful agents that can learn how to perform complicated tasks in an end-to-end fashion operating directly on raw unstructured data, e.g., images. However, the real world performance of such methods critically relies on the quality of the simulation environments used for training them. The main contribution of this paper is the development of a realistic simulation environment, by employing a state-of-the-art graphics engine, for training DRL agents that are able to control a drone for performing active shooting. In contrast with previous approaches, that solely relied on simplistic constrained datasets, the environment employed in this work supports a challenging open-world setting, providing a solid step towards developing effective RL methods for various drone control tasks. An appropriate reward shaping approach is also introduced in this work, ensuring that the agent will behave as expected, avoiding erratic movements, as demonstrated through the conducted experiments.
A. Tzimas, Nikolaos Passalis, Anastasios Tefas
ICME3
2020 Improving Visual Question Answering using Active Perception on Static Images
abstract
Visual Question Answering (VQA) is one of the most challenging emerging applications of deep learning. Providing powerful attention mechanisms is crucial for VQA, since the model must correctly identify the region of an image that is relevant to the question at hand. However, existing models analyze the input images at a fixed and typically small resolution, often leading to discarding valuable fine-grained details. To overcome this limitation, in this work we propose a reinforcement learning-based active perception approach that works by applying a series of transformation operations on the images (translation, zoom) in order to facilitate answering the question at hand. This allows for performing fine-grained analysis, effectively increasing the resolution at which the models process information. The proposed method is orthogonal to existing attention mechanisms and it can be combined with most existing VQA methods. The effectiveness of the proposed method is experimentally demonstrated on a challenging VQA dataset.
Theodoros Bozinis, Nikolaos Passalis, Anastasios Tefas
ICPR3
2020 A modified Single-Shot multibox Detector for beyond Real-Time Object Detection
abstract
This works focuses on examining the performance of the Single Shot Detector (SSD) model in resource restricted systems where maintaining the power of the full model comprises a significant prerequisite. The proposed SSD variations examine the behavior of lighter versions of SSD while propose measures to limit the unavoidable performance shortage. The outcomes of the conducted research demonstrate a remarkable trade-off between performance losses, speed improvement and the required resource reservation. Thus, the experimental results evidence the efficiency of the presented SSD alterations towards accomplishing higher frame rates and retaining the performance of the original model.
Georgios Orfanidis, Konstantinos Ioannidis, Stefanos Vrochidis, Anastasios Tefas, Ioannis Kompatsiaris
ICPR4
2020 Leveraging Quadratic Spherical Mutual Information Hashing for Fast Image Retrieval
abstract
Several deep supervised hashing techniques have been proposed to allow for querying large image databases. However, it is often overlooked that the process of information retrieval can be modeled using information-theoretic metrics, leading to optimizing various proxies for the problem at hand instead. Contrary to this, we propose a deep supervised hashing algorithm that optimizes the learned codes using an information-theoretic measure, the Quadratic Mutual Information (QMI). The proposed method is adapted to the needs of large-scale hashing and information retrieval leading to a novel information-theoretic measure, the Quadratic Spherical Mutual Information (QSMI), that is inspired by QMI, but leads to significant better retrieval precision. Indeed, the effectiveness of the proposed method is demonstrated under several different scenarios, using different datasets and network architectures, outperforming existing deep supervised image hashing techniques.
Nikolaos Passalis, Anastasios Tefas
ICPR2
2020 Efficient Online Subclass Knowledge Distillation for Image Classification
abstract
Deploying state-of-the-art deep learning models on embedded systems dictates certain storage and computation limitations. During the recent few years Knowledge Distillation (KD) has been recognized as a prominent approach to address this issue. That is, KD has been effectively proposed for training fast and compact deep learning models by transferring knowledge from more complex and powerful models. However, knowledge distillation, in its conventional form, involves multiple stages of training, rendering it a computationally and memory demanding procedure. In this paper, a novel single-stage self knowledge distillation method is proposed, namely Online Subclass Knowledge Distillation (OSKD), that aims at revealing the similarities inside classes, so as to improve the performance of any deep neural model in an online manner. Hence, as opposed to existing online distillation methods, we are able to acquire further knowledge from the model itself, without building multiple identical models or using multiple models to teach each other, rendering the proposed OSKD approach more efficient. The experimental evaluation on two datasets validates that the proposed method improves the classification performance.
Maria Tzelepi, Nikolaos Passalis, Anastasios Tefas
ICPR3
2020 Adaptive Initialization for Recurrent Photonic Networks using Sigmoidal Activations
abstract
Photonic Deep Learning (DL) accelerators are among the most promising approaches for providing fast and energy efficient neural network implementations for several applications. However, photonic accelerators require using different activation functions compared to those typically used in DL. This renders the training process especially difficult to tune, often requiring several trials just for selecting the appropriate initialization hyper-parameters for the network. This process becomes even more difficult for recurrent networks, where exploding gradient phenomena can further destabilize the training process. In this paper, we propose an adaptive data-driven initialization approach for recurrent photonic neural networks. The proposed method is activation-agnostic, while it takes into account the actual distribution of the data used to train the network, overcoming a number of significant limitations of existing approaches. The proposed method is simple and easy to implement, yet it leads to significant improvements in the performance of DL models, as it was experimentally demonstrated using two large-scale challenging time-series datasets.
Nikolaos Passalis, George Mourgias-Alexandris, Nikos Pleros, Anastasios Tefas
ISCAS4
2020 Multilayer Probabilistic Knowledge Transfer for Learning Image Representations
abstract
Probabilistic Knowledge Transfer (PKT) aims to transfer the knowledge encoded in the representations extracted from a layer of a large and complex neural network (teacher) into a smaller and faster one (student). However, PKT only transfers the knowledge between two layers of the networks, ignoring the potentially useful information encoded by the previous ones, reducing in this way the efficiency of PKT and the performance of the student model. In this paper, we propose a novel efficient multilayer PKT method that is capable of transferring the knowledge between the student and teacher networks by employing the representations extracted from multiple layers. The ability of the proposed multilayer PKT method to improve the knowledge transfer and increase the performance of the student model over other state-of-the-art methods is demonstrated using two image datasets.
Nikolaos Passalis, Maria Tzelepi, Anastasios Tefas
ISCAS3
2020 Efficient Adaptive Inference Leveraging Bag-of-Features-based Early Exits
abstract
Early exits provide an effective way of implementing adaptive computational graphs over deep learning models. In this way it is possible to adapt them on-the-fly to the available computational resources or even to the difficulty of each input sample, reducing the energy and computational power requirements in many embedded and mobile applications. However, performing this kind of adaptive inference also comes with several challenges, since the difficulty of each sample must be estimated and the most appropriate early exit must be selected. It is worth noting that existing approaches often lead to highly unbalanced distributions over the selected early exits, reducing the efficiency of the adaptive inference process. At the same time, only a few resources can be devoted to the aforementioned process, in order to ensure that an adequate speedup will be obtained. The main contribution of this work is to provide an easy to use and tune adaptive inference approach for early exits that can overcome some of these limitations. In this way, the proposed method allows for a) obtaining a more balanced inference distribution among the early exits, b) relying on a single and interpretable hyperparameter for tuning its behavior (ranging from faster inference to higher accuracy), and c) improving the performance of the networks (increasing the accuracy and reducing the time needed for inference). Indeed, the effectiveness of the proposed method over existing approaches is demonstrated using four different image datasets.
Nikolaos Passalis, Jenni Raitoharju, Moncef Gabbouj, Anastasios Tefas
MMSP4
2020 Leveraging Active Perception for Improving Embedding-based Deep Face Recognition
abstract
Even though recent advances in deep learning (DL) led to tremendous improvements for various computer and robotic vision tasks, existing DL approaches suffer from a significant limitation: they typically ignore that robots and cyber-physical systems are capable of interacting with the environment in order to better sense their surroundings. In this work we argue that perceiving the world through physical interaction, i.e., employing active perception, allows for both increasing the accuracy of DL models, as well as for deploying smaller and faster models. To this end, we propose an active perception-based face recognition approach, which is capable of simultaneously extracting discriminative embeddings, as well as predicting in which direction the robot must move in order to get a more discriminative view. To the best of our knowledge, we provide the first embedding-based active perception method for deep face recognition. As we experimentally demonstrate, the proposed method can indeed lead to significant improvements, increasing the face recognition accuracy up to 9%, as well as allowing for using overall smaller and faster models, reducing the number of parameters by over one order of magnitude.
Nikolaos Passalis, Anastasios Tefas
MMSP2
2020 Dense convolutional feature histograms for robust visual object tracking
Paraskevi Nousi, Anastasios Tefas, Ioannis Pitas
Image Vis. Comput.2
2020 Continuous drone control using deep reinforcement learning for frontal view person shooting
Nikolaos Passalis, Anastasios Tefas
Neural Comput. Appl.2
2020 K-Anonymity inspired adversarial attack and multiple one-class classification defense
Vasileios Mygdalis, Anastasios Tefas, Ioannis Pitas
Neural Networks2
2020 Initializing photonic feed-forward neural networks using auxiliary tasks
Nikolaos Passalis, George Mourgias-Alexandris, Nikos Pleros, Anastasios Tefas
Neural Networks4
2020 Class-specific discriminant regularization in real-time deep CNN models for binary classification problems
Maria Tzelepi, Anastasios Tefas
Neural Process. Lett.2
2020 Recurrent bag-of-features for visual information analysis
abstract
Deep Learning (DL) has provided powerful tools for visual information analysis. For example, Convolutional Neural Networks (CNNs) are excelling in complex and challenging image analysis tasks by extracting meaningful feature vectors with high discriminative power . However, these powerful feature vectors are crushed through the pooling layers of the network, that usually implement the pooling operation in a less sophisticated manner. This can lead to significant information loss, especially in cases where the informative content of the data is sequentially distributed over the spatial or temporal dimension, e.g., videos, which often require extracting fine-grained temporal information. A novel stateful recurrent pooling approach, that can overcome the aforementioned limitations, is proposed in this paper. The proposed method is inspired by the well-known Bag-of-Features (BoF) model, but employs a stateful trainable recurrent quantizer, instead of plain static quantization, allowing for efficiently processing sequential data and encoding both their temporal, as well as their spatial aspects. The effectiveness of the proposed Recurrent BoF model to enclose spatio-temporal information compared to other competitive methods is demonstrated using six different datasets and two different tasks.
Marios Krestenitis, Nikolaos Passalis, Alexandros Iosifidis, Moncef Gabbouj, Anastasios Tefas
Pattern Recognit.5
2020 Efficient adaptive inference for deep convolutional neural networks using hierarchical early exits
Nikolaos Passalis, Jenni Raitoharju, Anastasios Tefas, Moncef Gabbouj
Pattern Recognit.3
2020 Improving the performance of lightweight CNNs for binary classification using quadratic mutual information regularization
Maria Tzelepi, Anastasios Tefas
Pattern Recognit.2
2020 Variance-preserving deep metric learning for content-based image retrieval
abstract
Supervised deep metric learning led to spectacular results for several Content-based Information Retrieval (CBIR) applications. The success of these approaches slowly led to the belief that image retrieval and classification are just slightly different variations of the same problem. However, recent evidence suggests that learning highly discriminative representation for a (limited) set of training classes removes valuable information from the representation, potentially harming both the in-domain, as well as the out-of-domain retrieval precision. In this paper, we propose a regularized discriminative deep metric learning method that aims to not only learn a representation that allows for discriminating between different classes, but it is also capable of encoding the latent generative factors separately for each class, overcoming this limitation. This allows for modeling the in-class variance and, as a result, maintaining the ability to represent both sub-classes of the in-domain data, as well as objects that belong to classes outside the training domain. The effectiveness of the proposed method, over existing supervised and unsupervised representation/metric learning approaches , is demonstrated under different in-domain and out-of-domain setups and three challenging image datasets.
Nikolaos Passalis, Alexandros Iosifidis, Moncef Gabbouj, Anastasios Tefas
Pattern Recognit. Lett.4
2020 Temporal logistic neural Bag-of-Features for financial time series forecasting leveraging limit order book data
abstract
• Logistic Neural Bag-of-Features are employed for financial time series analysis. • The proposed method can be efficiently used in deep learning architectures. • An adaptive scaling method is proposed to ensure the smooth flow of information. • A logistic kernel is used to estimate the feature vector densities. • The proposed method outperforms the competitive methods on a large-scale dataset. Time series forecasting is a crucial component of many important applications, ranging from forecasting the stock markets to energy load prediction. The high-dimensionality, velocity and variety of the data collected in many of these applications pose significant and unique challenges that must be carefully addressed for each of them. In this work, a novel Temporal Logistic Neural Bag-of-Features approach, that can be used to tackle these challenges, is proposed. The proposed method can be effectively combined with deep neural networks , leading to powerful deep learning models for time series analysis. However, combining existing BoF formulations with deep feature extractors pose significant challenges: the distribution of the input features is not stationary, tuning the hyper-parameters of the model can be especially difficult and the normalizations involved in the BoF model can cause significant instabilities during the training process. The proposed method is capable of overcoming these limitations by a employing a novel adaptive scaling mechanism and replacing the classical Gaussian-based density estimation involved in the regular BoF model with a logistic kernel. The effectiveness of the proposed approach is demonstrated using extensive experiments on a large-scale limit order book dataset that consists of more than 4 million limit orders.
Nikolaos Passalis, Anastasios Tefas, Juho Kanniainen, Moncef Gabbouj, Alexandros Iosifidis
Pattern Recognit. Lett.2
2020 Deep Learning for Visual Content Analysis
Alexandros Iosifidis, Anastasios Tefas
Signal Process. Image Commun.2
2020 Deep autoencoders for attribute preserving face de-identification
Paraskevi Nousi, Sotirios Papadopoulos, Anastasios Tefas, Ioannis Pitas
Signal Process. Image Commun.3
2020 Re-identification framework for long term visual object tracking based on object detection and classification
Paraskevi Nousi, Danai Triantafyllidou, Anastasios Tefas, Ioannis Pitas
Signal Process. Image Commun.3
2020 Bag of Color Features for Color Constancy
abstract
In this paper, we propose a novel color constancy approach, called Bag of Color Features (BoCF), building upon Bag-of-Features pooling. The proposed method substantially reduces the number of parameters needed for illumination estimation. At the same time, the proposed method is consistent with the color constancy assumption stating that global spatial information is not relevant for illumination estimation and local information (edges, etc.) is sufficient. Furthermore, BoCF is consistent with color constancy statistical approaches and can be interpreted as a learning-based extension of many statistical approaches. To further improve the illumination estimation accuracy, we propose a novel attention mechanism for the BoCF model with two variants based on self-attention. BoCF approach and its variants achieve competitive, compared to the state of the art, results while requiring much fewer parameters on three benchmark datasets: ColorChecker RECommended, INTEL-TUT version 2, and NUS8.
Firas Laakom, Nikolaos Passalis, Jenni Raitoharju, Jarno Nikkanen, Anastasios Tefas, Alexandros Iosifidis, Moncef Gabbouj
IEEE Trans. Image Process.5
2020 Deep Adaptive Input Normalization for Time Series Forecasting
abstract
Deep learning (DL) models can be used to tackle time series analysis tasks with great success. However, the performance of DL models can degenerate rapidly if the data are not appropriately normalized. This issue is even more apparent when DL is used for financial time series forecasting tasks, where the nonstationary and multimodal nature of the data pose significant challenges and severely affect the performance of DL models. In this brief, a simple, yet effective, neural layer that is capable of adaptively normalizing the input time series, while taking into account the distribution of the data, is proposed. The proposed layer is trained in an end-to-end fashion using backpropagation and leads to significant performance improvements compared to other evaluated normalization schemes. The proposed method differs from traditional normalization methods since it learns how to perform normalization for a given task instead of using a fixed normalization scheme. At the same time, it can be directly applied to any new time series without requiring retraining. The effectiveness of the proposed method is demonstrated using a large-scale limit order book data set, as well as a load forecasting data set.
Nikolaos Passalis, Anastasios Tefas, Juho Kanniainen, Moncef Gabbouj, Alexandros Iosifidis
IEEE Trans. Neural Networks Learn. Syst.2
2019 Deep Convolutional Feature Histograms for Visual Object Tracking
abstract
Visual Object Tracking remains an open and challenging task in the Computer Vision field, requiring tracking algorithms to achieve a feeble balance between precision and speed performance. In this work, inspired by the classic Mean Shift algorithm for object tracking using histograms, as well as the recent advances of deep Convolutional Neural Networks (CNNs), we propose a novel tracker that incorporates elements from both worlds. Our tracker uses a deep CNN as the feature extraction backbone, which is capable of extracting semantically meaningful features from the target and its background, as well as a fully learnable Bag-of-Features mechanism which extracts histograms from those features. The tracker operates in a fully-convolutional fashion, allowing for the direct and efficient evaluation of multiple possible target locations. Extensive experimental results demonstrate the efficiency and effectiveness of the proposed tracker, allowing it to run at high speeds even on systems with lower computational capacity.
Paraskevi Nousi, Anastasios Tefas, Ioannis Pitas
ICASSP2
2019 Variance Preserving Initialization for Training Deep Neuromorphic Photonic Networks with Sinusoidal Activations
abstract
Photonic neuromorphic hardware can provide significant performance benefits for Deep Learning (DL) applications by accelerating and reducing the energy requirements of DL models. However, photonic neuromorphic architectures employ different activation elements than those traditionally used in DL, slowing down the convergence of the training process for such architectures. An initialization scheme that can be used to efficiently train deep photonic networks that employ quadratic sinusoidal activation functions is proposed in this paper. The proposed initialization scheme can overcome these limitations, leading to faster and more stable training of deep photonic neural networks. The ability of the proposed method to improve the convergence of the training process is experimentally demonstrated using two different DL architectures and two datasets.
Nikolaos Passalis, George Mourgias-Alexandris, Apostolos Tsakyridis, Nikos Pleros, Anastasios Tefas
ICASSP5
2019 Deep Temporal Logistic Bag-of-features for Forecasting High Frequency Limit Order Book Time Series
abstract
Forecasting time series has several applications in various domains. The vast amount of data that are available nowadays provide the opportunity to use powerful deep learning approaches, but at the same time pose significant challenges of high-dimensionality, velocity and variety. In this paper, a novel logistic formulation of the well-known Bag-of-Features model is proposed to tackle these challenges. The proposed method is combined with deep convolutional feature extractors and is capable of accurately modeling the temporal behavior of time series, forming powerful forecasting models that can be trained in an end-to-end fashion. The proposed method was extensively evaluated using a large-scale financial time series dataset, that consists of more than 4 million limit orders, outperforming other competitive methods.
Nikolaos Passalis, Anastasios Tefas, Juho Kanniainen, Moncef Gabbouj, Alexandros Iosifidis
ICASSP2
2019 Deep Reinforcement Learning for Financial Trading Using Price Trailing
abstract
Developing accurate financial analysis tools can be useful both for speculative trading, as well as for analyzing the behavior of markets and promptly responding to unstable conditions ensuring the smooth operation of the financial markets. This led to the development of various methods for analyzing and forecasting the behaviour of financial assets, ranging from traditional quantitative finance to more modern machine learning approaches. However, the volatile and unstable behavior of financial markets forbids the accurate prediction of future prices, reducing the performance of these approaches. In contrast, in this paper we propose a novel price trailing method that goes beyond traditional price forecasting by reformulating trading as a control problem, effectively overcoming the aforementioned limitations. The proposed method leads to developing robust agents that can withstand large amounts of noise, while still capturing the price trends and allowing for taking profitable decisions.
Konstantinos Saitas Zarkias, Nikolaos Passalis, Avraam Tsantekidis, Anastasios Tefas
ICASSP4
2019 Joint Lightweight Object Tracking and Detection for Unmanned Vehicles
abstract
In this paper, we address the problem of lightweight and effective visual object tracking and we present a real-time tracking system suitable for integration in embedded autonomous platforms. We propose a novel tracking framework for classification-based re-detection and tracking, with learnable management of tracking and detection results. The proposed framework includes a novel, very efficient object reidentification method, which filters the detection candidates and systematically corrects the tracking results. In our experiments, we demonstrate the effectiveness of the proposed system by comparing its performance against several other state-of-the art trackers and report the results on the UAV123 and UAV20L datasets. The results indicate that the proposed method is significantly more robust and accurate against recent state-of-the-art trackers, surpassing problems caused by real-world scenarios, while maintaining fast tracking speeds, making it suitable for use in real-time vision applications for autonomous robots, such as Unmanned Aerial Vehicles (UAVs).
Paraskevi Nousi, Danai Triantafyllidou, Anastasios Tefas, Ioannis Pitas
ICIP3
2019 Adaptive Inference Using Hierarchical Convolutional Bag-of-Features for Low-Power Embedded Platforms
abstract
Using early exits provide a straightforward way to implement models that can adapt on-the-fly to the available computational resources. However, early exits in many cases suffer from significant limitations, which often prohibit their practical application, especially when placed on convolutional layers with narrow receptive fields. In this work, we propose a method capable of overcoming these limitations by a) using a Bag-of-Features (BoF)-based pooling approach, that allows for keeping more information regarding the distribution of the extracted feature vectors, while also maintaining more spatial information and b) employing a simple, yet effective, hierarchical approach for designing the exits, allowing for efficiently re-using the information that was already extracted by the previous layers. It is experimentally demonstrated that the proposed approach leads to significant performance improvements, allowing early exits to be a more practical tool that can be used in many real-world embedded applications.
Nikolaos Passalis, Jenni Raitoharju, Anastasios Tefas, Moncef Gabbouj
ICIP3
2019 Class-Based Variational Representation Learning For Robust Image Retrieval
abstract
Supervised learning for Content-based Information Retrieval allows for obtaining discriminative representations that often excel within the training domain. However, recent evidence suggests that these representations can actually harm the retrieval precision for queries that do not belong to the domain of the training set compared to other, less discriminative representations. To avoid this behavior, we propose to learn discriminative representations which also encode the latent generative factors for each class. In this way, the proposed method is capable of maintaining (part of) the in-class variance, as well as being able to represent data that belong to classes that were not seen during the training by better learning the structure of the input space. The proposed method is evaluated under different in-domain and out-of-domain setups, significantly outperforming existing supervised and unsupervised representation learning approaches.
Nikolaos Passalis, Anastasios Tefas, Alexandros Iosifidis, Moncef Gabbouj
ICIP2
2019 Computational UAV Cinematography for Intelligent Shooting Based on Semantic Visual Analysis
abstract
Audiovisual coverage of sports events using Unmanned Aerial Vehicles (UAVs) is becoming increasingly popular. Intelligent audiovisual (A/V) shooting tools, accurately identifying the 2D region of cinematographic attention (RoCA) depicting rapidly moving target ensembles and automatically controlling the UAVs/cameras through visual content analysis, are thus needed. A novel algorithmic pipeline is proposed, implementing computational UAV cinematography for assisting sports coverage, based on semantic, human-centered visual analysis. Athlete and ball detection / tracking results as well as their spatial distribution on the image plane are the semantic features extracted from UAV video feed and exploited for RoCA extraction, based solely on present and past target detections. A PID controller visually controlling a real or virtual camera to track the RoCA and produce aesthetically pleasing shots, without exploiting 3D location-related information, is employed. The proposed method is evaluated on actual UAV footage from soccer matches and promising results are obtained.
Fotini Patrona, Ioannis Mademlis, Anastasios Tefas, Ioannis Pitas
ICIP3
2019 Discriminant Analysis Regularization in Lightweight Deep CNN Models
abstract
In this paper, we first propose lightweight deep CNN models, capable of effectively operating on-drone, in order to address various classification problems, i.e. crowd, football player, and bicycle detection, in the context of media coverage of specific sport events by drones with increased decisional autonomy. Subsequently, we propose a regularization technique, namely Discriminant Analysis regularization, aiming to enhance the generalization ability of the proposed models. The experimental evaluation validates the enhanced performance of the proposed regularizer.
Maria Tzelepi, Anastasios Tefas
ICIP2
2019 Semantic Map Annotation Through UAV Video Analysis Using Deep Learning Models in ROS
Efstratios Kakaletsis, Maria Tzelepi, Pantelis I. Kaplanoglou, Charalampos Symeonidis, Nikos Nikolaidis 0001, Anastasios Tefas, Ioannis Pitas
MMM (2)6
2019 Greedy Salient Dictionary Learning for Activity Video Summarization
Ioannis Mademlis, Anastasios Tefas, Ioannis Pitas
MMM (1)2
2019 Complete vector quantization of feedforward neural networks
Nikolaos Floropoulos, Anastasios Tefas
Neurocomputing2
2019 Deep reinforcement learning for controlling frontal person close-up shooting
Nikolaos Passalis, Anastasios Tefas
Neurocomputing2
2019 Interactive dimensionality reduction using similarity projections
Dimitris Spathis, Nikolaos Passalis, Anastasios Tefas
Knowl. Based Syst.3
2019 Long-term temporal averaging for stochastic optimization of deep neural networks
Nikolaos Passalis, Anastasios Tefas
Neural Comput. Appl.2
2019 Exploiting multiplex data relationships in Support Vector Machines
Vasileios Mygdalis, Anastasios Tefas, Ioannis Pitas
Pattern Recognit.2
2019 Discriminative clustering using regularized subspace learning
Nikolaos Passalis, Anastasios Tefas
Pattern Recognit.2
2019 Visual representation decoding from human brain activity using machine learning: A baseline study
Angeliki Papadimitriou, Nikolaos Passalis, Anastasios Tefas
Pattern Recognit. Lett.3
2019 Unsupervised Knowledge Transfer Using Similarity Embeddings
abstract
With the advent of deep neural networks, there is a growing interest in transferring the knowledge from a large and complex model to a smaller and faster one. In this brief, a method for unsupervised knowledge transfer (KT) between neural networks is proposed. To the best of our knowledge, the proposed method is the first method that utilizes similarity-induced embeddings to transfer the knowledge between any two layers of neural networks, regardless of the number of neurons in each of them. By this way, the knowledge is transferred without using any lossy dimensionality reduction transformations or requiring any information about the complex model, except for the activations of the layer used for KT. This is in contrast with most existing approaches that only generate soft-targets for training the smaller neural network or directly use the weights of the larger model. The proposed method is evaluated using six image data sets and it is demonstrated, through extensive experiments, that the knowledge of a neural network can be successfully transferred using different kinds of (synthetic or not) data, ranging from cross-domain data to just randomly generated data.
Nikolaos Passalis, Anastasios Tefas
IEEE Trans. Neural Networks Learn. Syst.2
2019 Training Lightweight Deep Convolutional Neural Networks Using Bag-of-Features Pooling
abstract
Convolutional neural networks (CNNs) are predominantly used for several challenging computer vision tasks achieving state-of-the-art performance. However, CNNs are complex models that require the use of powerful hardware, both for training and deploying them. To this end, a quantization-based pooling method is proposed in this paper. The proposed method is inspired from the bag-of-features model and can be used for learning more lightweight deep neural networks. Trainable radial basis function neurons are used to quantize the activations of the final convolutional layer, reducing the number of parameters in the network and allowing for natively classifying images of various sizes. The proposed method employs differentiable quantization and aggregation layers leading to an end-to-end trainable CNN architecture. Furthermore, a fast linear variant of the proposed method is introduced and discussed, providing new insight for understanding convolutional neural architectures. The ability of the proposed method to reduce the size of CNNs and increase the performance over other competitive methods is demonstrated using seven data sets and three different learning tasks (classification, regression, and retrieval).
Nikolaos Passalis, Anastasios Tefas
IEEE Trans. Neural Networks Learn. Syst.2
2019 Neurons With Paraboloid Decision Boundaries for Improved Neural Network Classification Performance
abstract
In mathematical terms, an artificial neuron computes the inner product of a d-dimensional input vector x with its weight vector w, compares it with a bias value w0and fires based on the result of this comparison. Therefore, its decision boundary is given by the equation wTx + w0= 0. In this paper, we propose replacing the linear hyperplane decision boundary of a neuron with a curved, paraboloid decision boundary. Thus, the decision boundary of the proposed paraboloid neuron is given by the equation (hTx + h0)2- ||x - p||22= 0, where h and h0denote the parameters of the directrix and p denotes the coordinates of the focus. Such paraboloid neural networks are proven to have superior recognition accuracy in a number of applications.
Nikolaos Tsapanos, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas
IEEE Trans. Neural Networks Learn. Syst.2
2018 Learning Deep Representations with Probabilistic Knowledge Transfer
Nikolaos Passalis, Anastasios Tefas
ECCV (11)2
2018 Regularized Svd-Based Video Frame Saliency for Unsupervised Activity Video Summarization
abstract
Storage, browsing and analysis of human activity videos can be significantly facilitated by automated video summarization. Unsupervised key-frame extraction remains the most widely applicable technique for summarizing activity videos. However, their specific properties make the problem difficult to solve. Typical relevant algorithms fall under the video frame clustering or the dictionary-of-representatives families, with salient dictionary learning having been recently proposed. Under this formulation, the video frames selected as key-frames are the ones which simultaneously best reconstruct the entire video and are salient compared to the rest. This paper improves upon such a method by replacing the video frame saliency estimation term with one based on Regularized SVD-based Low Rank Approximation, taking advantage of the well-established correlation between midrange matrix singular values and salient regions. Extensive empirical evaluation showcases the high performance of both the salient dictionary learning framework and the specific proposed method.
Ioannis Mademlis, Anastasios Tefas, Ioannis Pitas
ICASSP2
2018 Label Propagation on Facial Images Using Similarity and Dissimilarity Labelling Constraints
abstract
In this paper, a novel multimedia data (specifically facial images) label propagation method is presented that is based on the inclusion of labelling constraints in the objective function of the MLPP-CLP state of the art algorithm. The proposed method can incorporate pairwise facial image similarity and dissimilarity constraints into the objective function of the aforementioned method. Experiments which have been conducted on facial image labelling in three stereoscopic movies, confirm the increased labelling accuracy of the proposed method.
Efstratios Kakaletsis, Olga Zoidi, Ioannis Tsingalis, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas
ICIP4
2018 Learning Multi-Graph Regularization for SVM Classification
abstract
A classification method that emphasizes on learning the hyperplane that separates the training data with the maximum margin in a regularized space, is presented. In the proposed method, this regularized space is derived by exploiting multiple graph structures, in the SVM optimization process. Each of the employed graph structure carries some information concerning a geometric or semantic property about the training data, e.g., local neighborhood area and global geometric data relationships. The proposed method introduces information from each graph type to the standard SVM objective, as a projection of the SVM hyperplane to such a direction, where a specific property of the training data is highlighted. We show that each data property can be encoded in a regularized kernel matrix. Finally, response in the optimal classification space can be obtained by exploiting a weighted combination of multiple regularized kernel matrices. Experimental results in face recognition and object classification denote the effectiveness of the proposed method.
Vasileios Mygdalis, Anastasios Tefas, Ioannis Pitas
ICIP2
2018 Convolutional Neural Networks for Visual Information Analysis with Limited Computing Resources
abstract
Over the past decade, Deep Convolutional Neural Networks with heavy architectures and large numbers of parameters have achieved state-of-the-art results and eclipsed other methods in multiple visual analysis tasks, including object detection. However, the real-time requirements of such tasks directly conflict with the restricted computational capabilities of embedded systems, prohibiting the immediate deployment of bulky models, and necessitating their optimization for inference. Parameter pruning techniques reduce the number of parameters while reducing the input size leads to smaller internal representations, leading by extension to fewer computational operations. Furthermore, inference optimization schemes provided by Deep Learning frameworks can yield significant speed ups, for example by allowing half-precision floating point operations. We investigate the behavior of various model configurations in object detection tasks and perform a comparative study on inference optimization methods which aim to reduce the computational cost of Convolutional Neural Networks, while examining the effect of such methods on their performance, and propose architecture modifications for this purpose.
Paraskevi Nousi, Emmanouil Patsiouras, Anastasios Tefas, Ioannis Pitas
ICIP3
2018 Accelerating Similarity-Based Discriminant Analysis Using Class-Specific Prototypes
abstract
Reducing the dimensionality of data is among the most common preprocessing steps for a number of machine learning tasks. A linear similarity-based discriminant analysis (S-LDA) algorithm was recently proposed, significantly improving the discriminative ability of the learned low-dimensional representation over classical techniques, such as LDA. S-LDA builds upon the notion of similarity and allows for exploiting higher-order statistics, while being notably robust to outliers. However, optimizing S-LDA`s objective requires the calculation of the pairwise similarities between all the training data. This quickly becomes a significant burden and limits the ability of the method to scale to larger datasets. In this paper we propose using prototypes, that act as similarity anchors, to reduce the number of required similarity calculations. We experimentally demonstrate that the proposed technique is capable of reducing the optimization time by two orders of magnitude without harming the discriminative ability of the learned representations.
Nikolaos Passalis, Anastasios Tefas
ICIP2
2018 Neural Network Knowledge Transfer using Unsupervised Similarity Matching
abstract
Transferring the knowledge from a large and complex neural network to a smaller and faster one allows for deploying more lightweight and accurate networks. In this paper, we propose a novel method that is capable of transferring the knowledge between any two layers of two neural networks by matching the similarity between the extracted representations. The proposed method is model-agnostic overcoming several limitations of existing knowledge transfer techniques, since the knowledge is transferred between layers that can have different architecture and no information about the complex model is required, apart from the output of the layers employed for the knowledge transfer. Three image datasets are used to demonstrate the effectiveness of the proposed approach, including a large-scale dataset for learning a light-weight model for facial pose estimation that can be directly deployed on devices with limited computational resources, such as embedded systems for drones.
Nikolaos Passalis, Anastasios Tefas
ICPR2
2018 Visual Question Answering using Explicit Visual Attention
abstract
One of the most complex multi-model problems faced today is Visual Question Answering (VQA), which requires a machine to properly understand a question about a reference visual input, expressed in natural language, and then produce the answer to that question. In order to solve this problem and increase the probability of producing the correct answer, it is crucial to provide reliable attention information. However, existing methods only use implicitly trained attention models that are often unable to attend to the appropriate image region the question refers to, limiting their ability to provide the correct answer. To address this issue, we propose an explicitly trained attention model that is inspired by the theory of pictorial superiority effect. In this model, we use attention-oriented word embeddings that increase the efficiency of learning common representation spaces. The dataset that we use, the Visual7W dataset, is the only dataset that provides visual attention ground truth information. In this paper, we demonstrate the effectiveness of the proposed method over both implicit attention models and other state-of-art VQA techniques.
Vasileios Lioutas, Nikolaos Passalis, Anastasios Tefas
ISCAS3
2018 Efficient Camera Control using 2D Visual Information for Unmanned Aerial Vehicle-based Cinematography
abstract
Using Unmanned Aerial Vehicles (UAVs), also known as drones, for covering public sport events, such as bicycle races, is becoming increasingly popular. Even though the problem of controlling the flight path of a drone is well studied in the literature, little work has been done on controlling the shooting camera for producing professional grade video footage. In this work we propose a fast and efficient proportional-integral-derivative (PID) based control algorithm that rely solely on 2D visual information and we demonstrate that it is possible to accurately control the camera without inferring the 3D position of the target. To ensure that the proposed method will not exhibit undesired behavior, a genetic algorithm is used to tune its parameters using a properly defined fitness function. The proposed method is evaluated using two datasets that contain actual drone footage: a dataset that contains videos of a single cyclist, and a dataset that contains actually footage from a bicycle race event, the Giro D'Italia bicycle race.
Nikolaos Passalis, Anastasios Tefas, Ioannis Pitas
ISCAS2
2018 Semi-supervised subclass support vector data description for image and video classification
Vasileios Mygdalis, Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
Neurocomputing3
2018 Deep convolutional learning for Content Based Image Retrieval
Maria Tzelepi, Anastasios Tefas
Neurocomputing2
2018 A salient dictionary learning framework for activity video summarization via key-frame extraction
Ioannis Mademlis, Anastasios Tefas, Ioannis Pitas
Inf. Sci.2
2018 PySEF: A python library for similarity-based dimensionality reduction
Nikolaos Passalis, Anastasios Tefas
Knowl. Based Syst.2
2018 Caricature generation utilizing the notion of anti-face
Vlasis Gogousis, Anastasios Tefas
Multim. Tools Appl.2
2018 Exploiting tf-idf in deep Convolutional Neural Networks for Content Based Image Retrieval
Nikolaos Kondylidis, Maria Tzelepi, Anastasios Tefas
Multim. Tools Appl.3
2018 Learning bag-of-embedded-words representations for textual information retrieval
Nikolaos Passalis, Anastasios Tefas
Pattern Recognit.2
2018 Learning deep spatiotemporal features for video captioning
Eleftherios Daskalakis, Maria Tzelepi, Anastasios Tefas
Pattern Recognit. Lett.3
2018 Explicit ensemble attention learning for improving visual question answering
Vasileios Lioutas, Nikolaos Passalis, Anastasios Tefas
Pattern Recognit. Lett.3
2018 Fast constrained person identity label propagation in stereo videos using a pruned similarity matrix
Efstratios Kakaletsis, Olga Zoidi, Ioannis Tsingalis, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas
Signal Process. Image Commun.4
2018 Deep convolutional image retrieval: A general framework
Maria Tzelepi, Anastasios Tefas
Signal Process. Image Commun.2
2018 Positive and Negative Label Propagations
abstract
This paper extends the state-of-the-art label propagation (LP) framework in the propagation of negative labels. More specifically, the state-of-the-art LP methods propagate information of the form “the sample i should be assigned the label k.” The proposed method extends the state-of-theart framework by considering additional information of the form “the sample i should not be assigned the label k.” A theoretical analysis is presented in order to include negative LP in the problem formulation. Moreover, a method for selecting the negative labels in cases when they are not inherent from the data structure is presented. Furthermore, the incorporation of negative label information in two multigraph LP methods is presented. Finally, a discussion on the proposed algorithm extension to out of sample data, as well as scalability issues, is presented. Experimental results in various scenarios showed that the incorporation of negative label information increases, in all cases, the classification accuracy of the state of the art.
Olga Zoidi, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas
IEEE Trans. Circuits Syst. Video Technol.2
2018 Information Clustering Using Manifold-Based Optimization of the Bag-of-Features Representation
abstract
In this paper, a manifold-based dictionary learning method for the bag-of-features (BoF) representation optimized toward information clustering is proposed. First, the spectral representation, which unwraps the manifolds of the data and provides better clustering solutions, is formed. Then, a new dictionary is learned in order to make the histogram space, i.e., the space where the BoF historgrams exist, as similar as possible to the spectral space. The ability of the proposed method to improve the clustering solutions is demonstrated using a wide range of datasets: two image datasets, the 15-scene dataset and the Corel image dataset, one video dataset, the KTH dataset, and one text dataset, the RT-2k dataset. The proposed method improves both the internal and the external clustering criteria for two different clustering algorithms: 1) the -means and 2) the spectral clustering. Also, the optimized histogram space can be used to directly assign a new object to its cluster, instead of using the spectral space (which requires reapplying the spectral clustering algorithm or using incremental spectral clustering techniques). Finally, the learned representation is also evaluated using an information retrieval setup and it is demonstrated that improves the retrieval precision over the baseline BoF representation.
Nikolaos Passalis, Anastasios Tefas
IEEE Trans. Cybern.2
2018 Dimensionality Reduction Using Similarity-Induced Embeddings
abstract
The vast majority of dimensionality reduction (DR) techniques rely on the second-order statistics to define their optimization objective. Even though this provides adequate results in most cases, it comes with several shortcomings. The methods require carefully designed regularizers and they are usually prone to outliers. In this paper, a new DR framework that can directly model the target distribution using the notion of similarity instead of distance is introduced. The proposed framework, called similarity embedding framework (SEF), can overcome the aforementioned limitations and provides a conceptually simpler way to express optimization targets similar to existing DR techniques. Deriving a new DR technique using the SEF becomes simply a matter of choosing an appropriate target similarity matrix. A variety of classical tasks, such as performing supervised DR and providing out-of-sample extensions, as well as, new novel techniques, such as providing fast linear embeddings for complex techniques, are demonstrated in this paper using the proposed framework. Six data sets from a diverse range of domains are used to evaluate the proposed method and it is demonstrated that it can outperform many existing DR techniques.
Nikolaos Passalis, Anastasios Tefas
IEEE Trans. Neural Networks Learn. Syst.2
2017 Discriminatively Trained Autoencoders for Fast and Accurate Face Recognition
Paraskevi Nousi, Anastasios Tefas
EANN2
2017 Improving Face Pose Estimation Using Long-Term Temporal Averaging for Stochastic Optimization
Nikolaos Passalis, Anastasios Tefas
EANN2
2017 Summarization of human activity videos via low-rank approximation
abstract
Summarization of videos depicting human activities is a timely problem with important applications, e.g., in the domains of surveillance or film/TV production, that steadily becomes more relevant. Research on video summarization has mainly relied on global clustering or local (frame-by-frame) saliency methods to provide automated algorithmic solutions for key-frame extraction. This work presents a method based on selecting as key-frames video frames able to optimally reconstruct the entire video. The novelty lies in modelling the reconstruction algebraically as a Column Subset Selection Problem (CSSP), resulting in extracting key-frames that correspond to elementary visual building blocks. The problem is formulated under an optimization framework and approximately solved via a genetic algorithm. The proposed video summarization method is being evaluated using a publicly available annotated dataset and an objective evaluation metric. According to the quantitative results, it clearly outperforms the typical clustering approach.
Ioannis Mademlis, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas
ICASSP2
2017 Learning Bag-of-Features Pooling for Deep Convolutional Neural Networks
abstract
Convolutional Neural Networks (CNNs) are well established models capable of achieving state-of-the-art classification accuracy for various computer vision tasks. However, they are becoming increasingly larger, using millions of parameters, while they are restricted to handling images of fixed size. In this paper, a quantization-based approach, inspired from the well-known Bag-of-Features model, is proposed to overcome these limitations. The proposed approach, called Convolutional BoF (CBoF), uses RBF neurons to quantize the information extracted from the convolutional layers and it is able to natively classify images of various sizes as well as to significantly reduce the number of parameters in the network. In contrast to other global pooling operators and CNN compression techniques the proposed method utilizes a trainable pooling layer that it is end-to-end differentiable, allowing the network to be trained using regular back-propagation and to achieve greater distribution shift invariance than competitive methods. The ability of the proposed method to reduce the parameters of the network and increase the classification accuracy over other state-of-the-art techniques is demonstrated using three image datasets.
Nikolaos Passalis, Anastasios Tefas
ICCV2
2017 Summarization of human activity videos using a salient dictionary
abstract
Video summarization has become more prominent during the last decade, due to the massive amount of available digital video content. A video summarization algorithm is typically fed an input video and expected to extract a set of important key-frames which represent the entire content, convey semantic meaning and are significantly more concise than the original input. The most wide-spread approach relies on video frame clustering and extraction of the frames closest to the cluster centroids as key-frames. Such a process, although efficient, offloads the burden of semantic scene content modelling exclusively to the employed video frame description/representation scheme, while summarization itself is approached simply as a distance-based data partitioning problem. This work focuses on videos depicting human activities (e.g., from surveillance feeds) which display an attractive property, i.e., each video frame can be seen as a linear combination of elementary visual words (i.e., basic activity components). This is exploited so as to identify the video frames containing only the elementary visual building blocks, which ideally form a set of independent basis vectors that can linearly reconstruct the entire video. In this manner, the semantic content of the scene is considered by the video summarization process itself. The above process is modulated by a traditional distance-based video frame saliency estimation, biasing towards more spread content coverage and outlier inclusion, under a joint optimization framework derived from the Column Subset Selection Problem (CSSP). The proposed algorithm results in a final key-frame set which acts as as salient dictionary for the input video. Empirical evaluation conducted on a publicly available dataset suggest that the presented method outperforms both a baseline clustering-based approach and a state-of-the-art sparse dictionary learning-based algorithm.
Ioannis Mademlis, Anastasios Tefas, Ioannis Pitas
ICIP2
2017 Approximate kernel extreme learning machine for large scale data classification
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
Neurocomputing2
2017 Deep learning algorithms for discriminant autoencoding
Paraskevi Nousi, Anastasios Tefas
Neurocomputing2
2017 De-identifying facial images using singular value decomposition and projections
Panteleimon Chriskos, Olga Zoidi, Anastasios Tefas, Ioannis Pitas
Multim. Tools Appl.3
2017 Multimodal speaker clustering in full length movies
Ioannis Kapsouras, Anastasios Tefas, Nikos Nikolaidis 0001, Geoffroy Peeters, Elie-Laurent Benaroya, Ioannis Pitas
Multim. Tools Appl.2
2017 One-Class Classification Based on Extreme Learning and Geometric Class Information
Alexandros Iosifidis, Vasileios Mygdalis, Anastasios Tefas, Ioannis Pitas
Neural Process. Lett.3
2017 Neural Bag-of-Features learning
Nikolaos Passalis, Anastasios Tefas
Pattern Recognit.2
2017 Big Media Data Analysis
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas, Moncef Gabbouj
Signal Process. Image Commun.2
2017 Learning Neural Bag-of-Features for Large-Scale Image Retrieval
abstract
In this paper, the well-known bag-of-features (BoFs) model is generalized and formulated as a neural network that is composed of three layers: 1) a radial basis function (RBF) layer; 2) an accumulation layer; and 3) a fully connected layer. This formulation allows for decoupling the representation size from the number of used codewords, as well as for better modeling the feature distribution using a separate trainable scaling parameter for each RBF neuron. The resulting network, called retrieval-oriented neural BoF (RN-BoF), is trained using regular back propagation and allows for fast extraction of compact image representations. It is demonstrated that the RN-BoF model is capable of: 1) increasing the object encoding and retrieval speed; 2) reducing the extracted representation size; and 3) increasing the retrieval precision. A symmetry-aware spatial segmentation technique is also proposed to further reduce the encoding time and the storage requirements and allows the method to efficiently scale to large datasets. The proposed method is evaluated and compared to other state-of-the-art techniques using five different image datasets, including the large-scale YouTube Faces database.
Nikolaos Passalis, Anastasios Tefas
IEEE Trans. Syst. Man Cybern. Syst.2
2016 One class classification applied in facial image analysis
abstract
In this paper, we apply One-Class Classification methods in facial image analysis problems. We consider the cases where the available training data information originates from one class, or one of the available classes is of high importance. We propose a novel extension of the One-Class Extreme Learning Machines algorithm aiming at minimizing both the training error and the data dispersion and consider solutions that generate decision functions in the ELM space, as well as in ELM spaces of arbitrary dimensionality. We evaluate the performance in publicly available datasets. The proposed method compares favourably to other state-of-the-art choices.
Vasileios Mygdalis, Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
ICIP3
2016 Multi-view semantic temporal video segmentation
abstract
In this work, we propose a multi-view temporal video segmentation approach that employs a Gaussian scoring process for determining the best segmentation positions. By exploiting the semantic action information that the dense trajectories video description offers, this method can detect intra-shot actions as well, unlike shot boundary detection approaches. We compare the temporal segmentation results of the proposed method to both single-view and multi-view methods, and also compare the action recognition results obtained on ground truth video segments to the ones obtained on the proposed multi-view segments, on the IMPART multi-view action data set.
Thomas Theodoridis, Anastasios Tefas, Ioannis Pitas
ICIP2
2016 Exploiting local and global geometric data relationships in Support Vector Data Description
abstract
In this paper, we describe a one-class classification method based on Support Vector Data Description, which exploits multiple graph structures in its optimization process. We derive in a generic solution which can be employed for supervised one-class classification tasks. The devised method can produce linear or non-linear decision functions, depending on the adopted kernel function. In our experiments, we simultaneously adopted two graphs that describe local and global geometric training data relationships, respectively. We evaluated the proposed classifier in publicly available datasets, where its performance compared favorably against closely related methods.
Vasileios Mygdalis, Anastasios Tefas, Ioannis Pitas
ICPR2
2016 Bag of Embedded Words learning for text retrieval
abstract
The word embedding models are capable of capturing the semantic content of the textual words. The process of extracting a set of word embedding vectors from a text document is similar to the feature extraction step of the Bag-of-Features pipeline, which is usually used in computer vision tasks. That gives rise to the Bag-of-Embedded Words (BoEW) model. In this paper a novel learning technique that optimizes both the word embedding and the codebook of the BoEW model towards text retrieval is proposed. The proposed method adheres to the cluster hypothesis that states that points in the same cluster are likely to fulfill the same information need and it is demonstrated, using two text datasets, that can significantly increase the retrieval precision. Finally, the proposed technique uses smaller representations than the competitive representation methods, that allows to reduce both the retrieval time and the storage requirements.
Nikolaos Passalis, Anastasios Tefas
ICPR2
2016 Face detection based on deep convolutional neural networks exploiting incremental facial part learning
abstract
Deep learning methods are powerful approaches but often require expensive computations and lead to models of high complexity which need to be trained with large amounts of data. In this paper, we consider the problem of face detection and we propose a light-weight deep convolutional neural network that achieves a state-of-the-art recall rate at the challenging FDDB dataset. Our model is designed with a view to minimize both training and run time and outperforms the convolutional network used in [1] for the same task. Our model consists only of 113.864 free parameters whereas the previously proposed CNN for face detection had 60 million parameters. We propose a new training method that gradually increases the difficulty of both negative and positive examples and has proved to drastically improve training speed and accuracy. Our second approach, involves training a separate deep network to detect individual facial features whilst creating a model that combines the outputs of two different networks. Both methods are able to detect faces under severe occlusion and unconstrained pose variation and meet the difficulties and the large variations of real-world face detection.
Danai Triantafyllidou, Anastasios Tefas
ICPR2
2016 Exploiting supervised learning for finetuning deep CNNs in content based image retrieval
abstract
In this paper a novel CNN-based approach in the Content Based Image Retrieval domain that exploits supervised learning is proposed. We employ a deep CNN model to derive feature representations from the activations of the deepest layers and we refine the weights of the utilized layers in order to produce better image descriptors using information obtained from the available data labels. To this end, we adapt the pretrained model and we retrain it on the dataset so that each image representation comes closer in terms of Euclidean distance to its nearest relevant representations and moves away from the irrelevant ones. Experimental results on four publicly available datasets for image retrieval denote the effectiveness of the proposed method in enhancing the retrieval performance, outperforming other CNN-based retrieval techniques in three out of four datasets, as well as traditional handcrafted approaches.
Maria Tzelepi, Anastasios Tefas
ICPR2
2016 Movie shot selection preserving narrative properties
abstract
Automatic shot selection is an important aspect of movie summarization that is helpful both to producers and to audiences, e.g., for market promotion or browsing purposes. However, most of the related research has focused on shot selection based on low-level video content, which disregards semantic information, or on narrative properties extracted from text, which requires the movie script to be available. In this work, semantic shot selection based on the narrative prominence of movie characters in both the visual and the audio modalities is investigated, without the need for additional data such as a script. The output is a movie summary that only contains video frames from selected movie shots. Selection is controlled by a user-provided shot retention parameter, that removes key-frames/key-segments from the skim based on actor face appearances and speech instances. This novel process (Multimodal Shot Pruning, or MSP) is algebraically modelled as a multimodal matrix Column Subset Selection Problem, which is solved using an evolutionary computing approach.
Ioannis Mademlis, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas
MMSP2
2016 Laplacian one class extreme learning machines for human action recognition
abstract
A novel OCC method for human action recognition namely the Laplacian One Class Extreme Learning Machines is presented. The proposed method exploits local geometric data information within the OC-ELM optimization process. It is shown that emphasizing on preserving the local geometry of the data leads to a regularized solution, which models the target class more efficiently than the standard OC-ELM algorithm. The proposed method is extended to operate in feature spaces determined by the network hidden layer outputs, as well as in ELM spaces of arbitrary dimensions. Its superior performance against other OCC options is consistent among five publicly available human action recognition datasets.
Vasileios Mygdalis, Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
MMSP3
2016 Exploiting stereoscopic disparity for augmenting human activity recognition performance
Ioannis Mademlis, Alexandros Iosifidis, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas
Multim. Tools Appl.3
2016 Big Data Analysis for Media Production
abstract
A typical high-end film production generates several terabytes of data per day, either as footage from multiple cameras or as background information regarding the set (laser scans, spherical captures, etc). This paper presents solutions to improve the integration of the multiple data sources, and understand their quality and content, which are useful both to support creative decisions on-set (or near it) and enhance the postproduction process. The main cinema specific contributions, tested on a multisource production dataset made publicly available for research purposes, are the monitoring and quality assurance of multicamera set-ups, multisource registration and acceleration of 3-D reconstruction, anthropocentric visual analysis techniques for semantic content annotation, and integrated 2-D–3-D web visualization tools. We discuss as well improvements carried out in basic techniques for acceleration, clustering and visualization, which were necessary to deal with the very large multisource data, and can be applied to other big data problems in diverse application fields.
Josep Blat, Alun Evans, Hansung Kim 0001, Evren Imre, Lukás Polok, Viorela Ila, Nikos Nikolaidis 0001, Pavel Zemcík, Anastasios Tefas, Pavel Smrz, Adrian Hilton 0001, Ioannis Pitas
Proc. IEEE9
2016 Graph Embedded One-Class Classifiers for media data classification
Vasileios Mygdalis, Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
Pattern Recognit.3
2016 Graph Embedded Extreme Learning Machine
abstract
In this paper, we propose a novel extension of the extreme learning machine (ELM) algorithm for single-hidden layer feedforward neural network training that is able to incorporate subspace learning (SL) criteria on the optimization process followed for the calculation of the network's output weights. The proposed graph embedded ELM (GEELM) algorithm is able to naturally exploit both intrinsic and penalty SL criteria that have been (or will be) designed under the graph embedding framework. In addition, we extend the proposed GEELM algorithm in order to be able to exploit SL criteria in arbitrary (even infinite) dimensional ELM spaces. We evaluate the proposed approach on eight standard classification problems and nine publicly available datasets designed for three problems related to human behavior analysis, i.e., the recognition of human face, facial expression, and activity. Experimental results denote the effectiveness of the proposed approach, since it outperforms other ELM-based classification schemes in all the cases.
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
IEEE Trans. Cybern.2
2016 Multimodal Stereoscopic Movie Summarization Conforming to Narrative Characteristics
abstract
Video summarization is a timely and rapidly developing research field with broad commercial interest, due to the increasing availability of massive video data. Relevant algorithms face the challenge of needing to achieve a careful balance between summary compactness, enjoyability, and content coverage. The specific case of stereoscopic 3D theatrical films has become more important over the past years, but not received corresponding research attention. In this paper, a multi-stage, multimodal summarization process for such stereoscopic movies is proposed, that is able to extract a short, representative video skim conforming to narrative characteristics from a 3D film. At the initial stage, a novel, low-level video frame description method is introduced (frame moments descriptor) that compactly captures informative image statistics from luminance, color, optical flow, and stereoscopic disparity video data, both in a global and in a local scale. Thus, scene texture, illumination, motion, and geometry properties may succinctly be contained within a single frame feature descriptor, which can subsequently be employed as a building block in any key-frame extraction scheme, e.g., for intra-shot frame clustering. The computed key-frames are then used to construct a movie summary in the form of a video skim, which is post-processed in a manner that also considers the audio modality. The next stage of the proposed summarization pipeline essentially performs shot pruning, controlled by a user-provided shot retention parameter, that removes segments from the skim based on the narrative prominence of movie characters in both the visual and the audio modalities. This novel process (multimodal shot pruning) is algebraically modeled as a multimodal matrix column subset selection problem, which is solved using an evolutionary computing approach. Subsequently, disorienting editing effects induced by summarization are dealt with, through manipulation of the video skim. At the last step, the skim is suitably post-processed in order to reduce stereoscopic video defects that may cause visual fatigue.
Ioannis Mademlis, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas
IEEE Trans. Image Process.2
2016 Entropy Optimized Feature-Based Bag-of-Words Representation for Information Retrieval
abstract
In this paper, we present a supervised dictionary learning method for optimizing the feature-based Bag-of-Words (BoW) representation towards Information Retrieval. Following the cluster hypothesis, which states that points in the same cluster are likely to fulfill the same information need, we propose the use of an entropy-based optimization criterion that is better suited for retrieval instead of classification. We demonstrate the ability of the proposed method, abbreviated as EO-BoW, to improve the retrieval performance by providing extensive experiments on two multi-class image datasets. The BoW model can be applied to other domains as well, so we also evaluate our approach using a collection of 45 time-series datasets, a text dataset, and a video dataset. The gains are three-fold since the EO-BoW can improve the mean Average Precision, while reducing the encoding time and the database storage requirements. Finally, we provide evidence that the EO-BoW maintains its representation ability even when used to retrieve objects from classes that were not seen during the training.
Nikolaos Passalis, Anastasios Tefas
IEEE Trans. Knowl. Data Eng.2
2016 Visual Voice Activity Detection in the Wild
abstract
The visual voice activity detection (V-VAD) problem in unconstrained environments is investigated in this paper. A novel method for V-VAD in the wild, exploiting local shape and motion information appearing at spatiotemporal locations of interest for facial video segment description and the bag of words model for facial video segment representation, is proposed. Facial video segment classification is subsequently performed using the state-of-the-art classification algorithms. Experimental results on one publicly available V-VAD dataset denote the effectiveness of the proposed method, since it achieves better generalization performance in unseen users, when compared to the recently proposed state-of-the-art methods. Additional results on a new unconstrained dataset provide evidence that the proposed method can be effective even in such cases in which any other existing method fails.
Fotini Patrona, Alexandros Iosifidis, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas
IEEE Trans. Multim.3
2015 Enhancing class discrimination in Kernel Discriminant Analysis
abstract
In this paper, we propose an optimization scheme aiming at optimal nonlinear data projection, in terms of Fisher ratio maximization. To this end, we formulate an iterative optimization scheme consisting of two processing steps: optimal data projection calculation and optimal class representation determination. Compared to the standard approach employing the class mean vectors for class representation, the proposed optimization scheme increases class discrimination in the reduced-dimensionality feature space. We evaluate the proposed method in standard classification problems, as well as on the classification of human actions and face, and show that it is able to achieve better generalization performance, when compared to the standard approach.
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
ICASSP2
2015 Exploiting subclass information in one-class support vector machine for video summarization
abstract
In this paper, we propose a method for video summarization based on human activity description. We formulate this problem as the one of automatic video segment selection based on a learning process that employs salient video segment paradigms. For this one-class classification problem, we introduce a novel variant of the One-Class Support Vector Machine (OC-SVM) classifier that exploits subclass information in the OC-SVM optimization problem, in order to jointly minimize the data dispersion within each subclass and determine the optimal decision function. We evaluate the proposed approach in three Hollywood movies, where the performance of the proposed SOC-SVM algorithm is compared with that of the OC-SVM. Experimental results denote that the proposed approach is able to outperform OC-SVM-based video segment selection.
Vasileios Mygdalis, Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
ICASSP3
2015 Multimodal Speaker Diarization Utilizing Face Clustering Information
Ioannis Kapsouras, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas
ICIG (2)2
2015 Merging linear discriminant analysis with Bag of Words model for human action recognition
abstract
In this paper we propose a novel method for human action recognition, that unifies discriminative Bag of Words (BoW)-based video representation and discriminant subspace learning. An iterative optimization scheme is proposed for sequential discriminant BoWs-based action representation and code-book adaptation based on action discrimination in a reduced dimensionality feature space where action classes are better discriminated. Experiments on four publicly available action recognition data sets demonstrate that the proposed unified approach increases the discriminative ability of the obtained video representation, providing enhanced action classification performance.
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
ICIP2
2015 Large-scale nonlinear facial image classification based on approximate kernel Extreme Learning Machine
abstract
In this paper, we propose a scheme that can be used in large-scale nonlinear facial image classification problems. An approximate solution of the kernel Extreme Learning Machine classifier is formulated and evaluated. Experiments on two publicly available facial image datasets using two popular facial image representations illustrate the effectiveness and efficiency of the proposed approach. The proposed Approximate Kernel Extreme Learning Machine classifier is able to scale well in both time and memory, while achieving good generalization performance. Specifically, it is shown that it outperforms the standard ELM approach for the same time and memory requirements. Compared to the original kernel ELM approach, it achieves similar (or better) performance, while scaling well in both time and memory with respect to the training set cardinality.
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
ICIP2
2015 Facial image analysis based on two-dimensional linear discriminant analysis exploiting symmetry
abstract
In this paper a novel subspace learning technique is introduced for facial image analysis. The proposed technique takes into account the symmetry nature of facial images. This information is exploited by properly incorporating a symmetry constraint into the objective function of the Two-Dimensional Linear Discriminant Analysis (2DLDA) to determine symmetric projection vectors. The performance of the proposed Symmetric Two-Dimensional Linear Discriminant Analysis was evaluated on real face recognition databases. Experimental results highlight the superiority of the proposed technique in comparison to standard approach.
Konstantinos Papachristou, Anastasios Tefas, Ioannis Pitas
ICIP2
2015 Visual voice activity detection based on spatiotemporal information and bag of words
abstract
A novel method for Visual Voice Activity Detection (V-VAD) that exploits local shape and motion information appearing at spatiotemporal locations of interest for facial region video description and the Bag of Words (BoW) model for facial region video representation is proposed in this paper. Facial region video classification is subsequently performed based on Single-hidden Layer Feedforward Neural (SLFN) network trained by applying the recently proposed kernel Extreme Learning Machine (kELM) algorithm on training facial videos depicting talking and non-talking persons. Experimental results on two publicly available V-VAD data sets, denote the effectiveness of the proposed method, since better generalization performance in unseen users is achieved, compared to recently proposed state-of-the-art methods.
Fotini Patrona, Alexandros Iosifidis, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas
ICIP3
2015 Kernel matrix trimming for improved Kernel K-means clustering
abstract
The Kernel k-Means algorithm for clustering extends the classic k-Means clustering algorithm. It uses the kernel trick to implicitly calculate distances on a higher dimensional space, thus overcoming the classic algorithm's inability to handle data that are not linearly separable. Given a set of n elements to cluster, the n × n kernel matrix is calculated, which contains the dot products in the higher dimensional space of every possible combination of two elements. This matrix is then referenced to calculate the distance between an element and a cluster center, as per classic k-Means. In this paper, we propose a novel algorithm for zeroing elements of the kernel matrix, thus trimming the matrix, which results in reduced memory complexity and improved clustering performance.
Nikolaos Tsapanos, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas
ICIP2
2015 Distance-based human action recognition using optimized class representations
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
Neurocomputing2
2015 DropELM: Fast neural network regularization with Dropout and DropConnect
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
Neurocomputing2
2015 Subclass Graph Embedding and a Marginal Fisher Analysis paradigm
Anastasios Maronidis, Anastasios Tefas, Ioannis Pitas
Pattern Recognit.2
2015 A distributed framework for trimmed Kernel k-Means clustering
Nikolaos Tsapanos, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas
Pattern Recognit.2
2015 On the kernel Extreme Learning Machine classifier
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
Pattern Recognit. Lett.2
2015 Sparse extreme learning machine classifier exploiting intrinsic graphs
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
Pattern Recognit. Lett.2
2015 Facial image clustering in stereoscopic videos using double spectral analysis
Georgios Orfanidis, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas
Signal Process. Image Commun.2
2015 Class-Specific Reference Discriminant Analysis With Application in Human Behavior Analysis
abstract
In this paper, a novel nonlinear subspace learning technique for class-specific data representation is proposed. A novel data representation is obtained by applying nonlinear class-specific data projection to a discriminant feature space, where the data belonging to the class under consideration are enforced to be close to their class representation, while the data belonging to the remaining classes are enforced to be as far as possible from it. A class is represented by an optimized class vector, enhancing class discrimination in the resulting feature space. An iterative optimization scheme is proposed to this end, where both the optimal nonlinear data projection and the optimal class representation are determined in each optimization step. The proposed approach is tested on three problems relating to human behavior analysis: Face recognition, facial expression recognition, and human action recognition. Experimental results denote the effectiveness of the proposed approach, since the proposed class-specific reference discriminant analysis outperforms kernel discriminant analysis, kernel spectral regression, and class-specific kernel discriminant analysis, as well as support vector machine-based classification, in most cases.
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
IEEE Trans. Hum. Mach. Syst.2
2015 Graph Embedded Nonparametric Mutual Information for Supervised Dimensionality Reduction
abstract
In this paper, we propose a novel algorithm for dimensionality reduction that uses as a criterion the mutual information (MI) between the transformed data and their corresponding class labels. The MI is a powerful criterion that can be used as a proxy to the Bayes error rate. Furthermore, recent quadratic nonparametric implementations of MI are computationally efficient and do not require any prior assumptions about the class densities. We show that the quadratic nonparametric MI can be formulated as a kernel objective in the graph embedding framework. Moreover, we propose its linear equivalent as a novel linear dimensionality reduction algorithm. The derived methods are compared against the state-of-the-art dimensionality reduction algorithms with various classifiers and on various benchmark and real-life datasets. The experimental results show that nonparametric MI as an optimization objective for dimensionality reduction gives comparable and in most of the cases better results compared with other dimensionality reduction methods.
Dimitrios Bouzas, Nikolaos Arvanitopoulos, Anastasios Tefas
IEEE Trans. Neural Networks Learn. Syst.3
2014 Facial image clustering in stereo videos using local binary patterns and double spectral analysis
abstract
In this work we propose the use of local binary patterns in combination with double spectral analysis for facial image clustering applied to 3D (stereoscopic) videos. Double spectral clustering involves the fusion of two well known algorithms: Normalized cuts and spectral clustering in order to improve the clustering performance. The use of local binary patterns upon selected fiducial points on the facial images proved to be a good choice for describing images. The framework is applied on 3D videos and makes use of the additional information deriving from the existence of two channels, left and right for further improving the clustering results.
Georgios Orfanidis, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas
CIDM2
2014 Stereoscopic video description for human action recognition
abstract
In this paper, a stereoscopic video description method is proposed that indirectly incorporates scene geometry information derived from stereo disparity, through the manipulation of video interest points. This approach is flexible and able to cooperate with any monocular low-level feature descriptor. The method is evaluated on the problem of recognizing complex human actions in natural settings, using a publicly available action recognition database of unconstrained stereoscopic 3D videos, coming from Hollywood movies. It is compared both against competing depth-aware approaches and a state-of-the-art monocular algorithm. Experimental results denote that the proposed approach outperforms them and achieves state-of-the-art performance.
Ioannis Mademlis, Alexandros Iosifidis, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas
CIMSIVP3
2014 Minimum Variance Extreme Learning Machine for human action recognition
abstract
In this paper we propose an algorithm for Single-hidden Layer Feedforward Neural networks training. Based on the observation that the learning process of such networks can be considered to be a non-linear mapping of the training data to a high-dimensional feature space, followed by a data projection process to a low-dimensional space where classification is performed by a linear classifier, we extend the Extreme Learning Machine (ELM) algorithm in order to exploit the training data dispersion in its optimization process. The proposed Minimum Variance Extreme Learning Machine classifier is evaluated in human action recognition, where we compare its performance with that of other ELM-based classifiers, as well as the kernel Support Vector Machine classifier.
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
ICASSP2
2014 Human action recognition based on bag of features and multi-view neural networks
abstract
In this paper, we employ Single-hidden Layer Feedforward Neural networks in order to perform human action recognition based on multiple action representations. In order to determine both optimized network and action representation combination weights, we propose an optimization process that jointly minimizes the overall network training error and the within-class variance of the training data in the corresponding hidden layer spaces. The proposed approach has been evaluated by using the state-of-the-art Bag of Features-based action video representation on three publicly available action recognition databases, where it outperforms two commonly used video representation combination approaches, as well as the best single-descriptor classification outcome.
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
ICIP2
2014 Stereoscopic video shot clustering into semantic concepts based on visual and disparity information
abstract
In this paper, we propose a framework for clustering shots from stereoscopic videos into clusters that correspond to semantic concepts exploiting visual and disparity information. Various color, disparity and texture descriptors are applied to shot key frames for obtaining low-level representations. Self Organizing Maps are subsequently employed upon various combinations of these representations in order to determine a lattice of representative semantic concepts. Experimental results on performances and football stereoscopic videos show that the use of disparity information leads to better clustering compared to using visual information only.
Konstantinos Papachristou, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas
ICIP2
2014 Laplacian Support Vector Analysis for Subspace Discriminative Learning
abstract
In this paper we propose a novel dimensionality reduction method that is based on successive Laplacian SVM projections in orthogonal deflated subspaces. The proposed method, called Laplacian Support Vector Analysis, produces projection vectors, which capture the discriminant information that lies in the subspace orthogonal to the standard Laplacian SVMs. We show that the optimal vectors on these deflated subspaces can be computed by successively training a standard SVM with specially designed deflation kernels. The resulting normal vectors contain discriminative information that can be used for feature extraction. In our analysis, we derive an explicit form for the deflation matrix of the mapped features in both the initial and the Hilbert space by using the kernel trick and thus, we can handle linear and non-linear deflation transformations. Experimental results in several benchmark datasets illustrate the strength of our proposed algorithm.
Nikolaos Arvanitopoulos, Dimitrios Bouzas, Anastasios Tefas
ICPR3
2014 Random Walk Kernel Applications to Classification Using Support Vector Machines
abstract
Kernel Methods are algorithms that are widely used, mainly because they can implicitly perform a non-linear mapping of the input data to a high dimensional feature space. In this paper, novel Kernel Matrices, that reflect the general structure of data, are proposed for classification. The proposed Matrices exploit properties of the graph theory, which are generated using power iterations of already known Kernel Matrices and three approaches are presented. Experiments on various datasets are conducted and statistical tests are performed, comparing our proposed approach against current Kernel Matrices used on support vector machines. Also, experiments on real datasets for folk dance and activity recognition that highlight the superiority of our proposed method, are provided.
Vasileios Gavriilidis, Anastasios Tefas
ICPR2
2014 Semi-supervised Classification of Human Actions Based on Neural Networks
abstract
In this paper, we propose a novel algorithm for Single-hidden Layer Feed forward Neural networks training which is able to exploit information coming from both labeled and unlabeled data for semi-supervised action classification. We extend the Extreme Learning Machine algorithm by incorporating appropriate regularization terms describing geometric properties and discrimination criteria of the training data representation in the ELM space to this end. The proposed algorithm is evaluated on human action recognition, where its performance is compared with that of other (semi-)supervised classification schemes. Experimental results on two publicly available action recognition databases denote its effectiveness.
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
ICPR2
2014 Shot type characterization in 2D and 3D video content
abstract
Due to the enormous increase of video and image content on the web in the last decades, automatic video annotation became a necessity. The successful annotation of video and image content facilitate a successful indexing and retrieval in search databases. In this work we study a variety of possible shot type characterizations that can be assigned in a single video frame or still image. Possible ways to propagate these characterizations to a video segment (or to an entire shot) are also discussed. A method for the detection of Over-the-Shoulder shots in 3D (stereo) video is also proposed.
Ioannis Tsingalis, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas
MMSP2
2014 Regularized extreme learning machine for multi-view semi-supervised action recognition
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
Neurocomputing2
2014 Kernel Reference Discriminant Analysis
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
Pattern Recognit. Lett.2
2014 Discriminant Bag of Words based representation for human action recognition
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
Pattern Recognit. Lett.2
2014 Stereo object tracking with fusion of texture, color and disparity information
Olga Zoidi, Nikos Nikolaidis 0001, Anastasios Tefas, Ioannis Pitas
Signal Process. Image Commun.3
2014 Projected Gradients for Subclass Discriminant Nonnegative Subspace Learning
abstract
Current discriminant nonnegative matrix factorization (NMF) methods either do not guarantee convergence to a stationary limit point or assume a compact data distribution inside classes, thus ignoring intra class variance in extracting discriminant data samples representations. To address both limitations, we regard that data inside each class has a multimodal distribution, forming various subclasses and perform optimization using a projected gradients framework to ensure limit point stationarity. The proposed method combines appropriate clustering-based discriminant criteria in the NMF decomposition cost function, in order to find discriminant projections that enhance class separability in the reduced dimensional projection space, thus improving classification performance. The developed algorithms have been applied to facial expression, face and object recognition, and experimental results verified that they successfully identified discriminant parts, thus enhancing recognition performance.
Symeon Nikitidis, Anastasios Tefas, Ioannis Pitas
IEEE Trans. Cybern.2
2014 Maximum Margin Projection Subspace Learning for Visual Data Analysis
abstract
Visual pattern recognition from images often involves dimensionality reduction as a key step to discover a lower dimensional image data representation and obtain a more manageable problem. Contrary to what is commonly practiced today in various recognition applications where dimensionality reduction and classification are independently treated, we propose a novel dimensionality reduction method appropriately combined with a classification algorithm. The proposed method called maximum margin projection pursuit, aims to identify a low dimensional projection subspace, where samples form classes that are better discriminated, i.e., are separated with maximum margin. The proposed method is an iterative alternate optimization algorithm that computes the maximum margin projections exploiting the separating hyperplanes obtained from training a support vector machine classifier in the identified low dimensional space. Experimental results on both artificial data, as well as, on popular databases for facial expression, face and object recognition verified the superiority of the proposed method against various state-of-the-art dimensionality reduction algorithms.
Symeon Nikitidis, Anastasios Tefas, Ioannis Pitas
IEEE Trans. Image Process.2
2014 Symmetric Subspace Learning for Image Analysis
abstract
Subspace learning (SL) is one of the most useful tools for image analysis and recognition. A large number of such techniques have been proposed utilizing a priori knowledge about the data. In this paper, new subspace learning techniques are presented that use symmetry constraints in their objective functions. The rational behind this idea is to exploit the a priori knowledge that geometrical symmetry appears in several types of data, such as images, objects, faces, and so on. Experiments on artificial, facial expression recognition, face recognition, and object categorization databases highlight the superiority and the robustness of the proposed techniques, in comparison with standard SL techniques.
Konstantinos Papachristou, Anastasios Tefas, Ioannis Pitas
IEEE Trans. Image Process.2
2014 Person Identity Label Propagation in Stereo Videos
abstract
In this paper a novel method is introduced for propagating person identity labels on facial images extracted from stereo videos. It operates on image data with multiple representations and calculates a projection matrix that preserves locality information and a priori pairwise information, in the form of must-link and cannot-link constraints between the various data representations. The final data representation is a linear combination of the projections of all data representations. Moreover, the proposed method takes into account information obtained through data clustering. This information is exploited during the data propagation step in two ways: to regulate the similarity strength between the projected data and to indicate which samples should be selected for label propagation initialization. The performance of the proposed Multiple Locality Preserving Projections with Cluster-based Label Propagation (MLPP-CLP) method was evaluated on facial images extracted from stereo movies. Experimental results showed that the proposed method outperforms state of the art methods.
Olga Zoidi, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas
IEEE Trans. Multim.2
2013 Feature Comparison and Feature Fusion for Traditional Dances Recognition
Ioannis Kapsouras, Stylianos Karanikolos, Nikos Nikolaidis 0001, Anastasios Tefas
EANN (1)4
2013 Direct Multi-label Linear Discriminant Analysis
Maria Oikonomou, Anastasios Tefas
EANN (1)2
2013 Neural Networks for Digital Media Analysis and Description
Anastasios Tefas, Alexandros Iosifidis, Ioannis Pitas
EANN (1)1
2013 Active classification for human action recognition
abstract
In this paper, we propose a novel classification method involving two processing steps. Given a test sample, the training data residing to its neighborhood are determined. Classification is performed by a Single-hidden Layer Feedforward Neural network exploiting labeling information of the training data appearing in the test sample neighborhood and using the rest training data as unlabeled. By following this approach, the proposed classification method focuses the classification problem on the training data that are more similar to the test sample under consideration and exploits information concerning to the training set structure. Compared to both static classification exploiting all the available training data and dynamic classification involving data selection for classification, the proposed active classification method provides enhanced classification performance in two publicly available action recognition databases.
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
ICIP2
2013 Exploiting the SVM constraints in NMF with application in eating and drinking activity recognition
abstract
A novel method is introduced for exploiting the support vector machine constraints in nonnegative matrix factorization. The notion of the proposed method is to find the projection matrix that projects the data to a low-dimensional space so that the data projections between the two classes are separated with maximum margin. Experiments were performed for the task of eating and drinking activity classification. Experimental results showed that the proposed method achieves better classification performance than the state of the art nonnegative matrix factorization and discriminant nonnegative matrix factorization followed by support vector machines classification.
Olga Zoidi, Anastasios Tefas, Ioannis Pitas
ICIP2
2013 Learning sparse representations for view-independent human action recognition based on fuzzy distances
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
Neurocomputing2
2013 Using robust dispersion estimation in support vector machines
Nicholas Vretos, Anastasios Tefas, Ioannis Pitas
Pattern Recognit.2
2013 Dynamic action recognition based on dynemes and Extreme Learning Machine
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
Pattern Recognit. Lett.2
2013 Multi-view action recognition based on action volumes, fuzzy distances and cluster discriminant analysis
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
Signal Process.2
2013 Minimum Class Variance Extreme Learning Machine for Human Action Recognition
abstract
In this paper, we propose a novel method aiming at view-independent human action recognition. Action description is based on local shape and motion information appearing at spatiotemporal locations of interest in a video. Action representation involves fuzzy vector quantization, while action classification is performed by a feedforward neural network. A novel classification algorithm, called minimum class variance extreme learning machine, is proposed in order to enhance the action classification performance. The proposed method can successfully operate in situations that may appear in real application scenarios, since it does not set any assumption concerning the visual scene background and the camera view angle. Experimental results on five publicly available databases, aiming at different application scenarios, denote the effectiveness of both the adopted action recognition approach and the proposed minimum class variance extreme learning machine algorithm.
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
IEEE Trans. Circuits Syst. Video Technol.2
2013 Visual Object Tracking Based on Local Steering Kernels and Color Histograms
abstract
In this paper, we propose a visual object tracking framework, which employs an appearance-based representation of the target object, based on local steering kernel descriptors and color histogram information. This framework takes as input the region of the target object in the previous video frame and a stored instance of the target object, and tries to localize the object in the current frame by finding the frame region that best resembles the input. As the object view changes over time, the object model is updated, hence incorporating these changes. Color histogram similarity between the detected object and the surrounding background is employed for background subtraction. Experiments are conducted to test the performance of the proposed framework under various conditions. The proposed tracking scheme is proven to be successful in tracking objects under scale and rotation variations and partial occlusion, as well as in tracking rather slowly deformable articulated objects.
Olga Zoidi, Anastasios Tefas, Ioannis Pitas
IEEE Trans. Circuits Syst. Video Technol.2
2013 Multidimensional Sequence Classification Based on Fuzzy Distances and Discriminant Analysis
abstract
In this paper, we present a novel method aiming at multidimensional sequence classification. We propose a novel sequence representation, based on its fuzzy distances from optimal representative signal instances, called statemes. We also propose a novel modified clustering discriminant analysis algorithm minimizing the adopted criterion with respect to both the data projection matrix and the class representation, leading to the optimal discriminant sequence class representation in a low-dimensional space, respectively. Based on this representation, simple classification algorithms, such as the nearest subclass centroid, provide high classification accuracy. A three step iterative optimization procedure for choosing statemes, optimal discriminant subspace and optimal sequence class representation in the final decision space is proposed. The classification procedure is fast and accurate. The proposed method has been tested on a wide variety of multidimensional sequence classification problems, including handwritten character recognition, time series classification and human activity recognition, providing very satisfactory classification results.
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
IEEE Trans. Knowl. Data Eng.2
2013 On the Optimal Class Representation in Linear Discriminant Analysis
abstract
Linear discriminant analysis (LDA) is a widely used technique for supervised feature extraction and dimensionality reduction. LDA determines an optimal discriminant space for linear data projection based on certain assumptions, e.g., on using normal distributions for each class and employing class representation by the mean class vectors. However, there might be other vectors that can represent each class, to increase class discrimination. In this brief, we propose an optimization scheme aiming at the optimal class representation, in terms of Fisher ratio maximization, for LDA-based data projection. Compared with the standard LDA approach, the proposed optimization scheme increases class discrimination in the reduced dimensionality space and achieves higher classification rates in publicly available data sets.
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
IEEE Trans. Neural Networks Learn. Syst.2
2013 Multiplicative Update Rules for Concurrent Nonnegative Matrix Factorization and Maximum Margin Classification
abstract
The state-of-the-art classification methods which employ nonnegative matrix factorization (NMF) employ two consecutive independent steps. The first one performs data transformation (dimensionality reduction) and the second one classifies the transformed data using classification methods, such as nearest neighbor/centroid or support vector machines (SVMs). In the following, we focus on using NMF factorization followed by SVM classification. Typically, the parameters of these two steps, e.g., the NMF bases/coefficients and the support vectors, are optimized independently, thus leading to suboptimal classification performance. In this paper, we merge these two steps into one by incorporating maximum margin classification constraints into the standard NMF optimization. The notion behind the proposed framework is to perform NMF, while ensuring that the margin between the projected data of the two classes is maximal. The concurrent NMF factorization and support vector optimization are performed through a set of multiplicative update rules. In the same context, the maximum margin classification constraints are imposed on the NMF problem with additional discriminant constraints and respective multiplicative update rules are extracted. The impact of the maximum margin classification constraints on the NMF factorization problem is addressed in Section VI. Experimental results in several databases indicate that the incorporation of the maximum margin classification constraints into the NMF and discriminant NMF objective functions improves the accuracy of the classification.
Olga Zoidi, Anastasios Tefas, Ioannis Pitas
IEEE Trans. Neural Networks Learn. Syst.2
2012 Eating and drinking activity recognition based on discriminant analysis of fuzzy distances and activity volumes
abstract
Eating and drinking activity recognition can be considered a solitary research field in activity recognition area. The development of an application capable to identify human eating and drinking activity can be really useful in a smart home environment targeting to extend independent living of older persons in the early stages of dementia. In this paper a novel method aiming at eating and drinking activity recognition is presented. Activities are considered as a sequence of human body poses forming 3D volumes, in which the third dimension refers to time. Fuzzy Vector Quantization is performed to associate the 3D volume representation of an activity video with 3D volume prototypes and Linear Discriminant Analysis is used to map activity representations in a low dimensional discriminant feature space. In this space a simple Nearest Centroid classification procedure leads to very satisfactory classification results.
Alexandros Iosifidis, Ermioni Marami, Anastasios Tefas, Ioannis Pitas
ICASSP3
2012 Visual object tracking based on the object's salient features with application in automatic nutrition assistance
abstract
A novel method for object tracking in videos which can find application in eating and drinking activity recognition is proposed. The query object is detected in the first video frame, extracting a new query image. The initial query image along with the obtained query image are then compared with patches within a determined search region around the position of the detected object in the previous frame. For each image, the local steering kernels are extracted and the similarity between a query image and the patches of the video frame is measured by calculating the cosine similarity. The proposed method finds application in eating and drinking activity recognition.
Olga Zoidi, Anastasios Tefas, Ioannis Pitas
ICASSP2
2012 Discriminant action representation for view-invariant person identification
abstract
In this paper we propose a novel person identification method exploiting human motion information. Persons are described by using their poses during action execution. Identification process involves Fuzzy Vector Quantization and Discriminant Learning. In the case of multiple cameras used in the identification phase, single-view identification results combination is achieved by employing a Bayesian combination strategy. The proposed identification approach does not set the assumptions of known action class and number of capturing cameras in the identification phase. Experimental results on two publicly available video databases denote the effectiveness of the proposed approach.
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
ICIP2
2012 Exploiting subclass information in Support Vector Machines
Georgios Orfanidis, Anastasios Tefas
ICPR2
2012 Neural representation and learning for multi-view human action recognition
abstract
In this paper we propose a novel method aiming at view-independent multi-view action recognition. Instead of combining the information provided by all the cameras forming the camera setup, for action representation and classification, we perform single-view action representation and classification to all the available videos depicting the person under consideration independently. Action representation involves a self organizing neural network training followed by fuzzy vector quantization. Action classification is performed by a feedforward neural network which is trained for view-invariant action recognition. Multiple action classification results combination based on Bayesian learning, in the recognition phase, results to high action recognition accuracy. The performance of the proposed action recognition method is evaluated on two publicly available databases, aiming at different application scenarios.
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
IJCNN2
2012 Multi-view human movement recognition based on fuzzy distances and linear discriminant analysis
Alexandros Iosifidis, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas
Comput. Vis. Image Underst.2
2012 Subclass discriminant Nonnegative Matrix Factorization for facial image analysis
Symeon Nikitidis, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas
Pattern Recognit.2
2012 Shape matching using a binary search tree structure of weak classifiers
Nikolaos Tsapanos, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas
Pattern Recognit.2
2012 Activity-Based Person Identification Using Fuzzy Representation and Discriminant Learning
abstract
In this paper, a novel view invariant person identification method based on human activity information is proposed. Unlike most methods proposed in the literature, in which “walk” (i.e., gait) is assumed to be the only activity exploited for person identification, we incorporate several activities in order to identify a person. A multicamera setup is used to capture the human body from different viewing angles. Fuzzy vector quantization and linear discriminant analysis are exploited in order to provide a discriminant activity representation. Person identification, activity recognition, and viewing angle specification results are obtained for all the available cameras independently. By properly combining these results, a view-invariant activity-independent person identification method is obtained. The proposed approach has been tested in challenging problem setups, simulating real application situations. Experimental results are very promising.
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
IEEE Trans. Inf. Forensics Secur.2
2012 View-Invariant Action Recognition Based on Artificial Neural Networks
abstract
In this paper, a novel view invariant action recognition method based on neural network representation and recognition is proposed. The novel representation of action videos is based on learning spatially related human body posture prototypes using self organizing maps. Fuzzy distances from human body posture prototypes are used to produce a time invariant action representation. Multilayer perceptrons are used for action classification. The algorithm is trained using data from a multi-camera setup. An arbitrary number of cameras can be used in order to recognize actions using a Bayesian framework. The proposed method can also be applied to videos depicting interactions between humans, without any modification. The use of information captured from different viewing angles leads to high classification performance. The proposed method is the first one that has been tested in challenging experimental setups, a fact that denotes its effectiveness to deal with most of the open issues in action recognition.
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas
IEEE Trans. Neural Networks Learn. Syst.2
2011 Optimizing Linear Discriminant Error Correcting Output Codes Using Particle Swarm Optimization
Dimitrios Bouzas, Nikolaos Arvanitopoulos, Anastasios Tefas
ICANN (2)3
2011 Facial expression recognition using clustering discriminant Non-negative Matrix Factorization
abstract
Non-negative Matrix Factorization (NMF) is among the most popular subspace methods widely used in a variety of image processing problems. Recently, a discriminant NMF method that incorporates Linear Discriminant Analysis criteria and achieves an efficient decomposition of the provided data to its discriminant parts has been proposed. However, this approach poses several limitations since it assumes that the underline data distribution forms compact sets which is often unrealistic. To remedy this limitation we regard that data inside each class form various number of clusters and apply a Clustering based Discriminant Analysis. The proposed method combines appropriate discriminant constraints in the NMF decomposition cost function in order to address the problem of finding discriminant projections that enhance class separability in the reduced dimensional projection space. Experimental results performed on the Cohn-Kanade database verified the effectiveness of the proposed method in the facial expression recognition task.
Symeon Nikitidis, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas
ICIP2
2011 Improving subspace learning for facial expression recognition using person dependent and geometrically enriched training sets
Anastasios Maronidis, Dimitris Bolis, Anastasios Tefas, Ioannis Pitas
Neural Networks3
2010 Using Particle Swarm Optimization for scaling and rotation invariant face detection
abstract
Common face detection algorithms exhaustively search in all possible locations in the image for precisely located, frontal faces. In this paper, a novel face detection algorithm based on Particle Swarm Optimization (PSO) method for searching in the image is proposed. The algorithm uses a linear Support Vector Machine (SVM) as fast and accurate classifier and searches for a face in four dimensions: plane, orientation of the face, size of the face. Using PSO, the exhaustive search in all possible combinations of the 4D coordinates can be avoided, saving time and decreasing the computational complexity. Moreover, linear SVMs are proved to be a powerful and fast classifier for demanding applications. Experimental results under real recording conditions in the BioID and VALID database are very promising and indicate the potential use of the proposed approach to real applications.
Ermioni Marami, Anastasios Tefas
IEEE Congress on Evolutionary Computation2
2010 Improving the Robustness of Subspace Learning Techniques for Facial Expression Recognition
Dimitris Bolis, Anastasios Maronidis, Anastasios Tefas, Ioannis Pitas
ICANN (1)3
2010 Frontal View Recognition Using Spectral Clustering and Subspace Learning Methods
Anastasios Maronidis, Anastasios Tefas, Ioannis Pitas
ICANN (1)2
2010 Neural Networks Training for Weapon Selection in First-Person Shooter Games
Stelios Petrakis, Anastasios Tefas
ICANN (3)2
2010 Dynamic Shape Learning and Forgetting
Nikolaos Tsapanos, Anastasios Tefas, Ioannis Pitas
ICANN (3)2
2010 Subclass Error Correcting Output Codes Using Fisher's Linear Discriminant Ratio
abstract
Error-Correcting Output Codes (ECOC) with sub-classes reveal a common way to solve multi-class classification problems. According to this approach, a multi-class problem is decomposed into several binary ones based on the maximization of the mutual information (MI) between the classes and their respective labels. The MI is modelled through the fast quadratic mutual information (FQMI) procedure. However, FQMI is not applicable on large datasets due to its high algorithmic complexity. In this paper we propose Fisher's Linear Discriminant Ratio (FLDR) as an alternative decomposition criterion which is of much less computational complexity and achieves in most experiments conducted better classification performance. Furthermore, we compare FLDR against FQMI for facial expression recognition over the Cohn-Kanade database.
Nikolaos Arvanitopoulos, Dimitrios Bouzas, Anastasios Tefas
ICPR3
2010 Optimizing subclass discriminant Error Correcting Output Codes using particle swarm optimization
abstract
Error-Correcting Output Codes (ECOC) reveal a common way to model multi-class classification problems. According to this state of the art technique, a multi-class problem is decomposed into several binary ones. Additionally, on the ECOC framework we can apply the subclass technique (sub-ECOC), where by splitting the initial classes of the problem we create larger but easier to solve ECOC configurations. The multi-class problem's decomposition is achieved via a discriminant tree creation procedure. This discriminant tree's creation is controlled by a triplet of thresholds that define a set of user defined splitting standards. The selection of the thresholds plays a major role in the classification performance. In our work we show that by optimizing these thresholds via particle swarm optimization we improve significantly the classification performance. Moreover, using Support Vector Machines (SVMs) as classifiers we can optimize in the same time both the thresholds of sub-ECOC and the parameters C and σ of the SVMs, resulting in even better classification performance. Extensive experiments in both real and artificial data illustrate the superiority of the proposed approach in terms of performance.
Dimitrios Bouzas, Nikolaos Arvanitopoulos, Anastasios Tefas
IJCNN3
2010 Online shape learning using binary search trees
Nikolaos Tsapanos, Anastasios Tefas, Ioannis Pitas
Image Vis. Comput.2
2010 Salient feature and reliable classifier selection for facial expression classification
Marios Kyperountas, Anastasios Tefas, Ioannis Pitas
Pattern Recognit.2
2009 Human identification from human movements
abstract
In this paper a multi-modal method for human identification that exploits the discrimination power of several movement types performed from the same human is proposed. Utilizing a fuzzy vector quantization (FVQ) and linear discriminant analysis (LDA) based algorithm, an unknown movement is first classified, and, then, the person performing the movement is recognized from a movement specific person classifier. In case that the unknown person performs more than one movements, a multi-modal algorithm combines the results of the individual classifiers to yield the final decision for the id of the unknown human. Using a publicly available database, we provide promising results regarding the discrimination power of the different movements for the human identification task, as well as we indicate that the combination of the individual classifiers may increase the robustness of the human recognition algorithm.
Nikolaos Gkalelis, Anastasios Tefas, Ioannis Pitas
ICIP2
2009 Pairwise facial expression classification
abstract
This paper presents a novel facial expression recognition methodology. In order to classify the expression of a test face to one of seven pre-determined facial expression classes, multiple two-class classification tasks are carried out. For each such task, a unique set of features is identified that is enhanced, in terms of its ability to help produce a proper separation between the two specific classes. The selection of these sets of features is accomplished by making use of a class separability measure that is utilized in an iterative process. Fisher's linear discriminant is employed in order to produce the separation between each pair of classes and train each two-class classifier. In order to combine the classification results from all two-class classifiers, the `voting' classifier-decision fusion process is employed. The standard JAFFE database is utilized in order to evaluate the performance of this algorithm. Experimental results show that the proposed methodology provides a good solution to the facial expression recognition problem.
Marios Kyperountas, Anastasios Tefas, Ioannis Pitas
MMSP2
2008 Motivating class-specific nonlinear projections for single and multiple view face verification
abstract
In this paper we motivate the use of class-specific nonlinear subspace methods for face verification. The problem of face verification is considered as a two-class problem (genuine versus impostor class). The typical Fisher's linear discriminant analysis (FLDA) gives only one or two projections in a two-class problem. This is a very strict limitation to the search of discriminant dimensions. As for the FLDA for N class problems (N > 2) the transformation is not person specific. In order to remedy these limitations of FLDA, exploit the individuality of human faces and take into consideration the fact that the distribution of facial images, under different viewpoints, illumination variations and facial expression is highly complex and non-linear, novel kernel discriminant algorithms are used. The new method was tested in the face verification problem using single and multiple view datasets and found to outperform other commonly used kernel approaches.
Georgios Goudelis, Stefanos Zafeiriou, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas
ICIP3
2008 Face recognition via adaptive discriminant clustering
abstract
This paper presents a methodology that tackles the face recognition problem by accommodating multiple clustering steps. At each clustering step, the test and training faces are projected to a discriminant space and the projected training data are partitioned into clusters using the k-means algorithm. Then a subset of the training data clusters is selected, based on how similar the faces in these clusters are to the test face. In the clustering step that follows a new discriminant space is defined by processing this subset and both the test and training data are projected to this space. This process is repeated until one final cluster is selected and the most similar, to the test face, face class contained is set as the identity match. The UMIST and XM2VTS face databases have been used to evaluate the algorithm and results indicate that the proposed framework provides a promising solution to the face recognition problem.
Marios Kyperountas, Anastasios Tefas, Ioannis Pitas
ICIP2
2008 Sparse human movement representation and recognition
abstract
In this paper a novel method for human movement representation and recognition is proposed. A movement type is regarded as a unique combination of basic movement patterns, the so-called dynemes. The fuzzy c-mean (FCM) algorithm is used to identify the dynemes in the input space and allow the expression of a posture in terms of these dynemes. In the so-called dyneme space, the sparse posture representations of a movement are combined to represent the movement as a single point in that space, and linear discriminant analysis (LDA) is further employed to increase movement type discrimination and compactness of representation. This method allows for simple Mahalanobis or cosine distance comparison of movements, taking implicitly into account time shifts and internal speed variations, and, thus, aiding the design of a real-time movement recognition algorithm.
Nikolaos Gkalelis, Anastasios Tefas, Ioannis Pitas
MMSP2
2008 Dynamic training using multistage clustering for face recognition
Marios Kyperountas, Anastasios Tefas, Ioannis Pitas
Pattern Recognit.2
2008 Combining Fuzzy Vector Quantization With Linear Discriminant Analysis for Continuous Human Movement Recognition
abstract
In this paper, a novel method for continuous human movement recognition based on fuzzy vector quantization (FVQ) and linear discriminant analysis (LDA) is proposed. We regard a movement as a unique combination of basic movement patterns, the so-called dynemes. The proposed algorithm combines FVQ and LDA to discover the most discriminative dynemes as well as represent and discriminate the different human movements in terms of these dynemes. This method allows for simple Mahalanobis or cosine distance comparison of not aligned human movements, taking into account implicitly time shifts and internal speed variations, and, thus, aiding the design of a real-time continuous human movement recognition algorithm. The effectiveness and robustness of this method is shown by experimental results on a standard dataset with videos captured under real conditions, and on a new video dataset created using motion capture data.
Nikolaos Gkalelis, Anastasios Tefas, Ioannis Pitas
IEEE Trans. Circuits Syst. Video Technol.2
2008 Automated Facial Pose Extraction From Video Sequences Based on Mutual Information
abstract
Estimation of the facial pose in video sequences is one of the major issues in many vision systems such as face-based biometrics, scene understanding for humans, and others. The proposed method uses a novel pose estimation algorithm based on mutual information to extract any required facial poses from video sequences. The method extracts the poses automatically and classifies them according to view angle. Experimental results on the XM2VTS video database and on a new database created for the needs of this research indicated a pose classification rate of 99.2% while it was shown that it outperforms a principal component analysis reconstruction method that was used as a benchmark.
Georgios Goudelis, Anastasios Tefas, Ioannis Pitas
IEEE Trans. Circuits Syst. Video Technol.2
2007 Discriminant Graph Structures for Face Verification
abstract
Elastic graph matching is one of the most well known techniques for frontal face recognition/verification and one of the few techniques that can be combined successfully with fully automatic face localization and alignment methods. In this paper, we propose an algorithm for finding the most discriminant features upon a person's face and a person-specific graph is placed in the spatial coordinates that correspond to these discriminant features. We illustrate the improvements in performance by applying the proposed method in frontal face verification using the XM2VTS database.
Stefanos Zafeiriou, Anastasios Tefas, Ioannis Pitas
ICASSP (1)2
2007 A Novel Kernel Discriminant Analysis for Face Verification
abstract
In this paper a novel non-linear subspace method for face verification is proposed. The problem of face verification is considered as a two-class problem (genuine versus impostor class). The typical Fisher's linear discriminant analysis (FLDA) gives only one or two projections in a two-class problem. This is a very strict limitation to the search of discriminant dimensions. As for the FLDA for N class problems (N is greater than two) the transformation is not person specific. In order to remedy these limitations of FLDA, exploit the individuality of human faces and take into consideration the fact that the distribution of facial images, under different viewpoints, illumination variations and facial expression is highly complex and non-linear, novel kernel discriminant algorithms are proposed. The new methods are tested in the face verification problem using the XM2VTS database where it is verified that they outperform other commonly used kernel approaches.
Georgios Goudelis, Stefanos Zafeiriou, Anastasios Tefas, Ioannis Pitas
ICIP (4)3
2007 Face Verification using Locally Linear Discriminant Models
abstract
When linear discriminant analysis (LDA) is employed, the correct classification of a sample heavily depends on having an adequately large training set. This is often not possible in practical applications, such as person verification, where the lack of sufficient training samples causes improper estimation of a linear separation hyper-plane between the two classes. To overcome this shortcoming a novel algorithm that can handle the verification problem more efficiently than traditional LDA is presented. The dimensionality of the samples is reduced by breaking them down, thus creating subsets of smaller dimensionality feature vectors, and applying discriminant analysis on each subset. The resulting discriminant weight sets are themselves weighted under a normalization criterion, making the discriminant functions continuous in this sense. A series of simulations that formulate the face verification problem illustrate the cases for which our method outperforms traditional LDA and various statistical observations are made about the discriminant coefficients that are generated.
Marios Kyperountas, Anastasios Tefas, Ioannis Pitas
ICIP (4)2
2007 The discriminant elastic graph matching algorithm applied to frontal face verification
Stefanos Zafeiriou, Anastasios Tefas, Ioannis Pitas
Pattern Recognit.2
2007 Class-Specific Kernel-Discriminant Analysis for Face Verification
abstract
In this paper, novel nonlinear subspace methods for face verification are proposed. The problem of face verification is considered as a two-class problem (genuine versus impostor class). The typical Fisher's linear discriminant analysis (FLDA) gives only one or two projections in a two-class problem. This is a very strict limitation to the search of discriminant dimensions. As for the FLDA forNclass problems (Nis greater than two), the transformation is not person specific. In order to remedy these limitations of FLDA, exploit the individuality of human faces and take into consideration the fact that the distribution of facial images, under different viewpoints, illumination variations, and facial expression is highly complex and nonlinear, novel kernel-discriminant algorithms are proposed. The new methods are tested in the face verification problem using the XM2VTS, AR, ORL, Yale, and UMIST databases where it is verified that they outperform other commonly used kernel approaches such as kernel-PCA (KPCA), kernel direct discriminant analysis (KDDA), complete kernel Fisher's discriminant analysis (CKFDA), the two-class KDDA, CKFDA, and other two-class and multiclass variants of kernel-discriminant analysis based on Fisher's criterion.
Georgios Goudelis, Stefanos Zafeiriou, Anastasios Tefas, Ioannis Pitas
IEEE Trans. Inf. Forensics Secur.3
2007 Learning Discriminant Person-Specific Facial Models Using Expandable Graphs
abstract
In this paper, a novel algorithm for finding discriminant person-specific facial models is proposed and tested for frontal face verification. The most discriminant features of a person's face are found and a deformable model is placed in the spatial coordinates that correspond to these discriminant features. The discriminant deformable models, for verifying the person's identity, that are learned through this procedure are elastic graphs that are dense in the facial areas considered discriminant for a specific person and sparse in other less significant facial areas. The discriminant graphs are enhanced by a discriminant feature selection method for the graph nodes in order to find the most discriminant jet features. The proposed approach significantly enhances the performance of elastic graph matching in frontal face verification
Stefanos Zafeiriou, Anastasios Tefas, Ioannis Pitas
IEEE Trans. Inf. Forensics Secur.2
2007 Minimum Class Variance Support Vector Machines
abstract
In this paper, a modified class of support vector machines (SVMs) inspired from the optimization of Fisher's discriminant ratio is presented, the so-called minimum class variance SVMs (MCVSVMs). The MCVSVMs optimization problem is solved in cases in which the training set contains less samples that the dimensionality of the training vectors using dimensionality reduction through principal component analysis (PCA). Afterward, the MCVSVMs are extended in order to find nonlinear decision surfaces by solving the optimization problem in arbitrary Hilbert spaces defined by Mercer's kernels. In that case, it is shown that, under kernel PCA, the nonlinear optimization problem is transformed into an equivalent linear MCVSVMs problem. The effectiveness of the proposed approach is demonstrated by comparing it with the standard SVMs and other classifiers, like kernel Fisher discriminant analysis in facial image characterization problems like gender determination, eyeglass, and neutral facial expression detection.
Stefanos Zafeiriou, Anastasios Tefas, Ioannis Pitas
IEEE Trans. Image Process.2
2007 Weighted Piecewise LDA for Solving the Small Sample Size Problem in Face Verification
abstract
A novel algorithm that can be used to boost the performance of face-verification methods that utilize Fisher's criterion is presented and evaluated. The algorithm is applied to similarity, or matching error, data and provides a general solution for overcoming the "small sample size" (SSS) problem, where the lack of sufficient training samples causes improper estimation of a linear separation hyperplane between the classes. Two independent phases constitute the proposed method. Initially, a set of weighted piecewise discriminant hyperplanes are used in order to provide a more accurate discriminant decision than the one produced by the traditional linear discriminant analysis (LDA) methodology. The expected classification ability of this method is investigated throughout a series of simulations. The second phase defines proper combinations for person-specific similarity scores and describes an outlier removal process that further enhances the classification ability. The proposed technique has been tested on the M2VTS and XM2VTS frontal face databases. Experimental results indicate that the proposed framework greatly improves the face-verification performance.
Marios Kyperountas, Anastasios Tefas, Ioannis Pitas
IEEE Trans. Neural Networks2
2006 Exploiting discriminant information in nonnegative matrix factorization with application to frontal face verification
abstract
In this paper, two supervised methods for enhancing the classification accuracy of the Nonnegative Matrix Factorization (NMF) algorithm are presented. The idea is to extend the NMF algorithm in order to extract features that enforce not only the spatial locality, but also the separability between classes in a discriminant manner. The first method employs discriminant analysis in the features derived from NMF. In this way, a two-phase discriminant feature extraction procedure is implemented, namely NMF plus Linear Discriminant Analysis (LDA). The second method incorporates the discriminant constraints inside the NMF decomposition. Thus, a decomposition of a face to its discriminant parts is obtained and new update rules for both the weights and the basis images are derived. The introduced methods have been applied to the problem of frontal face verification using the well-known XM2VTS database. Both methods greatly enhance the performance of NMF for frontal face verification.
Stefanos Zafeiriou, Anastasios Tefas, Ioan Buciu, Ioannis Pitas
IEEE Trans. Neural Networks2
2005 Methods for improving discriminant analysis for face authentication
abstract
A novel algorithm that can be used to boost the performance of face authentication methods that utilize Fisher's criterion is presented. The algorithm is applied to matching error data and provides a general solution for overcoming the "small sample size" (SSS) problem, where the lack of sufficient training samples causes improper estimation of a linear separation hyperplane between the classes. Two independent phases constitute the proposed method. Initially, a set of locally linear discriminant models is used in order to calculate discriminant weights in a more accurate way than the traditional linear discriminant analysis (LDA) methodology. Additionally, defective discriminant coefficients are identified and reestimated. The second phase defines proper combinations for person-specific matching scores and describes an outlier removal process that enhances the classification ability. Our technique was tested on the M2VTS and XM2VTS frontal face databases. Experimental results indicate that the proposed framework greatly improves the authentication algorithm's performance.
Marios Kyperountas, Anastasios Tefas, Ioannis Pitas
ICASSP (2)2
2005 Enhanced transform-domain correlation-based audio watermarking
abstract
Various watermarking techniques have been proposed so far, aiming at the copyright protection of audio signals. Little effort has been made, however, in taking under consideration the spectrum of the watermark sequence itself and exploiting its frequency properties. An enhanced audio watermarking technique, based on correlation detection, is introduced in this paper, where high-frequency chaotic watermarks are multiplicatively embedded in the low frequencies of the DFT domain. A series of experiments have been conducted to demonstrate both detection reliability and robustness against attacks.
Anastasios Tefas, Alexia Giannoula, Nikos Nikolaidis 0001, Ioannis Pitas
ICASSP (2)1
2005 Exploiting discriminant information in elastic graph matching
abstract
In this paper, we investigate the use of discriminant techniques in the elastic graph matching (EGM) algorithm. First we use discriminant analysis in the feature vectors of the nodes in order to find the most discriminant features. The similarity measure for discriminant feature vectors and the node deformation are combined in a discriminant manner in order to form a local similarity measure between nodes. Moreover, the local similarity values at the nodes of the elastic graph, are weighted by coefficients that are also derived by some discriminant analysis in order to form a total similarity measure between faces. We illustrate the improvements in performance in frontal face verification using a modified multiscale morphological analysis.
Stefanos Zafeiriou, Anastasios Tefas, Ioannis Pitas
ICIP (3)2
2005 Blind Robust Watermarking Schemes for Copyright Protection of 3D Mesh Objects
abstract
In this paper, two novel methods suitable for blind 3D mesh object watermarking applications are proposed. The first method is robust against 3D rotation, translation, and uniform scaling. The second one is robust against both geometric and mesh simplification attacks. A pseudorandom watermarking signal is cast in the 3D mesh object by deforming its vertices geometrically, without altering the vertex topology. Prior to watermark embedding and detection, the object is rotated and translated so that its center of mass and its principal component coincide with the origin and the z-axis of the Cartesian coordinate system. This geometrical transformation ensures watermark robustness to translation and rotation. Robustness to uniform scaling is achieved by restricting the vertex deformations to occur only along the r coordinate of the corresponding (r, theta, phi) spherical coordinate system. In the first method, a set of vertices that correspond to specific angles theta is used for watermark embedding. In the second method, the samples of the watermark sequence are embedded in a set of vertices that correspond to a range of angles in the theta domain in order to achieve robustness against mesh simplifications. Experimental results indicate the ability of the proposed method to deal with the aforementioned attacks.
Stefanos Zafeiriou, Anastasios Tefas, Ioannis Pitas
IEEE Trans. Vis. Comput. Graph.2
2004 A blind robust watermarking scheme for copyright protection of 3d mesh models
Stefanos Zafeiriou, Anastasios Tefas, Ioannis Pitas
ICIP2
2003 Watermarking of 3D models using principal component analysis
abstract
A novel method for 3D model watermarking, robust to geometric distortions such as rotation, translation and scaling, is proposed. A ternary watermark is embedded in the vertex topology of a 3D model. A transformation of the model to an invariant space is proposed prior to watermark embedding. Simulation results indicate the ability of the proposed method to deal with the aforementioned attacks giving very good results.
Andreas Kalivas, Anastasios Tefas, Ioannis Pitas
ICASSP (5)2
2003 Improving the detection reliability of correlation-based watermarking techniques
abstract
The performance of watermarking schemes based on correlation detection is closely related to the frequency characteristics of the watermark sequence. In order to improve both detection reliability and robustness against attacks, embedding of watermarks with high-frequency spectrum, in the low frequencies of the DFT domain, is introduced in this paper and theoretical analysis of correlation based watermarking techniques with multiplicative embedding is performed. The proposed watermarking framework is successfully applied to audio signals, demonstrating its superiority with respect to both robustness and inaudibility. Experiments are conducted, in order to verify the validity of the theoretical analysis results.
Alexia Giannoula, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas
ICME2
2003 Watermarking of 3D models using principal component analysis
abstract
A novel method for 3D model watermarking robust to geometric distortions such as rotation, translation and scaling is proposed. A ternary watermark is embedded in the vertex topology of a 3D model. A transformation of the model to an invariant space is proposed prior to watermark embedding. Simulation results indicate the ability of the proposed method to deal with the aforementioned attacks giving very good results.
Andreas Kalivas, Anastasios Tefas, Ioannis Pitas
ICME2
2002 3D image watermarking robust to geometric distortions
abstract
A novel blind method for 3D image watermarking robust against geometric distortions is proposed. A ternary watermark is embedded in a grayscale or a color 3D volume. Construction of watermarks having appropriate structure enables fast and robust watermark detection even after several geometric distortions of the watermarked volume. Simulation results indicate the ability of the proposed method to deal with the aforementioned attacks. The proposed method is also robust against lossy compression up to a certain compression ratio. Experiments conducted indicate the superiority of the proposed method.
Anastasios Tefas, Giorgos Louizis, Ioannis Pitas
ICASSP1
2002 Copyright protection of 3D images using watermarks of specific spatial structure
abstract
A novel blind method for 3D image watermarking, robust against geometric distortions, is proposed. A ternary watermark is embedded in a grayscale or a color 3D volume. Construction of watermarks having appropriate structure enables fast and robust watermark detection even after several geometric distortions of the watermarked volume. Simulation results indicate the ability of the proposed method to deal with the aforementioned attacks. The proposed method is also robust against lossy compression up to a certain compression ratio.
Giorgos Louizis, Anastasios Tefas, Ioannis Pitas
ICME (2)2
2002 Watermark detection: benchmarking perspectives
abstract
Benchmarking of watermarking algorithms is a complicated task that requires examination of a set of mutually dependent performance factors (algorithm complexity, decoding/detection performance, and perceptual quality). This paper will focus on detection/decoding performance evaluation and try to summarize its basic principles. A methodology for deriving the corresponding performance metrics will also be provided.
Nikos Nikolaidis 0001, Vassilios Solachidis, Anastasios Tefas, Ioannis Pitas
ICME (2)3
2002 Face verification using elastic graph matching based on morphological signal decomposition
Anastasios Tefas, Constantine Kotropoulos, Ioannis Pitas
Signal Process.1
2001 Robust spatial image watermarking using progressive detection
abstract
A novel method for image watermarking robust to geometric distortions is proposed. A binary watermark is embedded in a grayscale or a color host image. The ability of progressive watermark detection enables fast and robust watermark detection even after several geometric distortions of the watermarked image. Simulation results indicate the ability of the proposed method to deal with the aforementioned attacks. Experiments conducted using the Stirmark benchmarking tests, indicate the superiority of the proposed method.
Anastasios Tefas, Ioannis Pitas
ICASSP1
2001 Bernoulli shift generated chaotic watermarks: theoretic investigation
abstract
The paper statistically analyzes the behaviour of chaotic watermark signals generated by n-way Bernoulli shift maps. For this purpose, a simple blind copyright protection watermarking system is considered. The analysis involves theoretical evaluation of the system detection reliability, when a correlator detector is used. The aim of the paper is twofold: (i) to introduce the n-way Bernoulli shift generated chaotic watermarks and theoretically contemplate their properties with respect to detection reliability and (ii) to establish theoretically their potential superiority against the widely used pseudorandom watermarks. Experimental verification of the theoretical analysis results is also performed.
Sofia Tsekeridou, Vassilios Solachidis, Nikos Nikolaidis 0001, Athanasios Nikolaidis, Anastasios Tefas, Ioannis Pitas
ICASSP5
2001 A survey on watermarking application scenarios and related attacks
abstract
A thorough investigation on all possible scenarios where digital imperceptible watermarking is applicable is presented. All previously proposed watermarking schemes fall to at least one of the referenced application categories. Possible attacks are divided into categories and application scenarios are presented, always referring to the watermarking parameters involved.
Athanasios Nikolaidis, Sofia Tsekeridou, Anastasios Tefas, Vassilios Solachidis
ICIP (3)3
2001 A benchmarking protocol for watermarking methods
abstract
A benchmarking system for watermarking algorithms is described. The proposed benchmarking system can be used to evaluate the performance of watermarking methods used for copyright protection, authentication, fingerprinting, etc. Although the system described is used for image watermarking, the general framework can be used, by introducing a different set of attacks, for benchmarking of video and audio data.
Nikos Nikolaidis 0001, Sofia Tsekeridou, Anastasios Tefas, Vassilios Solachidis, Athanasios Nikolaidis, Ioannis Pitas
ICIP (3)3
2001 Using Support Vector Machines to Enhance the Performance of Elastic Graph Matching for Frontal Face Authentication
abstract
A novel method for enhancing the performance of elastic graph matching in frontal face authentication is proposed. The starting point is to weigh the local similarity values at the nodes of an elastic graph according to their discriminatory power. Powerful and well-established optimization techniques are used to derive the weights of the linear combination. More specifically, we propose a novel approach that reformulates Fisher's discriminant ratio to a quadratic optimization problem subject to a set of inequality constraints by combining statistical pattern recognition and support vector machines (SVM). Both linear and nonlinear SVM are then constructed to yield the optimal separating hyperplanes and the optimal polynomial decision surfaces, respectively. The method has been applied to frontal face authentication on the M2VTS database. Experimental results indicate that the performance of morphological elastic graph matching is highly improved by using the proposed weighting technique.
Anastasios Tefas, Constantine Kotropoulos, Ioannis Pitas
IEEE Trans. Pattern Anal. Mach. Intell.1
2001 Image authentication techniques for surveillance applications
abstract
In automatic video surveillance (VS) systems, the issue of authenticating the video content is of primary importance. Given the ease with which digital images and videos can be manipulated, practically they do not have any value as legal proof, if the possibility of authenticating their content is not provided. In this paper, the problem of authenticating video surveillance image sequences is considered. After an introduction motivating the need for a watermarking-based authentication of VS sequences, a brief survey of the main watermarking-based authentication techniques is presented and the requirements that an authentication algorithm should satisfy for VS applications, are discussed. A novel algorithm which is suitable for VS visual data authentication is also presented and the results obtained by applying it to test data are discussed.
Franco Bartolini, Anastasios Tefas, Mauro Barni, Ioannis Pitas
Proc. IEEE2
2001 Statistical analysis of a watermarking system based on Bernoulli chaotic sequences
Sofia Tsekeridou, Vassilios Solachidis, Nikos Nikolaidis 0001, Athanasios Nikolaidis, Anastasios Tefas, Ioannis Pitas
Signal Process.5
2000 Face authentication by using elastic graph matching and support vector machines
abstract
A novel method for enhancing the performance of elastic graph matching in face authentication is proposed. The starting point is to weigh the local matching errors at the nodes of an elastic graph according to their discriminatory power. We propose a novel approach to discriminant analysis that re-formulates Fisher's linear discriminant ratio to a quadratic optimization problem subject to inequality constraints by combining statistical pattern recognition and support vector machines. The method is applied to frontal face authentication on the M2VTS database.
Anastasios Tefas, Constantine Kotropoulos, Ioannis Pitas
ICASSP1
2000 Using Support Vector Machines for Face Authentication Based on Elastic Graph Matching
abstract
A novel method for enhancing the performance of elastic graph matching in face authentication is proposed. Our objective is to weigh the local matching errors at the nodes of an elastic graph according to their discriminatory power. We propose a novel approach to discriminant analysis that re-formulates Fisher's linear discriminant ratio to a quadratic optimization problem subject to inequality constraints by combining statistical pattern recognition and support vector machines. The method is applied to frontal face authentication on the M2VTS database.
Anastasios Tefas, Constantine Kotropoulos, Ioannis Pitas
ICIP1
2000 Multi-Bit Image Watermarking Robust to Geometric Distortions
abstract
A novel method for multi-bit image watermarking robust to geometric distortions is proposed. A binary watermark is embedded in a grayscale or a color host image. The ability of progressive watermark detection enables fast and robust watermark detection even after several geometric distortions of the watermarked image. The embedding of a multi-bit message and its robust decoding after the watermark detection is also discussed. Simulation results indicate the ability of the proposed method to deal with the aforementioned attacks.
Anastasios Tefas, Ioannis Pitas
ICIP1
2000 Comparison of Face Verification Results on the XM2VTS Database
abstract
Presents results of the face verification contest that was organized in conjunction with International Conference on Pattern Recognition 2000. Participants had to use identical data sets from a large, publicly available multimodal database XM2VTSDB. Training and evaluation was carried out according to an a priori known protocol. Verification results of all tested algorithms have been collected and made public on the XM2VTSDB website, facilitating large scale experiments on classifier combination and fusion. Tested methods included, among others, representatives of the most common approaches to face verification -elastic graph matching, Fisher's linear discriminant and support vector machines.
Jiri Matas, Miroslav Hamouz, Kenneth Jonsson, Josef Kittler, Yongping Li, Constantine Kotropoulos, Anastasios Tefas, Ioannis Pitas, Teewoon Tan, Hong Yan 0001, Fabrizio Smeraldi, N. Capdevielle, Wulfram Gerstner, Yousri Abdeljaoued, Josef Bigün, Souheil Ben Yacoub, Eddy Mayoraz
ICPR7
2000 Image authentication using chaotic mixing systems
abstract
A novel method for image authentication is proposed. A watermark signal is embedded in a grayscale or a color host image. The watermark key controls a set of parameters of a chaotic system used for the watermark generation. The use of chaotic mixing increases the security of the proposed method and provides the additional feature of imperceptible encryption of the image owner logo in the host image. The method succeeds in detecting any alteration made in a watermarked image. The proposed method is robust in high quality lossy image compression. It provides the user not only with a measure for the authenticity of the test image but also with an image map that highlights the unaltered image regions when selective tampering has been made.
Anastasios Tefas, Ioannis Pitas
ISCAS1
2000 Morphological elastic graph matching applied to frontal face authentication under well-controlled and real conditions
Constantine Kotropoulos, Anastasios Tefas, Ioannis Pitas
Pattern Recognit.2
2000 Frontal face authentication using morphological elastic graph matching
abstract
A novel dynamic link architecture based on multiscale morphological dilation-erosion is proposed for frontal face authentication. Instead of a set of Gabor filters tuned to different orientations and scales, multiscale morphological operations are employed to yield a feature vector at each node of the reference grid. Linear projection algorithms for feature selection and automatic weighting of the nodes according to their discriminatory power succeed to increase the authentication capability of the method. The performance of the morphological dynamic link architecture is evaluated in terms of the receiver operating characteristic in the M2VTS face image database. The comparison with other frontal face authentication algorithms indicates that the morphological dynamic link architecture with discriminatory power coefficients is the best algorithm with respect to the equal error rate achieved.
Constantine Kotropoulos, Anastasios Tefas, Ioannis Pitas
IEEE Trans. Image Process.2
2000 Frontal Face Authentication Using Discriminating Grids with Morphological Feature Vectors
abstract
A novel elastic graph matching procedure based on multiscale morphological operations, the so called morphological dynamic link architecture, is developed for frontal face authentication. Fast algorithms for implementing mathematical morphology operations are presented. Feature selection by employing linear projection algorithms is proposed. Discriminatory power coefficients that weigh the matching error at each grid node are derived. The performance of morphological dynamic link architecture in frontal face authentication is evaluated in terms of the receiver operating characteristic on the M2VTS face image database. Preliminary results for face recognition using the proposed technique are also presented.
Constantine Kotropoulos, Anastasios Tefas, Ioannis Pitas
IEEE Trans. Multim.2
1999 Compensating for variable recording conditions in frontal face authentication algorithms
abstract
This paper addresses the problem of compensating for variable recording conditions such as changes in illumination, scale differences, and varying face position. It is well known that the performance of any face authentication/recognition algorithm deteriorates significantly in the presence of the aforementioned conditions as well as the expression variations. The use of simple and powerful pre-processing techniques aiming at compensating for variable recording conditions prior to the application of any authentication algorithm is proposed. It is shown that such an approach overcomes indeed the image variations and guarantees an almost stable performance for the Morphological Dynamic Link Architecture developed within the European research project M2VTS.
Anastasios Tefas, Yann Menguy, Constantine Kotropoulos, Gaël Richard, Ioannis Pitas, Philip Lockwood
ICASSP1
1998 Variants of Dynamic Link Architecture Based on Mathematical Morphology for Frontal Face Authentication
abstract
Two novel variants of dynamic link architecture that are based on mathematical morphology and incorporate coefficients which weigh the contribution of each node in elastic graph matching according to its discriminatory power are developed. They are the so called Morphological Dynamic Link Architecture and the Morphological Signal Decomposition-Dynamic Lint Architecture. The proposed variants are tested for face authentication in a cooperative scenario where the candidates claim an identity to be checked. Their performance is evaluated in terms of their receiver operating characteristic and the equal error rate achieved in M2VTS database. An equal error rate in the range 3.7-6.8% is reported.
Anastasios Tefas, Constantine Kotropoulos, Ioannis Pitas
CVPR1
1998 Face Verification based on Morphological Shape Decomposition
Anastasios Tefas, Constantine Kotropoulos, Ioannis Pitas
FG1
1998 Face authentication using variants of elastic graph matching based on mathematical morphology that incorporate local discriminant coefficients
abstract
Two novel variants of dynamic link architecture that are based on mathematical morphology and incorporate local coefficients which weigh the contribution of each node according to its discriminatory power in elastic graph matching are proposed, namely, the morphological dynamic link architecture and the morphological signal decomposition-dynamic link architecture. They are tested for face authentication in a cooperative scenario where the candidates claim an identity to be checked. Their performance is evaluated in terms of their receiver operating characteristics and the equal error rate achieved in the M2VTS database. An equal error rate of 6.6%-6.8% is reported.
Constantine Kotropoulos, Anastasios Tefas, Ioannis Pitas
ICASSP2
1998 Frontal Face Authentication using Variants of Dynamic Link Matching Based on Mathematical Morphology
abstract
Two variants of dynamic link matching based on mathematical morphology are developed and tested for frontal face authentication, namely, the morphological dynamic link architecture and the morphological signal decomposition-dynamic link architecture. Local coefficients which weigh the contribution of each node in elastic graph matching according to its discriminatory power are derived. The performance of the proposed algorithms is evaluated in terms of their receiver operating characteristic and the equal error rate (EER) achieved in the M2VTS database. The comparison with other frontal face authentication algorithms developed within M2VTS project indicates that morphological dynamic link architecture with discriminatory power coefficients is ranked as the best algorithm in terms of the EER.
Constantine Kotropoulos, Anastasios Tefas, Ioannis Pitas
ICIP (1)2