EDBT 2026 Demo / reviewers in the wild / expert
Alexandros Iosifidis
dblp:01/9539
· DBLP profile ↗
164ranked-venue papers
44as first author
68since 2021 · last 2027
0000-0003-4807-1345ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 98 · 24 first-author · 43 since 2021Graphics, computer vision, multimedia, augmented reality and games · 65 · 17 first-author · 19 since 2021Computer networks · 9 · 9 since 2021Systems, architecture and hardware · 4 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Security and privacy · 2 · 2 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Data-driven generative digital twins for wind-farm flows from highly sparse measurementsabstractReconstructing high-fidelity flow fields from sparse measurements remains a central challenge in applied sensing and real-time flow control. This work addresses the problem in a real-world setting, focusing on turbulent wind-farm flows. To span a range of reconstruction paradigms, we compare three approaches: gappy proper orthogonal decomposition (GPOD), a linear reduced-order baseline; the shallow recurrent decoder (SHRED), a state-of-the-art deep learning method; and a proposed observation-guided generative framework (GGenAI). The methods are systematically evaluated on a wind-farm dataset across sensor densities ranging from 0.12% to 6.65% of the spatial domain. The results reveal a clear performance hierarchy. In the extremely sparse regime ( < 0.49% coverage), GPOD yields relatively low global reconstruction error; however, qualitative analysis indicates that this performance reflects recovery of the mean field rather than instantaneous dynamics. In this same regime, SHRED exhibits behavior similar to GPOD. As sensor coverage increases (around 0.49%), SHRED begins to recover meaningful coarse flow structures. However, the reconstructions remain overly smooth and fail to capture instantaneous turbulent features. In contrast, GGenAI undergoes an earlier transition toward instantaneous-like reconstruction, and at higher sensor densities (beyond approximately 1% coverage) it consistently outperforms both baselines. To further assess physical consistency, we introduce a POD-based reduced-order reference field representing the dominant energy-containing flow structures. GGenAI exhibits improved agreement with this reduced-order reference, emphasizing that reconstruction quality must be interpreted in relation to relevant spatial scales. Overall, the study demonstrates that our GGenAI approach extends to complex turbulent flows, offering a promising pathway toward data-driven digital twins. Sajad Salavati, Christoffer Hansen, Henrik Karstoft, Alexandros Iosifidis, Mahdi Abkar |
Expert Syst. Appl. | 4 |
| 2026 | Continual low-rank scaled dot-product attentionabstract• A new Continual Inference Transformer that explores low-rank attention is proposed. • We propose two new ways to compute data-driven landmarks for the low-rank attention. • The proposed model leads to up to three orders of magnitude computation reduction. Transformers are widely used for their ability to capture data relations in sequence processing, with great success for a wide range of static tasks. However, the computational and memory footprint of their main component, i.e., the Scaled Dot-product Attention, is commonly overlooked. This makes their adoption infeasible in applications involving stream data processing with constraints in response latency, computational and memory resources. Some works have proposed methods to lower the computational cost of Transformers by using low-rank approximations, sparsity in attention, and efficient formulations for Continual Inference. In this paper, we introduce a new formulation of the Scaled Dot-product Attention based on the Nyström approximation that is suitable for Continual Inference. In experiments on Online Audio Classification and Online Action Detection tasks, the proposed Continual Scaled Dot-product Attention can lower the number of operations by up to three orders of magnitude compared to the original Transformers while retaining the predictive performance of competing models. Ginés Carreto Picón, Illia Oleksiienko, Lukas Hedegaard, Arian Bakhtiarnia, Alexandros Iosifidis |
Neural Networks | 5 |
| 2026 | Quasi black-box variational inference for Bayesian learning
Martin Magris, Mostafa Shabani, Alexandros Iosifidis |
Pattern Recognit. | 3 |
| 2026 | Corrigendum to "Augmented bilinear network for incremental multi-stock time-series classification" [Pattern Recognit. 141 (2023) 109604]abstractThe authors regret that an implementation error was identified in the trading simulation presented in Section 4.5 of the original paper. Specifically, the computation of trade returns incorrectly reversed the use of bid and ask prices, which resulted in inflated profitability estimates. This Corrigendum describes the correction, presents the revised results, and details a new simulation study incorporating a signal-strength–based trading rule to evaluate model robustness. The authors apologise for any inconvenience or misunderstanding caused. Mostafa Shabani, Dat Thanh Tran, Juho Kanniainen, Alexandros Iosifidis |
Pattern Recognit. | 4 |
| 2025 | Delay Bound Relaxation with Deep Learning-based Haptic Estimation for Tactile InternetabstractHaptic teleoperation typically demands sub-millisecond latency and ultra-high reliability (99.999%) in Tactile Internet. At a 1 kHz haptic signal sampling rate, this translates into an extremely high packet transmission rate, posing significant challenges for timely delivery and introducing substantial complexity and overhead in radio resource allocation. To address this critical challenge, we introduce a novel DL model that estimates force feedback using multi-modal input, i.e. both force measurements from the remote side and local operator motion signals. The DL model can capture complex temporal features of haptic time-series with the use of CNN and LSTM layers, followed by a transformer encoder, and autoregressively produce a highly accurate estimation of the next force values for different teleoperation activities. By ensuring that the estimation error is within a predefined threshold, the teleoperation system can safely relax its strict delay requirements. This enables the batching and transmission of multiple haptic packets within a single resource block, improving resource efficiency and facilitating scheduling in resource allocation. Through extensive simulations, we evaluated network performance in terms of reliability and capacity. Results show that, for both dynamic and rigid object interactions, the proposed method increases the number of reliably served users by up to 66%. Georgios Kokkinis, Alexandros Iosifidis, Qi Zhang 0013 |
GLOBECOM | 2 |
| 2025 | Deep Reinforcement Learning-Based Video-Haptic Radio Resource Slicing in Tactile InternetabstractEnabling video-haptic radio resource slicing in the Tactile Internet requires a sophisticated strategy to meet the distinct requirements of video and haptic data, ensure their synchronized transmission, and address the stringent latency demands of haptic feedback. This paper introduces a Deep Reinforcement Learning-based radio resource slicing framework that addresses video-haptic teleoperation challenges by dynamically balancing radio resources between the video and haptic modalities. The proposed framework employs a refined reward function that considers latency, packet loss, data rate, and the synchronization requirements of both modalities to optimize resource allocation. By catering to the specific service requirements of video-haptic teleoperation, the proposed framework achieves up to a 25 % increase in user satisfaction over existing methods, while maintaining effective resource slicing with execution intervals up to 50 ms. Georgios Kokkinis, Alexandros Iosifidis, Qi Zhang 0013 |
ICC | 2 |
| 2025 | Dynamic Semantic Compression for CNN Inference in Multi-Access Edge Computing: A Graph Reinforcement Learning-Based AutoencoderabstractThis paper studies the computational offloading of CNN inference in dynamic multi-access edge computing (MEC) networks. To address the uncertainties in communication time and edge servers’ available capacity, we propose a novel semantic compression method, autoencoder-based CNN architecture (AECNN), for effective semantic extraction and compression in partial offloading. In the semantic encoder, we introduce a feature compression module based on the channel attention mechanism in CNNs, to compress intermediate data by selecting the most informative features. Additionally, to further reduce communication overhead, we leverage entropy encoding to remove the statistical redundancy in the compressed data. In the semantic decoder, we design a lightweight decoder to reconstruct the intermediate data through learning from the received compressed data to improve accuracy. To effectively trade-off communication, computation, and inference accuracy, we design a reward function and formulate the offloading problem of CNN inference as a maximization problem with the goal of maximizing the average inference accuracy and throughput over the long term. To address this maximization problem, we propose a graph reinforcement learning-based AECNN (GRL-AECNN) method, which outperforms existing works DROO-AECNN, GRL-BottleNet++ and GRL-DeepJSCC under different dynamic scenarios. This highlights the advantages of GRL-AECNN in offloading decision-making for CNN inference tasks in dynamic MEC. Nan Li 0064, Alexandros Iosifidis, Qi Zhang 0013 |
IEEE Trans. Wirel. Commun. | 2 |
| 2024 | Accurate Gigapixel Crowd Counting by Iterative Zooming and RefinementabstractThe increasing prevalence of gigapixel resolutions has presented new challenges for crowd counting. Such resolutions are far beyond the memory and computation limits of current GPUs, and available deep neural network architectures and training procedures are not designed for such massive inputs. Although several methods have been proposed to address these challenges, they are either limited to downsampling the input image to a small size, or borrowing from other gigapixel tasks, which are not tailored for crowd counting. In this paper, we propose a novel method called GigaZoom, which iteratively zooms into the densest areas of the image and refines coarser density maps with finer details. We show that GigaZoom obtains the state-of-the-art for gigapixel crowd counting and improves the accuracy of the next best method by 42%. Arian Bakhtiarnia, Qi Zhang 0013, Alexandros Iosifidis |
ICASSP | 3 |
| 2024 | Uimt: A Framework for Improving Unimodal Inference via Multimodal TrainingabstractThe field of multimodal learning is developing rapidly, with emergence of many novel models and applications. Still, works proposing unimodal and multimodal models are generally disjoint, and either focus on fully-unimodal or fully-multimodal scenarios. Nevertheless, oftentimes in real-world applications data of multiple modalities are available during training while only one of them can be utilized during inference due to associated computational costs, or complexity of utilizing additional sensors. In this work, we develop a framework for improving inference of arbitrary unimodal models with multimodal training, without incurring any additional computational cost at inference time, but benefiting from the advantages of multimodal training. We show that our framework is applicable to different architecture types: transformers, 3D CNNs, and 2D+1D CNNs. To showcase this generality we evaluate our approach on tasks of dynamic hand gesture recognition based on RGB and Depth, audiovisual emotion recognition based, and audio-video-text based sentiment analysis. Our approach consistently outperforms the conventionally trained unimodal counterparts. We additionally investigate how within our framework training of multimodal models can benefit from unimodal, modality-specific learning signals. Utilizing the same variety of architectures as mentioned above, we show how models trained with additional supervision from each isolated modality outperform a multimodal-only counterpart. Kateryna Chumachenko, Alexandros Iosifidis, Moncef Gabbouj |
ICIP | 2 |
| 2024 | MAVAD: Audio-Visual Dataset and Method for Anomaly Detection in Traffic VideosabstractThis paper introduces the first audio-visual dataset for traffic anomaly detection called MAVAD, taken from real-world scenes, with a diverse range of illumination conditions. In addition, a novel anomaly detection method is proposed which combines visual and audio features extracted from video sequences by means of cross-attention. We demonstrate that the addition of audio improves anomaly detection performance by up to 5.2%. Moreover, the impact of image anonymization is evaluated, showing only a minor decrease in performance averaging at 1.7%. Blazej Leporowski, Arian Bakhtiarnia, Nicole Bonnici, Adrian Muscat, Luca Zanella, Yiming Wang 0002, Alexandros Iosifidis |
ICIP | 7 |
| 2024 | Uncertainty-Aware AB3DMOT by Variational 3D Object DetectionabstractAutonomous driving needs to rely on high-quality 3D object detection to ensure safe navigation in the world. Uncertainty estimation is an effective tool to provide statistically accurate predictions, while the associated detection uncertainty can be used to implement a more safe navigation protocol or include the user in the loop. In this paper, we propose a Variational Neural Network-based TANet 3D object detector to generate 3D object detections with uncertainty and introduce these detections to an uncertainty-aware AB3DMOT tracker. This is done by applying a linear transformation to the estimated uncertainty matrix, which is subsequently used as a measurement noise for the adopted Kalman filter. We implement two ways to estimate output uncertainty, i.e., internally, by computing the variance of the CNN outputs and then propagating the uncertainty through the post-processing, and externally, by associating the final predictions of different samples and computing the covariance of each predicted box. In experiments, we show that the external uncertainty estimation leads to better results, outperforming both internal uncertainty estimation and classical tracking approaches. Furthermore, we propose a method to initialize the Variational 3D object detector with a pretrained TANet model, which leads to the best performing models. Illia Oleksiienko, Alexandros Iosifidis |
ICIP | 2 |
| 2024 | Superpixel-based Anomaly Detection for Irregular Textures with a Focus on Pixel-level AccuracyabstractRecent anomaly detection methods achieve high performance on commonly used image and pixel-level metrics. However, due to the imbalance in the number of normal and abnormal pixels commonly encountered in anomaly detection problems, commonly adopted pixel-level performance metrics cannot effectively evaluate model performance. This paper proposes a novel approach for anomaly detection within the irregular texture domain, focusing on pixel-level accuracy metrics suitable for such imbalanced problems. The proposed Superpixel-based Coupled-hypersphere-based Feature Adaptation (Sp-CFA) method leverages the intermediate adaptive representation of superpixels to enable superior pixel-level anomaly detection performance. We demonstrate superior performance over the irregular texture classes within the MVTec AD benchmark dataset, KSDD2 dataset, and an X-ray dataset of manufactured fibrous products. Mehdi Rafiei, Toby P. Breckon, Alexandros Iosifidis |
IJCNN | 3 |
| 2024 | Vpit: real-time embedded single object 3D tracking using voxel pseudo imagesabstractAbstract In this paper, we propose a novel voxel-based 3D single object tracking (3D SOT) method called Voxel Pseudo Image Tracking (VPIT). VPIT is the first method that uses voxel pseudo images for 3D SOT. The input point cloud is structured by pillar-based voxelization, and the resulting pseudo image is used as an input to a 2D-like Siamese SOT method. The pseudo image is created in the Bird’s-eye View (BEV) coordinates; and therefore, the objects in it have constant size. Thus, only the object rotation can change in the new coordinate system and not the object scale. For this reason, we replace multi-scale search with a multi-rotation search, where differently rotated search regions are compared against a single target representation to predict both position and rotation of the object. Experiments on KITTI [1] Tracking dataset show that VPIT is the fastest 3D SOT method and maintains competitive Success and Precision values. Application of a SOT method in a real-world scenario meets with limitations such as lower computational capabilities of embedded devices and a latency-unforgiving environment, where the method is forced to skip certain data frames if the inference speed is not high enough. We implement a real-time evaluation protocol and show that other methods lose most of their performance on embedded devices; while, VPIT maintains its ability to track the object. Illia Oleksiienko, Paraskevi Nousi, Nikolaos Passalis, Anastasios Tefas, Alexandros Iosifidis |
Neural Comput. Appl. | 5 |
| 2024 | Manifold Gaussian Variational Bayes on the Precision MatrixabstractWe propose an optimization algorithm for variational inference (VI) in complex models. Our approach relies on natural gradient updates where the variational space is a Riemann manifold. We develop an efficient algorithm for gaussian variational inference whose updates satisfy the positive definite constraint on the variational covariance matrix. Our manifold gaussian variational Bayes on the precision matrix (MGVBP) solution provides simple update rules, is straightforward to implement, and the use of the precision matrix parameterization has a significant computational advantage. Due to its black-box nature, MGVBP stands as a ready-to-use solution for VI in complex models. Over five data sets, we empirically validate our feasible approach on different statistical and econometric models, discussing its performance with respect to baseline methods. Martin Magris, Mostafa Shabani, Alexandros Iosifidis |
Neural Comput. | 3 |
| 2024 | Structured pruning adaptersabstractAdapters are a parameter-efficient alternative to fine-tuning, which augment a frozen base network to learn new tasks. Yet, the inference of the adapted model is often slower than the corresponding fine-tuned model. To improve on this, we introduce the concept of Structured Pruning Adapters (SPAs), a family of compressing, task-switching network adapters, that accelerate and specialize networks using tiny parameter sets and structured pruning. Specifically, we propose the Structured Pruning Low-rank Adapter (SPLoRA) and the Structured Pruning Residual Adapter (SPPaRA) and evaluate them on a suite of pruning methods, architectures, and image recognition benchmarks. Compared to regular structured pruning with fine-tuning, SPLoRA improves image recognition accuracy by 6.9% on average for ResNet50 while using half the parameters at 90% pruned weights. Alternatively, a SPLoRA augmented model can learn adaptations with 17× fewer parameters at 70% pruning with 1.6% lower accuracy. For ViT-b/16 models, SPLoRA improves accuracy by an average of 43%-points at 75% pruned weights while learning 6.8× fewer parameters. Our experimental code and Python library of adapters are available at www.github.com/lukashedegaard/structured-pruning-adapters. Lukas Hedegaard, Aman Alok, Juby Jose, Alexandros Iosifidis |
Pattern Recognit. | 4 |
| 2024 | Reducing redundancy in the bottleneck representation of autoencodersabstractAutoencoders (AEs) are a type of unsupervised neural networks, which can be used to solve various tasks, e.g., dimensionality reduction, image compression, and image denoising. An AE has two goals: (i) compress the original input to a low-dimensional space at the bottleneck of the network topology using an encoder, (ii) reconstruct the input from the representation at the bottleneck using a decoder. Both encoder and decoder are optimized jointly by minimizing a distortion-based loss which implicitly forces the model to keep only the information in input data required to reconstruct them and to reduce redundancies. In this paper, we propose a scheme to explicitly penalize feature redundancies in the bottleneck representation. To this end, we propose an additional loss term, based on the pairwise covariances of the network units, which complements the data reconstruction loss forcing the encoder to learn a more diverse and richer representation of the input. We tested our approach across different tasks, namely dimensionality reduction, image compression, and image denoising. Experimental results show that the proposed loss leads consistently to superior performance compared to using the standard AE loss. Firas Laakom, Jenni Raitoharju, Alexandros Iosifidis, Moncef Gabbouj |
Pattern Recognit. Lett. | 3 |
| 2023 | WLD-Reg: A Data-Dependent Within-Layer Diversity RegularizerabstractNeural networks are composed of multiple layers arranged in a hierarchical structure jointly trained with a gradient-based optimization, where the errors are back-propagated from the last layer back to the first one. At each optimization step, neurons at a given layer receive feedback from neurons belonging to higher layers of the hierarchy. In this paper, we propose to complement this traditional 'between-layer' feedback with additional 'within-layer' feedback to encourage the diversity of the activations within the same layer. To this end, we measure the pairwise similarity between the outputs of the neurons and use it to model the layer's overall diversity. We present an extensive empirical study confirming that the proposed approach enhances the performance of several state-of-the-art neural network models in multiple tasks. The code is publically available at https://github.com/firasl/AAAI-23-WLD-Reg. Firas Laakom, Jenni Raitoharju, Alexandros Iosifidis, Moncef Gabbouj |
AAAI | 3 |
| 2023 | Dynamic Split Computing for Efficient Deep EDGE IntelligenceabstractDeploying deep neural networks (DNNs) on IoT and mobile devices is a challenging task due to their limited computational resources. Thus, demanding tasks are often entirely offloaded to edge servers which can accelerate inference, however, it also causes communication cost and evokes privacy concerns. In addition, this approach leaves the computational capacity of end devices unused. Split computing is a paradigm where a DNN is split into two sections; the first section is executed on the end device, and the output is transmitted to the edge server where the final section is executed. Here, we introduce dynamic split computing, where the optimal split location is dynamically selected based on the state of the communication channel. By using natural bottlenecks that already exist in modern DNN architectures, dynamic split computing avoids retraining and hyperparameter optimization, and does not have any negative impact on the final accuracy of DNNs. Through extensive experiments, we show that dynamic split computing achieves faster inference in edge computing environments where the data rate and server load vary over time. Arian Bakhtiarnia, Nemanja Milosevic, Qi Zhang 0013, Dragana Bajovic, Alexandros Iosifidis |
ICASSP | 5 |
| 2023 | Attention-Based Feature Compression for CNN Inference Offloading in Edge ComputingabstractThis paper studies the computational offloading of CNN inference in device-edge co-inference systems. Inspired by the emerging paradigm semantic communication, we propose a novel autoencoder-based CNN architecture (AECNN), for effective feature extraction at end-device. We design a feature compression module based on the channel attention method in CNN, to compress the intermediate data by selecting the most important features. To further reduce communication overhead, we can use entropy encoding to remove the statistical redundancy in the compressed data. At the receiver, we design a lightweight decoder to reconstruct the intermediate data through learning from the received compressed data to improve accuracy. To fasten the convergence, we use a step-by-step approach to train the neural networks obtained based on ResNet-50 architecture. Experimental results show that AECNN can compress the intermediate data by more than 256 × with only about 4% accuracy loss, which outperforms the state-of-the-art work, BottleNet++. Compared to offloading inference task directly to edge server, AECNN can complete inference task earlier, in particular, under poor wireless channel condition, which highlights the effectiveness of AECNN in guaranteeing higher accuracy within time constraint. Nan Li 0064, Alexandros Iosifidis, Qi Zhang 0013 |
ICC | 2 |
| 2023 | Continual Transformers: Redundancy-Free Attention for Online Inference
Lukas Hedegaard, Arian Bakhtiarnia, Alexandros Iosifidis |
ICLR | 3 |
| 2023 | PromptMix: Text-to-image diffusion models enhance the performance of lightweight networksabstractMany deep learning tasks require annotations that are too time consuming for human operators, resulting in small dataset sizes. This is especially true for dense regression problems such as crowd counting which requires the location of every person in the image to be annotated. Techniques such as data augmentation and synthetic data generation based on simulations can help in such cases. In this paper, we introduce PromptMix, a method for artificially boosting the size of existing datasets, that can be used to improve the performance of lightweight networks. First, synthetic images are generated in an end-to-end data-driven manner, where text prompts are extracted from existing datasets via an image captioning deep network, and subsequently introduced to text-to-image diffusion models. The generated images are then annotated using one or more high-performing deep networks, and mixed with the real dataset for training the lightweight network. By extensive experiments on five datasets and two tasks, we show that PromptMix can significantly increase the performance of lightweight networks by up to 26%. Arian Bakhtiarnia, Qi Zhang 0013, Alexandros Iosifidis |
IJCNN | 3 |
| 2023 | Class-Specific Variational Auto-Encoder for Content-Based Image RetrievalabstractUsing a discriminative representation obtained by supervised deep learning methods showed promising results on diverse Content-Based Image Retrieval (CBIR) problems. However, existing methods exploiting labels during training try to discriminate all available classes, which is not ideal in cases where the retrieval problem focuses on a class of interest. In this paper, we propose a regularized loss for Variational Auto-Encoders (VAEs) forcing the model to focus on a given class of interest. As a result, the model learns to discriminate the data belonging to the class of interest from any other possibility, making the learnt latent space of the VAE suitable for class-specific retrieval tasks. The proposed Class-Specific Variational Auto-Encoder (CS-VAE) is evaluated on three public and one custom datasets, and its performance is compared with that of three related VAE-based methods. Experimental results show that the proposed method outperforms its competition in both in-domain and out-of-domain retrieval problems. Mehdi Rafiei, Alexandros Iosifidis |
IJCNN | 2 |
| 2023 | Recognition of Defective Mineral Wool Using Pruned ResNet ModelsabstractMineral wool production is a non-linear process that makes it hard to control the final quality. Therefore, having a nondestructive method to analyze the product quality and recognize defective products is critical. For this purpose, we developed a visual quality control system for mineral wool. X-ray images of wool specimens were collected to create a training set of defective and non-defective samples. Afterward, we developed several recognition models based on the ResNet architecture to find the most efficient model. In order to have a light-weight and fast inference model for real-life applicability, two structural pruning methods are applied to the classifiers. Considering the low quantity of the dataset, cross-validation and augmentation methods are used during the training. As a result, we obtained a model with more than 98% accuracy, which in comparison to the current procedure used at the company, it can recognize 20% more defective products. Mehdi Rafiei, Dat Thanh Tran, Alexandros Iosifidis |
INDIN | 3 |
| 2023 | Improving Online non-destructive Moisture Content Estimation using Data Augmentation by Feature Space Interpolation with Variational AutoencodersabstractData augmentation techniques have proven to be highly effective for many types of problems. However, the development of data augmentation for continuous input-output mappings in regression problems has not received much attention. Insufficient training data remains a significant challenge in machine learning, especially for industrial applications, as the cost of experimentation on the production line can be prohibitively expensive. This acts as a barrier to adoption of machine learning methods in industrial applications. In this study, we propose a data augmentation method called feature space interpolation for continuous input-output regression problems based on discontinuous data sets with clear gaps in the data. The proposed method is applied to a dataset of industrial drying of bulky filter media products. It is shown, that augmenting the original dataset by generated synthetic data points in the gap of the dataset by interpolation in the latent space of a well-trained variational autoencoder (VAE) can improve the performance of state-of-the-art of bulky filter media product moisture content estimation models, as measured by the mean absolute error and mean squared error by 4.82% and 6.32% respectively, and outperforms baseline generative data augmentation methods such as latent space sampling from VAEs. Christian Wewer, Alexandros Iosifidis |
INDIN | 2 |
| 2023 | Learning Distinct Features Helps, Provably
Firas Laakom, Jenni Raitoharju, Alexandros Iosifidis, Moncef Gabbouj |
ECML/PKDD (2) | 3 |
| 2023 | Predicting the trading behavior of socially connected investors: Graph neural network approach with implications to market surveillanceabstractDespite the success of machine learning models, the literature lacks their applications to identify the exploitation of non-public information. We address this gap by developing a tool, which ranks investors based on their suspiciousness. We achieve this by predicting the future trading decisions of retail investors based on their social connections in an insider network. Particularly, the high predictability of the investor’s trading behavior with data on her/his social neighborhood can indicate that she/he takes advantage of her/his social connections and trades on non-public information. Our system captures complex and cyclical patterns in investor social networks with Graph Networks trained with full investor-level transaction data. We show that using data on social neighborhoods significantly improves the model’s performance in predicting investors’ trading behavior. The results are robust regarding the different groups of investors, trading windows, trading directions, and outliers. We demonstrate the tool by ranking suspicious investors and companies, with 12 out of 153 companies having a statistically significant concentration of directors (p-value < 0.05) with suspicious trading behavior. The system provides regulators with a valuable tool for prioritizing market surveillance efforts. Kestutis Baltakys, Margarita Baltakiene, Negar Heidari, Alexandros Iosifidis, Juho Kanniainen |
Expert Syst. Appl. | 4 |
| 2023 | Qualifying and raising anti-money laundering alarms with deep learning
Rasmus Jensen, Alexandros Iosifidis |
Expert Syst. Appl. | 2 |
| 2023 | Predicting the state of synchronization of financial time series using cross recurrence plotsabstractAbstract Cross-correlation analysis is a powerful tool for understanding the mutual dynamics of time series. This study introduces a new method for predicting the future state of synchronization of the dynamics of two financial time series. To this end, we use the cross recurrence plot analysis as a nonlinear method for quantifying the multidimensional coupling in the time domain of two time series and for determining their state of synchronization. We adopt a deep learning framework for methodologically addressing the prediction of the synchronization state based on features extracted from dynamically sub-sampled cross recurrence plots. We provide extensive experiments on several stocks, major constituents of the S &P100 index, to empirically validate our approach. We find that the task of predicting the state of synchronization of two time series is in general rather difficult, but for certain pairs of stocks attainable with very satisfactory performance (84% F1-score, on average). Mostafa Shabani, Martin Magris, George Tzagkarakis, Juho Kanniainen, Alexandros Iosifidis |
Neural Comput. Appl. | 5 |
| 2023 | Continual spatio-temporal graph convolutional networksabstractGraph-based reasoning over skeleton data has emerged as a promising approach for human action recognition. However, the application of prior graph-based methods, which predominantly employ whole temporal sequences as their input, to the setting of online inference entails considerable computational redundancy. In this paper, we tackle this issue by reformulating the Spatio-Temporal Graph Convolutional Neural Network as a Continual Inference Network, which can perform step-by-step predictions in time without repeat frame processing. To evaluate our method, we create a continual version of ST-GCN, CoST-GCN, alongside two derived methods with different self-attention mechanisms, CoAGCN and CoS-TR. We investigate weight transfer strategies and architectural modifications for inference acceleration, and perform experiments on the NTU RGB+D 60, NTU RGB+D 120, and Kinetics Skeleton 400 datasets. Retaining similar predictive accuracy, we observe up to 109× reduction in time complexity, on-hardware accelerations of 26×, and reductions in maximum allocated memory of 52% during online inference. Lukas Hedegaard, Negar Heidari, Alexandros Iosifidis |
Pattern Recognit. | 3 |
| 2023 | Augmented bilinear network for incremental multi-stock time-series classificationabstractDeep Learning models have become dominant in tackling financial time-series analysis problems, overturning conventional machine learning and statistical methods. Most often, a model trained for one market or security cannot be directly applied to another market or security due to differences inherent in the market conditions. In addition, as the market evolves over time, it is necessary to update the existing models or train new ones when new data is made available. This scenario, which is inherent in most financial forecasting applications, naturally raises the following research question: How to efficiently adapt a pre-trained model to a new set of data while retaining performance on the old data, especially when the old data is not accessible? In this paper, we propose a method to efficiently retain the knowledge available in a neural network pre-trained on a set of securities and adapt it to achieve high performance in new ones. In our method, the prior knowledge encoded in a pre-trained neural network is maintained by keeping existing connections fixed, and this knowledge is adjusted for the new securities by a set of augmented connections, which are optimized using the new data. The auxiliary connections are constrained to be of low rank. This not only allows us to rapidly optimize for the new task but also reduces the storage and run-time complexity during the deployment phase. The efficiency of our approach is empirically validated in the stock mid-price movement prediction problem using a large-scale limit order book dataset. Experimental results show that our approach enhances prediction performance as well as reduces the overall number of network parameters. Mostafa Shabani, Dat Thanh Tran, Juho Kanniainen, Alexandros Iosifidis |
Pattern Recognit. | 4 |
| 2023 | Graph-embedded subspace support vector data descriptionabstractIn this paper, we propose a novel subspace learning framework for one-class classification. The proposed framework presents the problem in the form of graph embedding. It includes the previously proposed subspace one-class techniques as its special cases and provides further insight on what these techniques actually optimize. The framework allows to incorporate other meaningful optimization goals via the graph preserving criterion and reveals a spectral solution and a spectral regression-based solution as alternatives to the previously used gradient-based technique. We combine the subspace learning framework iteratively with Support Vector Data Description applied in the subspace to formulate Graph-Embedded Subspace Support Vector Data Description. We experimentally analyzed the performance of newly proposed different variants. We demonstrate improved performance against the baselines and the recently proposed subspace learning methods for one-class classification. Fahad Sohrab, Alexandros Iosifidis, Moncef Gabbouj, Jenni Raitoharju |
Pattern Recognit. | 2 |
| 2022 | Continual 3D Convolutional Neural Networks for Real-time Processing of Videos
Lukas Hedegaard, Alexandros Iosifidis |
ECCV (4) | 2 |
| 2022 | Graph Reinforcement Learning-based CNN Inference Offloading in Dynamic Edge ComputingabstractThis paper studies the computational offloading of CNN inference in dynamic multi-access edge computing (MEC) networks. To address the uncertainties in communication time and Edge servers' available capacity, we use early-exit mechanism to terminate the computation earlier to meet the deadline of inference tasks. We design a reward function to trade off the communication, computation and inference accuracy, and formu-late the offloading problem of CNN inference as a maximization problem with the goal of maximizing the average inference accuracy and throughput in long term. To solve the maxi-mization problem, we propose a graph reinforcement learning-based early-exit mechanism (GRLE), which outperforms the state-of-the-art work, deep reinforcement learning-based online offloading (DROO) and its enhanced method, DROO with early-exit mechanism (DROOE), under different dynamic scenarios. The experimental results show that G RLE achieves the average accuracy up to 3.41 x over graph reinforcement learning (GRL) and 1.45x over DROOE, which shows the advantages of GRLE for offloading decision-making in dynamic MEC. Nan Li 0064, Alexandros Iosifidis, Qi Zhang 0013 |
GLOBECOM | 2 |
| 2022 | Distributed Deep Learning Inference Acceleration using Seamless Collaboration in Edge ComputingabstractThis paper studies inference acceleration using distributed convolutional neural networks (CNNs) in collaborative edge computing. To ensure inference accuracy in inference task partitioning, we consider the receptive-field when performing segment-based partitioning. To maximize the parallelization between the communication and computing processes, thereby minimizing the total inference time of an inference task, we design a novel task collaboration scheme in which the overlapping zone of the sub-tasks on secondary edge servers (ESs) is executed on the host ES, named as HALP. We further extend HALP to the scenario of multiple tasks. Experimental results show that HALP can accelerate CNN inference in VGG-16 by 1.7-2.0x for a single task and 1.7-1.8x for 4 tasks per batch on GTX 1080TI and JETSON AGX Xavier, which outperforms the state-of-the-art work MoDNN. Moreover, we evaluate the service reliability under time-variant channel, which shows that HALP is an effective solution to ensure high service reliability with strict service deadline. Nan Li 0064, Alexandros Iosifidis, Qi Zhang 0013 |
ICC | 2 |
| 2022 | Receptive Field-based Segmentation for Distributed CNN Inference Acceleration in Collaborative Edge ComputingabstractThis paper studies inference acceleration using distributed convolutional neural networks (CNNs) in collaborative edge computing network. To avoid inference accuracy loss in inference task partitioning, we propose receptive field-based segmentation (RFS). To reduce the computation time and communication overhead, we propose a novel collaborative edge computing using fused-layer parallelization to partition a CNN model into multiple blocks of convolutional layers. In this scheme, the collaborative edge servers (ESs) only need to exchange small fraction of the sub-outputs after computing each fused block. In addition, to find the optimal solution of partitioning a CNN model into multiple blocks, we use dynamic programming, named as dynamic programming for fused-layer parallelization (DPFP). The experimental results show that DPFP can accelerate inference of VGG-16 up to 73% compared with the pre-trained model, which outperforms the existing work MoDNN in all tested scenarios. Moreover, we evaluate the service reliability of DPFP under time-variant channel, which shows that DPFP is an effective solution to ensure high service reliability with strict service deadline. Nan Li 0064, Alexandros Iosifidis, Qi Zhang 0013 |
ICC | 2 |
| 2022 | Self-attention fusion for audiovisual emotion recognition with incomplete dataabstractIn this paper, we consider the problem of multi-modal data analysis with a use case of audiovisual emotion recognition. We propose an architecture capable of learning from raw data and describe three variants of it with distinct modality fusion mechanisms. While most of the previous works consider the ideal scenario of presence of both modalities at all times during inference, we evaluate the robustness of the model in the unconstrained settings where one modality is absent or noisy, and propose a method to mitigate these limitations in a form of modality dropout. Most importantly, we find that following this approach not only improves performance drastically under the absence/noisy representations of one modality, but also improves the performance in a standard ideal setting, outperforming the competing methods. Kateryna Chumachenko, Alexandros Iosifidis, Moncef Gabbouj |
ICPR | 2 |
| 2022 | Generalized Reference Kernel for One-class ClassificationabstractIn this paper, we formulate a new generalized reference kernel hoping to improve the original base kernel using a set of reference vectors. Depending on the selected reference vectors, our formulation shows similarities to approximate kernels, random mappings, and Non-linear Projection Trick. Focusing on small-scale one-class classification, our analysis and experimental results show that the new formulation provides approaches to regularize, adjust the rank, and incorporate additional information into the kernel itself, leading to improved one-class classification accuracy. Jenni Raitoharju, Alexandros Iosifidis |
IJCNN | 2 |
| 2022 | Product Quality Control in Assembly Machine under Data Restricted SettingsabstractEvaluating the product quality in an assembly machine is critical yet time-consuming since, in product assessment in batch manufacturing, a certain amount of products should be investigated in an invasive manner. However, continuous manufacturing ensures product quality assessment during assembly with high efficiency and traceability. This paper proposes a quality assessment method for an industrial use case. First, the data is prepared based on two indicators and expert knowledge. Then two data classification approaches (one-class classification and binary classification) are applied to evaluate the products’ quality by analysing the related data. Finally, the most efficient model is selected to predict the product labels and deviate anomalies from normal products. For the studied use case and the limited number of products, the binary classifier guarantees to detect 100% of defective products. The proposed approach can provide the engineers and operators with understandable extracted process knowledge, and can therefore be adapted to a high-speed manufacturing line where large data volume and process complexity can be problematic. Fatemeh Kakavandi, Roger De Reus, Cláudio Gomes 0001, Negar Heidari, Alexandros Iosifidis, Peter Gorm Larsen |
INDIN | 5 |
| 2022 | OpenDR: An Open Toolkit for Enabling High Performance, Low Footprint Deep Learning for RoboticsabstractExisting Deep Learning (DL) frameworks typically do not provide ready-to-use solutions for robotics, where very specific learning, reasoning, and embodiment problems exist. Their relatively steep learning curve and the different methodologies employed by DL compared to traditional approaches, along with the high complexity of DL models, which often leads to the need of employing specialized hardware accelerators, further increase the effort and cost needed to employ DL models in robotics. Also, most of the existing DL methods follow a static inference paradigm, as inherited by the traditional computer vision pipelines, ignoring active perception, which can be employed to actively interact with the environment in order to increase perception accuracy. In this paper, we present the Open Deep Learning Toolkit for Robotics (OpenDR). OpenDR aims at developing an open, non-proprietary, efficient, and modular toolkit that can be easily used by robotics companies and research institutions to efficiently develop and deploy AI and cognition technologies to robotics applications, providing a solid step towards addressing the aforementioned challenges. We also detail the design choices, along with an abstract interface that was created to overcome these challenges. This interface can describe various robotic tasks, spanning beyond traditional DL cognition and inference, as known by existing frameworks, incorporating openness, homogeneity and robotics-oriented perception e.g., through active perception, as its core design principles. Nikolaos Passalis, S. Pedrazzi, Robert Babuska, Wolfram Burgard, D. Dias, F. Ferro, Moncef Gabbouj, Ole Green, Alexandros Iosifidis, Erdal Kayacan, Jens Kober, O. Michel, Nikos Nikolaidis 0001, Paraskevi Nousi, Roel Pieters, Maria Tzelepi, Abhinav Valada, Anastasios Tefas |
IROS | 9 |
| 2022 | Collaborative edge computing for distributed CNN inference acceleration using receptive field-based segmentationabstractThis paper studies inference acceleration using distributed CNNs in collaborative edge computing network. To ensure no inference accuracy loss in task partitioning, we propose receptive field-based segmentation. To reduce the computation time and communication overhead, we propose a novel collaborative edge computing using fused-layer parallelization to partition a CNN model into multiple blocks. To find the optimal partition of a CNN model, we use dynamic programming, named as DPFP. To address computation heterogeneity of edge servers (ESs), we design a low-complexity search algorithm which can select the optimal subset of collaborative ESs for inference. The experimental results show that DPFP can accelerate inference up to 71% for ResNet-50 and 73% for VGG-16 compared to running the pre-trained models, which outperforms the existing works MoDNN and DeepSlicing. Moreover, we propose an analytical method to estimate the speedup ratio of different GPU platforms by using FLOPs and effective computing capacity. Furthermore, we evaluate the service failure probability under time-variant channel and variation of image sizes, which shows that DPFP is effective to ensure high service reliability with strict service deadline. Nan Li 0064, Alexandros Iosifidis, Qi Zhang 0013 |
Comput. Networks | 2 |
| 2022 | Remote Multilinear Compressive Learning With Adaptive CompressionabstractMultilinear compressive learning (MCL) is an efficient signal acquisition and learning paradigm for multidimensional signals. The level of signal compression affects the detection or classification performance of an MCL model, with higher compression rates often associated with lower inference accuracy. However, higher compression rates are more amenable to a wider range of applications, especially those that require low operating bandwidth and minimal energy consumption such as Internet of Things (IoT) applications. Many communication protocols provide support for adaptive data transmission to maximize the throughput and minimize energy consumption. By developing compressive sensing and learning models that can operate with an adaptive compression rate, we can maximize the informational content throughput of the whole application. In this article, we propose a novel optimization scheme that enables such a feature for MCL models. Our proposal enables the practical implementation of adaptive compressive signal acquisition and inference systems. Experimental results demonstrated that the proposed approach can significantly reduce the amount of computations required during the training phase of remote learning systems but also improve the informational content throughput via adaptive-rate sensing. Dat Thanh Tran, Moncef Gabbouj, Alexandros Iosifidis |
IEEE Internet Things J. | 3 |
| 2022 | Editorial to special issue on cross-media learning for visual question answering
Shaohua Wan 0001, Chen Chen 0001, Alexandros Iosifidis |
Image Vis. Comput. | 3 |
| 2022 | Single-layer vision transformers for more accurate early exits with less overheadabstractDeploying deep learning models in time-critical applications with limited computational resources, for instance in edge computing systems and IoT networks, is a challenging task that often relies on dynamic inference methods such as early exiting. In this paper, we introduce a novel architecture for early exiting based on the vision transformer architecture, as well as a fine-tuning strategy that significantly increase the accuracy of early exit branches compared to conventional approaches while introducing less overhead. Through extensive experiments on image and audio classification as well as audiovisual crowd counting, we show that our method works for both classification and regression problems, and in both single- and multi-modal settings. Additionally, we introduce a novel method for integrating audio and visual modalities within early exits in audiovisual data analysis, that can lead to a more fine-grained dynamic inference. Arian Bakhtiarnia, Qi Zhang 0013, Alexandros Iosifidis |
Neural Networks | 3 |
| 2022 | Feedforward neural networks initialization based on discriminant learningabstractIn this paper, a novel data-driven method for weight initialization of Multilayer Perceptrons and Convolutional Neural Networks based on discriminant learning is proposed. The approach relaxes some of the limitations of competing data-driven methods, including unimodality assumptions, limitations on the architectures related to limited maximal dimensionalities of the corresponding projection spaces, as well as limitations related to high computational requirements due to the need of eigendecomposition on high-dimensional data. We also consider assumptions of the method on the data and propose a way to account for them in a form of a new normalization layer. The experiments on three large-scale image datasets show improved accuracy of the trained models compared to competing random-based and data-driven weight initialization methods, as well as better convergence properties in certain cases. Kateryna Chumachenko, Alexandros Iosifidis, Moncef Gabbouj |
Neural Networks | 2 |
| 2022 | Saliency-Based Multilabel Linear Discriminant AnalysisabstractLinear discriminant analysis (LDA) is a classical statistical machine-learning method, which aims to find a linear data transformation increasing class discrimination in an optimal discriminant subspace. Traditional LDA sets assumptions related to the Gaussian class distributions and single-label data annotations. In this article, we propose a new variant of LDA to be used in multilabel classification tasks for dimensionality reduction on original data to enhance the subsequent performance of any multilabel classifier. A probabilistic class saliency estimation approach is introduced for computing saliency-based weights for all instances. We use the weights to redefine the between-class and within-class scatter matrices needed for calculating the projection matrix. We formulate six different variants of the proposed saliency-based multilabel LDA (SMLDA) based on different prior information on the importance of each instance for their class(es) extracted from labels and features. Our experiments show that the proposed SMLDA leads to performance improvements in various multilabel classification problems compared to several competing dimensionality reduction methods. Lei Xu 0036, Jenni Raitoharju, Alexandros Iosifidis, Moncef Gabbouj |
IEEE Trans. Cybern. | 3 |
| 2021 | Multi-Exit Vision Transformer for Dynamic Inference
Arian Bakhtiarnia, Qi Zhang 0013, Alexandros Iosifidis |
BMVC | 3 |
| 2021 | ECINN: Efficient Counterfactuals from Invertible Neural Networks
Frederik Hvilshøj, Alexandros Iosifidis, Ira Assent |
BMVC | 2 |
| 2021 | Learning to ignore: rethinking attention in CNNs
Firas Laakom, Kateryna Chumachenko, Jenni Raitoharju, Alexandros Iosifidis, Moncef Gabbouj |
BMVC | 4 |
| 2021 | Robust channel-wise illumination estimation
Firas Laakom, Jenni Raitoharju, Jarno Nikkanen, Alexandros Iosifidis, Moncef Gabbouj |
BMVC | 4 |
| 2021 | Ensembling Object Detectors for Image and Video Data AnalysisabstractIn this paper, we propose a method for ensembling the outputs of multiple object detectors for improving detection performance and precision of bounding boxes on image data. We further extend it to video data by proposing a two-stage tracking-based scheme for detection refinement. The proposed method can be used as a standalone approach for improving object detection performance, or as a part of a framework for faster bounding box annotation in unseen datasets, assuming that the objects of interest are those present in some common public datasets. Kateryna Chumachenko, Jenni Raitoharju, Alexandros Iosifidis, Moncef Gabbouj |
ICASSP | 3 |
| 2021 | Augmenting Transferred Representations for Stock ClassificationabstractStock classification is a challenging task due to high levels of noise and volatility of stocks returns. In this paper we show that using transfer learning can help with this task, by pre-training a model to extract universal features on the full universe of stocks of the S&P500 index and then transferring it to another model to directly learn a trading rule. Transferred models present more than double the risk-adjusted returns than their counterparts trained from zero. In addition, we propose the use of data augmentation on the feature space defined as the output of a pre-trained model (i.e. augmenting the aggregated time-series representation). We compare this augmentation approach with the standard one, i.e. augmenting the time-series in the input space. We show that augmentation methods on the feature space leads to 20% increase in risk-adjusted return compared to a model trained with transfer learning but without augmentation. Elizabeth Fons, Paula Dawson, Xiaojun Zeng, John A. Keane, Alexandros Iosifidis |
ICASSP | 5 |
| 2021 | Progressive Spatio-Temporal Graph Convolutional Network for Skeleton-Based Human Action RecognitionabstractGraph convolutional networks have been very successful in skeleton- based human action recognition where the sequence of skeletons is modeled as a graph. However, most of the graph convolutional network-based methods in this area train a deep feed-forward network with a fixed topology that leads to high computational complexity and restricts their application in low computation scenarios. In this paper, we propose a method to automatically find a compact and problem-specific topology for spatio-temporal graph convolutional networks in a progressive manner. Experimental results on two widely used datasets for skeleton-based human action recognition indicate that the proposed method has competitive or even better classification performance compared to the state-of-the-art methods while it has much lower computational complexity. Negar Heidari, Alexandros Iosifidis |
ICASSP | 2 |
| 2021 | Improving the Accuracy of Early Exits in Multi-Exit Architectures via Curriculum LearningabstractDeploying deep learning services for time-sensitive and resource-constrained settings such as IoT using edge computing systems is a challenging task that requires dynamic adjustment of inference time. Multi-exit architectures allow deep neural networks to terminate their execution early in order to adhere to tight deadlines at the cost of accuracy. To mitigate this cost, in this paper we introduce a novel method called Multi-Exit Curriculum Learning that utilizes curriculum learning, a training strategy for neural networks that imitates human learning by sorting the training samples based on their difficulty and gradually introducing them to the network. Experiments on CIFAR-10 and CIFAR-100 datasets and various configurations of multi-exit architectures show that our method consistently improves the accuracy of early exits compared to the standard training approach. Arian Bakhtiarnia, Qi Zhang 0013, Alexandros Iosifidis |
IJCNN | 3 |
| 2021 | On the spatial attention in spatio-temporal graph convolutional networks for skeleton-based human action recognitionabstractGraph convolutional networks (GCNs) achieved promising performance in skeleton-based human action recognition by modeling a sequence of skeletons as a spatio-temporal graph. Most of the recently proposed GCN-based methods improve the performance by learning the graph structure at each layer of the network using spatial attention applied on a predefined graph Adjacency matrix that is optimized jointly with model's parameters in an end-to-end manner. In this paper, we analyze the spatial attention used in spatio-temporal GCN layers and propose a symmetric spatial attention for better reflecting the symmetric property of the relative positions of the human body joints when executing actions. We also highlight the connection of spatio-temporal GCN layers employing additive spatial attention to bilinear layers, and we propose the spatiotemporal bilinear network (ST-BLN) which does not require the use of predefined Adjacency matrices and allows for more flexible design of the model. Experimental results show that the three models lead to effectively the same performance. Moreover, by exploiting the flexibility provided by the proposed ST-BLN, one can increase the efficiency of the model. Negar Heidari, Alexandros Iosifidis |
IJCNN | 2 |
| 2021 | Monte Carlo Dropout Ensembles for Robust Illumination EstimationabstractComputational color constancy is a preprocessing step used in many camera systems. The main aim is to discount the effect of the illumination on the colors in the scene and restore the original colors of the objects. Recently, several deep learning-based approaches have been proposed to solve this problem and they often led to state-of-the-art performance in terms of average errors. However, for extreme samples, these methods fail and lead to high errors. In this paper, we address this limitation by proposing to aggregate different deep learning methods according to their output uncertainty. We estimate the relative uncertainty of each approach using Monte Carlo dropout and the final illumination estimate is obtained as the sum of the different model estimates weighted by the log-inverse of their corresponding uncertainties. The proposed framework leads to state-of-the-art performance on INTEL-TAU dataset. Firas Laakom, Jenni Raitoharju, Alexandros Iosifidis, Jarno Nikkanen, Moncef Gabbouj |
IJCNN | 3 |
| 2021 | Speech Command Recognition in Computationally Constrained Environments with a Quadratic Self-Organized Operational LayerabstractAutomatic classification of speech commands has revolutionized human computer interactions in robotic applications. However, employed recognition models usually follow the methodology of deep learning with complicated networks which are memory and energy hungry. So, there is a need to either squeeze these complicated models or use more efficient lightweight models in order to be able to implement the resulting classifiers on embedded devices. In this paper, we pick the second approach and propose a network layer to enhance the speech command recognition capability of a lightweight network and demonstrate the result via experiments. The employed method borrows the ideas of Taylor expansion and quadratic forms to construct a better representation of features in both input and hidden layers. This richer representation results in recognition accuracy improvement as shown by extensive experiments on Google speech commands (GSC) and synthetic speech commands (SSC) datasets. Mohammad Soltanian, Junaid Malik, Jenni Raitoharju, Alexandros Iosifidis, Serkan Kiranyaz, Moncef Gabbouj |
IJCNN | 4 |
| 2021 | Progressive Spatio-Temporal Bilinear Network with Monte Carlo Dropout for Landmark-based Facial Expression Recognition with Uncertainty EstimationabstractDeep neural networks have been widely used for feature learning in facial expression recognition systems. However, small datasets and large intra-class variability can lead to overfitting. In this paper, we propose a method which learns an optimized compact network topology for real-time facial expression recognition utilizing localized facial landmark features. Our method employs a spatio-temporal bilinear layer as backbone to capture the motion of facial landmarks during the execution of a facial expression effectively. Besides, it takes advantage of Monte Carlo Dropout to capture the model’s uncertainty which is of great importance to analyze and treat uncertain cases. The performance of our method is evaluated on three widely used datasets and it is comparable to that of video-based state-of-the-art methods while it has much less complexity. Negar Heidari, Alexandros Iosifidis |
MMSP | 2 |
| 2021 | Automatic Main Character Recognition for Photographic StudiesabstractMain characters in images are the most important humans that catch the viewer’s attention upon first look, and they are emphasized by properties such as size, position, color saturation, and sharpness of focus. Identifying the main character in images plays an important role in traditional photographic studies and media analysis, but the task is performed manually and is, thus, slow and laborious. Furthermore, selection of main characters can be sometimes subjective. In this paper, we analyze the feasibility of solving the main character recognition needed for photographic studies automatically and propose a method for identifying the main characters. The proposed method uses machine learning based human pose estimation along with traditional computer vision approaches for this task. We approach the task as a binary classification problem where each detected human is classified either as a main character or not. To evaluate both the subjectivity of the task and the performance of our method, we collected a dataset of 300 varying images from multiple sources and asked five people, a photographic researcher and four other persons, to annotate the main characters. Our analysis showed a relatively high agreement between different annotators. The proposed method achieved a promising F1 score of 0.83 on the full image set and 0.96 on a subset evaluated as most clear and important cases by the photographic researcher. Mert Seker, Anssi Männistö, Alexandros Iosifidis, Jenni Raitoharju |
MMSP | 3 |
| 2021 | Pain detection using batch normalized discriminant restricted Boltzmann machine layers
Reza Kharghanian, Ali Peiravi, Farshad Moradi, Alexandros Iosifidis |
J. Vis. Commun. Image Represent. | 4 |
| 2021 | Exploiting heterogeneity in operational neural networks by synaptic plasticityabstractAbstract The recently proposed network model, Operational Neural Networks (ONNs), can generalize the conventional Convolutional Neural Networks (CNNs) that are homogenous only with a linear neuron model. As a heterogenous network model, ONNs are based on a generalized neuron model that can encapsulateanyset of non-linear operators to boost diversity and to learn highly complex and multi-modal functions or spaces with minimal network complexity and training data. However, the default search method to find optimal operators in ONNs, the so-called Greedy Iterative Search (GIS) method, usually takes several training sessions to find a single operator set per layer. This is not only computationally demanding, also the network heterogeneity is limited since the same set of operators will then be used for all neurons in each layer. To address this deficiency and exploit a superior level of heterogeneity, in this study the focus is drawn on searching the best-possible operator set(s) for the hidden neurons of the network based on the “Synaptic Plasticity” paradigm that poses the essential learning theory in biological neurons. During training, each operator set in the library can be evaluated by their synaptic plasticity level, ranked from the worst to the best, and an “elite” ONN can then be configured using the top-ranked operator sets found at each hidden layer. Experimental results over highly challenging problems demonstrate that the elite ONNs even with few neurons and layers can achieve a superior learning performance than GIS-based ONNs and as a result, the performance gap over the CNNs further widens. Serkan Kiranyaz, Junaid Malik, Habib Ben Abdallah, Turker Ince, Alexandros Iosifidis, Moncef Gabbouj |
Neural Comput. Appl. | 5 |
| 2021 | Visualising deep network time-series representations
Blazej Leporowski, Alexandros Iosifidis |
Neural Comput. Appl. | 2 |
| 2021 | Self-organized Operational Neural Networks with Generative NeuronsabstractOperational Neural Networks (ONNs) have recently been proposed to address the well-known limitations and drawbacks of conventional Convolutional Neural Networks (CNNs) such as network homogeneity with the sole linear neuron model. ONNs are heterogeneous networks with a generalized neuron model. However the operator search method in ONNs is not only computationally demanding, but the network heterogeneity is also limited since the same set of operators will then be used for all neurons in each layer. Moreover, the performance of ONNs directly depends on the operator set library used, which introduces a certain risk of performance degradation especially when the optimal operator set required for a particular task is missing from the library. In order to address these issues and achieve an ultimate heterogeneity level to boost the network diversity along with computational efficiency, in this study we propose Self-organized ONNs (Self-ONNs) with generative neurons that can adapt (optimize) the nodal operator of each connection during the training process. Moreover, this ability voids the need of having a fixed operator set library and the prior operator search within the library in order to find the best possible set of operators. We further formulate the training method to back-propagate the error through the operational layers of Self-ONNs. Experimental results over four challenging problems demonstrate the superior learning capability and computational efficiency of Self-ONNs over conventional ONNs and CNNs. Serkan Kiranyaz, Junaid Malik, Habib Ben Abdallah, Turker Ince, Alexandros Iosifidis, Moncef Gabbouj |
Neural Networks | 5 |
| 2021 | Speed-up and multi-view extensions to subclass discriminant analysisabstractIn this paper, we propose a speed-up approach for subclass discriminant analysis and formulate a novel efficient multi-view solution to it. The speed-up approach is developed based on graph embedding and spectral regression approaches that involve eigendecomposition of the corresponding Laplacian matrix and regression to its eigenvectors. We show that by exploiting the structure of the between-class Laplacian matrix, the eigendecomposition step can be substituted with a much faster process. Furthermore, we formulate a novel criterion for multi-view subclass discriminant analysis and show that an efficient solution to it can be obtained in a similar manner to the single-view case. We evaluate the proposed methods on nine single-view and nine multi-view datasets and compare them with related existing approaches. Experimental results show that the proposed solutions achieve competitive performance, often outperforming the existing methods. At the same time, they significantly decrease the training time. Kateryna Chumachenko, Jenni Raitoharju, Alexandros Iosifidis, Moncef Gabbouj |
Pattern Recognit. | 3 |
| 2021 | Multimodal subspace support vector data descriptionabstractIn this paper, we propose a novel method for projecting data from multiple modalities to a new subspace optimized for one-class classification. The proposed method iteratively transforms the data from the original feature space of each modality to a new common feature space along with finding a joint compact description of data coming from all the modalities. For data in each modality, we define a separate transformation to map the data from the corresponding feature space to the new optimized subspace by exploiting the available information from the class of interest only. We also propose different regularization strategies for the proposed method and provide both linear and non-linear formulations. The proposed Multimodal Subspace Support Vector Data Description outperforms all the competing methods using data from a single modality or fusing data from all modalities in four out of five datasets. Fahad Sohrab, Jenni Raitoharju, Alexandros Iosifidis, Moncef Gabbouj |
Pattern Recognit. | 3 |
| 2021 | Supervised Domain Adaptation: A Graph Embedding Perspective and a Rectified Experimental ProtocolabstractDomain Adaptation is the process of alleviating distribution gaps between data from different domains. In this paper, we show that Domain Adaptation methods using pair-wise relationships between source and target domain data can be formulated as a Graph Embedding in which the domain labels are incorporated into the structure of the intrinsic and penalty graphs. Specifically, we analyse the loss functions of three existing state-of-the-art Supervised Domain Adaptation methods and demonstrate that they perform Graph Embedding. Moreover, we highlight some generalisation and reproducibility issues related to the experimental setup commonly used to demonstrate the few-shot learning capabilities of these methods. To assess and compare Supervised Domain Adaptation methods accurately, we propose a rectified evaluation protocol, and report updated benchmarks on the standard datasets Office31 (Amazon, DSLR, and Webcam), Digits (MNIST, USPS, SVHN, and MNIST-M) and VisDA (Synthetic, Real). Lukas Hedegaard, Omar Ali Sheikh-Omar, Alexandros Iosifidis |
IEEE Trans. Image Process. | 3 |
| 2021 | Deep Multi-View Learning to RankabstractWe study the problem of learning to rank from multiple information sources. Though multi-view learning and learning to rank have been studied extensively leading to a wide range of applications, multi-view learning to rank as a synergy of both topics has received little attention. The aim of the paper is to propose a composite ranking method while keeping a close correlation with the individual rankings simultaneously. We present a generic framework for multi-view subspace learning to rank (MvSL2R), and two novel solutions are introduced under the framework. The first solution captures information of feature mappings from within each view as well as across views using autoencoder-like networks. Novel feature embedding methods are formulated in the optimization of multi-view unsupervised and discriminant autoencoders. Moreover, we introduce an end-to-end solution to learning towards both the joint ranking objective and the individual rankings. The proposed solution enhances the joint ranking with minimum view-specific ranking loss, so that it can achieve the maximum global view agreements in a single optimization process. The proposed method is evaluated on three different ranking problems, i.e., university ranking, multi-view lingual text ranking, and image data ranking, providing superior results compared to related methods. Guanqun Cao, Alexandros Iosifidis, Moncef Gabbouj, Vijay Raghavan 0001, Raju N. Gottumukkala |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2021 | Hypersphere-Based Weight Imprinting for Few-Shot Learning on Embedded DevicesabstractWeight imprinting (WI) was recently introduced as a way to perform gradient descent-free few-shot learning. Due to this, WI was almost immediately adapted for performing few-shot learning on embedded neural network accelerators that do not support back-propagation, e.g., edge tensor processing units. However, WI suffers from many limitations, e.g., it cannot handle novel categories with multimodal distributions and special care should be given to avoid overfitting the learned embeddings on the training classes since this can have a devastating effect on classification accuracy (for the novel categories). In this article, we propose a novel hypersphere-based WI approach that is capable of training neural networks in a regularized, imprinting-aware way effectively overcoming the aforementioned limitations. The effectiveness of the proposed method is demonstrated using extensive experiments on three image data sets. Nikolaos Passalis, Alexandros Iosifidis, Moncef Gabbouj, Anastasios Tefas |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2021 | Multilinear Compressive LearningabstractCompressive learning (CL) is an emerging topic that combines signal acquisition via compressive sensing (CS) and machine learning to perform inference tasks directly on a small number of measurements. Many data modalities naturally have a multidimensional or tensorial format, with each dimension or tensor mode representing different features such as the spatial and temporal information in video sequences or the spatial and spectral information in hyperspectral images. However, in existing CL frameworks, the CS component utilizes either random or learned linear projection on the vectorized signal to perform signal acquisition, thus discarding the multidimensional structure of the signals. In this article, we propose multilinear CL (MCL), a framework that takes into account the tensorial nature of multidimensional signals in the acquisition step and builds the subsequent inference model on the structurally sensed measurements. Our theoretical complexity analysis shows that the proposed framework is more efficient compared to its vector-based counterpart in both memory and computation requirement. With extensive experiments, we also empirically show that our MCL framework outperforms the vector-based framework in object classification and face recognition tasks, and scales favorably when the dimensionalities of the original signals increase, making it highly efficient for high-dimensional multidimensional signals. Dat Thanh Tran, Mehmet Yamac, Aysen Degerli, Moncef Gabbouj, Alexandros Iosifidis |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2020 | Adaptive Normalization for Forecasting Limit Order Book Data Using Convolutional Neural NetworksabstractDeep learning models are capable of achieving state-of-the-art performance on a wide range of time series analysis tasks. However, their performance crucially depends on the employed normalization scheme, while they are usually unable to efficiently handle non-stationary features without first appropriately pre-processing them. These limitations impact the performance of deep learning models, especially when used for forecasting financial time series, due to their non-stationary and multimodal nature. In this paper we propose a data-driven adaptive normalization layer which is capable of learning the most appropriate normalization scheme that should be applied on the data. To this end, the proposed method first identifies the distribution from which the data were generated and then it dynamically shifts and scales them in order to facilitate the task at hand. The proposed nor-malization scheme is fully differentiable and it is trained in an end-to-end fashion along with the rest of the parameters of the model. The proposed method leads to significant performance improvements over several competitive normalization approaches, as demonstrated using a large-scale limit order book dataset. Nikolaos Passalis, Anastasios Tefas, Juho Kanniainen, Moncef Gabbouj, Alexandros Iosifidis |
ICASSP | 5 |
| 2020 | Text-To-Image Synthesis Method Evaluation Based On Visual PatternsabstractA commonly used evaluation metric for text-to-image synthesis is the Inception score (IS) [1], which has been shown to be a quality metric that correlates well with human judgment. However, IS does not reveal properties of the generated images indicating the ability of a text-to-image synthesis method to correctly convey semantics of the input text descriptions. In this paper, we introduce an evaluation metric and a visual evaluation method allowing for the simultaneous estimation of the realism, variety and semantic accuracy of generated images. The proposed method uses a pre-trained Inception network [2] to produce high dimensional representations for both real and generated images. These image representations are then visualized in a 2-dimensional feature space defined by the t-distributed Stochastic Neighbor Embedding (t-SNE) [3]. Visual concepts are determined by clustering the real image representations, and are subsequently used to evaluate the similarity of the generated images to the real ones by classifying them to the closest visual concept. The resulting classification accuracy is shown to be a effective gauge for the semantic accuracy of text-to-image synthesis methods. William Lund Sommer, Alexandros Iosifidis |
ICASSP | 2 |
| 2020 | Incremental Fast Subclass Discriminant AnalysisabstractThis paper proposes an incremental solution to Fast Subclass Discriminant Analysis (fastSDA). We present an exact and an approximate linear solution, along with an approximate kernelized variant. Extensive experiments on eight image datasets with different incremental batch sizes show the superiority of the proposed approach in terms of training time and accuracy being equal or close to fastSDA solution and outperforming other methods. Kateryna Chumachenko, Jenni Raitoharju, Moncef Gabbouj, Alexandros Iosifidis |
ICIP | 4 |
| 2020 | Probabilistic Color ConstancyabstractIn this paper, we propose a novel unsupervised color constancy method, called Probabilistic Color Constancy (PCC). We define a framework for estimating the illumination of a scene by weighting the contribution of different image regions using a graph-based representation of the image. To estimate the weight of each (super-)pixel, we rely on two assumptions: (Super-)pixels with similar colors contribute similarly and darker (super-)pixels contribute less. The resulting system has one global optimum solution. The proposed method achieves competitive performance, compared to the state-of-the-art, on INTEL-TAU dataset. Firas Laakom, Jenni Raitoharju, Alexandros Iosifidis, Uygar Tuna, Jarno Nikkanen, Moncef Gabbouj |
ICIP | 3 |
| 2020 | Subset Sampling for Progressive Neural Network LearningabstractProgressive Neural Network Learning is a class of algorithms that incrementally construct the network's topology and optimize its parameters based on the training data. While this approach exempts the users from the manual task of designing and validating multiple network topologies, it often requires an enormous number of computations. In this paper, we propose to speed up this process by exploiting subsets of training data at each incremental training step. Three different sampling strategies for selecting the training samples according to different criteria are proposed and evaluated. We also propose to perform online hyperparameter selection during the network progression, which further reduces the overall training time. Experimental results in object, scene and face recognition problems demonstrate that the proposed approach speeds up the optimization procedure considerably while operating on par with the baseline approach exploiting the entire training set throughout the training process. Dat Thanh Tran, Moncef Gabbouj, Alexandros Iosifidis |
ICIP | 3 |
| 2020 | Temporal Attention-Augmented Graph Convolutional Network for Efficient Skeleton-Based Human Action RecognitionabstractGraph convolutional networks (GCNs) have been very successful in modeling non-Euclidean data structures, like sequences of body skeletons forming actions modeled as spatiotemporal graphs. Most GCN-based action recognition methods use deep feed-forward networks with high computational complexity to process all skeletons in an action. This leads to a high number of floating point operations (ranging from 16G to 100G FLOPs) to process a single sample, making their adoption in restricted computation application scenarios infeasible. In this paper, we propose a temporal attention module (TAM) for increasing the efficiency in skeleton-based action recognition by selecting the most informative skeletons of an action at the early layers of the network. We incorporate the TAM in a lightweight GCN topology to further reduce the overall number of computations. Experimental results on two benchmark datasets show that the proposed method outperforms with a large margin the baseline GCN-based method while having ×2.9 less number of computations. Moreover, it performs on par with the state-of-the-art with up to ×9.6 less number of computations. Negar Heidari, Alexandros Iosifidis |
ICPR | 2 |
| 2020 | Supervised Domain Adaptation using Graph EmbeddingabstractGetting deep convolutional neural networks to perform well requires a large amount of training data. When the available labelled data is small, it is often beneficial to use transfer learning to leverage a related larger dataset (source) in order to improve the performance on the small dataset (target). Among the transfer learning approaches, domain adaptation methods assume that distributions between the two domains are shifted and attempt to realign them. In this paper, we consider the domain adaptation problem from the perspective of multi-view graph embedding and dimensionality reduction. Instead of solving the generalised eigenvalue problem to perform the embedding, we formulate the graph-preserving criterion as a loss in the neural network and learn a domain-invariant feature transformation in an end-to-end fashion. We show that the proposed approach leads to a powerful Domain Adaptation framework which generalises the prior methods CCSA and d-SNE, and enables simple and effective loss designs; an LDA-inspired instantiation of the framework leads to performance on par with the state-of-the-art on the most widely used Domain Adaptation benchmarks, Office31 and MNIST to USPS datasets. Lukas Hedegaard, Omar Ali Sheikh-Omar, Alexandros Iosifidis |
ICPR | 3 |
| 2020 | Not all domains are equally complex: Adaptive Multi-Domain LearningabstractDeep learning approaches are highly specialized and require training separate models for different tasks. Multidomain learning looks at ways to learn a multitude of different tasks, each coming from a different domain, at once. The most common approach in multi-domain learning is to form a domain agnostic model, the parameters of which are shared among all domains, and learn a small number of extra domain-specific parameters for each individual new domain. However, different domains come with different levels of difficulty; parameterizing the models of all domains using an augmented version of the domain agnostic model leads to unnecessarily inefficient solutions, especially for easy to solve tasks. We propose an adaptive parameterization approach to deep neural networks for multidomain learning. The proposed approach performs on par with the original approach while reducing by far the number of parameters, leading to efficient multi-domain learning solutions. Ali Senhaji, Jenni Raitoharju, Moncef Gabbouj, Alexandros Iosifidis |
ICPR | 4 |
| 2020 | Data Normalization for Bilinear Structures in High-Frequency Financial Time-seriesabstractFinancial time-series analysis and forecasting have been extensively studied over the past decades, yet still remain as a very challenging research topic. Since the financial market is inherently noisy and stochastic, a majority of financial time-series of interests are non-stationary, and often obtained from different modalities. This property presents great challenges and can significantly affect the performance of the subsequent analysis/forecasting steps. Recently, the Temporal Attention augmented Bilinear Layer (TABL) has shown great performances in tackling financial forecasting problems. In this paper, by taking into account the nature of bilinear projections in TABL networks, we propose Bilinear Normalization (BiN), a simple, yet efficient normalization layer to be incorporated into TABL networks to tackle potential problems posed by non-stationarity and multimodalities in the input series. Our experiments using a large scale Limit Order Book (LOB) consisting of more than 4 million order events show that BiN-TABL outperforms TABL networks using other state-of-the-arts normalization schemes by a large margin. Dat Thanh Tran, Juho Kanniainen, Moncef Gabbouj, Alexandros Iosifidis |
ICPR | 4 |
| 2020 | Progressive Operational Perceptrons with MemoryabstractGeneralized Operational Perceptron (GOP) was proposed to generalize the linear neuron model used in the traditional Multilayer Perceptron (MLP) by mimicking the synaptic connections of biological neurons showing nonlinear neurochemical behaviours. Previously, Progressive Operational Perceptron (POP) was proposed to train a multilayer network of GOPs which is formed layer-wise in a progressive manner. While achieving superior learning performance over other types of networks, POP has a high computational complexity. In this work, we propose POPfast, an improved variant of POP that signicantly reduces the computational complexity of POP, thus accelerating the training time of GOP networks. In addition, we also propose major architectural modications of POPfast that can augment the progressive learning process of POP by incorporating an information preserving, linear projection path from the input to the output layer at each progressive step. The proposed extensions can be interpreted as a mechanism that provides direct information extracted from the previously learned layers to the network, hence the term “memory”. This allows the network to learn deeper architectures and better data representations. An extensive set of experiments in human action, object, facial identity and scene recognition problems demonstrates that the proposed algorithms can train GOP networks much faster than POPs while achieving better performance compared to original POPs and other related algorithms. Dat Thanh Tran, Serkan Kiranyaz, Moncef Gabbouj, Alexandros Iosifidis |
Neurocomputing | 4 |
| 2020 | Operational neural networksabstractAbstract Feed-forward, fully connected artificial neural networks or the so-called multi-layer perceptrons are well-known universal approximators. However, their learning performance varies significantly depending on the function or the solution space that they attempt to approximate. This is mainly because of their homogenous configuration based solely on the linear neuron model. Therefore, while they learn very well those problems with a monotonous, relatively simple and linearly separable solution space, they may entirely fail to do so when the solution space is highly nonlinear and complex. Sharing the same linear neuron model with two additional constraints (local connections and weight sharing), this is also true for the conventional convolutional neural networks (CNNs) and it is, therefore, not surprising that in many challenging problems only the deep CNNs with a massive complexity and depth can achieve the required diversity and the learning performance. In order to address this drawback and also to accomplish a more generalized model over the convolutional neurons, this study proposes a novel network model, called operational neural networks (ONNs), which can be heterogeneous and encapsulate neurons with any set of operators to boost diversity and to learn highly complex and multi-modal functions or spaces with minimal network complexity and training data. Finally, the training method to back-propagate the error through the operational layers of ONNs is formulated. Experimental results over highly challenging problems demonstrate the superior learning capabilities of ONNs even with few neurons and hidden layers. Serkan Kiranyaz, Turker Ince, Alexandros Iosifidis, Moncef Gabbouj |
Neural Comput. Appl. | 3 |
| 2020 | Recurrent bag-of-features for visual information analysisabstractDeep Learning (DL) has provided powerful tools for visual information analysis. For example, Convolutional Neural Networks (CNNs) are excelling in complex and challenging image analysis tasks by extracting meaningful feature vectors with high discriminative power . However, these powerful feature vectors are crushed through the pooling layers of the network, that usually implement the pooling operation in a less sophisticated manner. This can lead to significant information loss, especially in cases where the informative content of the data is sequentially distributed over the spatial or temporal dimension, e.g., videos, which often require extracting fine-grained temporal information. A novel stateful recurrent pooling approach, that can overcome the aforementioned limitations, is proposed in this paper. The proposed method is inspired by the well-known Bag-of-Features (BoF) model, but employs a stateful trainable recurrent quantizer, instead of plain static quantization, allowing for efficiently processing sequential data and encoding both their temporal, as well as their spatial aspects. The effectiveness of the proposed Recurrent BoF model to enclose spatio-temporal information compared to other competitive methods is demonstrated using six different datasets and two different tasks. Marios Krestenitis, Nikolaos Passalis, Alexandros Iosifidis, Moncef Gabbouj, Anastasios Tefas |
Pattern Recognit. | 3 |
| 2020 | Class mean vector component and discriminant analysis
Alexandros Iosifidis |
Pattern Recognit. Lett. | 1 |
| 2020 | Variance-preserving deep metric learning for content-based image retrievalabstractSupervised deep metric learning led to spectacular results for several Content-based Information Retrieval (CBIR) applications. The success of these approaches slowly led to the belief that image retrieval and classification are just slightly different variations of the same problem. However, recent evidence suggests that learning highly discriminative representation for a (limited) set of training classes removes valuable information from the representation, potentially harming both the in-domain, as well as the out-of-domain retrieval precision. In this paper, we propose a regularized discriminative deep metric learning method that aims to not only learn a representation that allows for discriminating between different classes, but it is also capable of encoding the latent generative factors separately for each class, overcoming this limitation. This allows for modeling the in-class variance and, as a result, maintaining the ability to represent both sub-classes of the in-domain data, as well as objects that belong to classes outside the training domain. The effectiveness of the proposed method, over existing supervised and unsupervised representation/metric learning approaches , is demonstrated under different in-domain and out-of-domain setups and three challenging image datasets. Nikolaos Passalis, Alexandros Iosifidis, Moncef Gabbouj, Anastasios Tefas |
Pattern Recognit. Lett. | 2 |
| 2020 | Temporal logistic neural Bag-of-Features for financial time series forecasting leveraging limit order book dataabstract• Logistic Neural Bag-of-Features are employed for financial time series analysis. • The proposed method can be efficiently used in deep learning architectures. • An adaptive scaling method is proposed to ensure the smooth flow of information. • A logistic kernel is used to estimate the feature vector densities. • The proposed method outperforms the competitive methods on a large-scale dataset. Time series forecasting is a crucial component of many important applications, ranging from forecasting the stock markets to energy load prediction. The high-dimensionality, velocity and variety of the data collected in many of these applications pose significant and unique challenges that must be carefully addressed for each of them. In this work, a novel Temporal Logistic Neural Bag-of-Features approach, that can be used to tackle these challenges, is proposed. The proposed method can be effectively combined with deep neural networks , leading to powerful deep learning models for time series analysis. However, combining existing BoF formulations with deep feature extractors pose significant challenges: the distribution of the input features is not stationary, tuning the hyper-parameters of the model can be especially difficult and the normalizations involved in the BoF model can cause significant instabilities during the training process. The proposed method is capable of overcoming these limitations by a employing a novel adaptive scaling mechanism and replacing the classical Gaussian-based density estimation involved in the regular BoF model with a logistic kernel. The effectiveness of the proposed approach is demonstrated using extensive experiments on a large-scale limit order book dataset that consists of more than 4 million limit orders. Nikolaos Passalis, Anastasios Tefas, Juho Kanniainen, Moncef Gabbouj, Alexandros Iosifidis |
Pattern Recognit. Lett. | 5 |
| 2020 | Human experts vs. machines in taxa recognition
Johanna Ärje, Jenni Raitoharju, Alexandros Iosifidis, Ville Tirronen, Kristian Meissner, Moncef Gabbouj, Serkan Kiranyaz, Salme Kärkkäinen |
Signal Process. Image Commun. | 3 |
| 2020 | Deep Learning for Visual Content Analysis
Alexandros Iosifidis, Anastasios Tefas |
Signal Process. Image Commun. | 1 |
| 2020 | Bag of Color Features for Color ConstancyabstractIn this paper, we propose a novel color constancy approach, called Bag of Color Features (BoCF), building upon Bag-of-Features pooling. The proposed method substantially reduces the number of parameters needed for illumination estimation. At the same time, the proposed method is consistent with the color constancy assumption stating that global spatial information is not relevant for illumination estimation and local information (edges, etc.) is sufficient. Furthermore, BoCF is consistent with color constancy statistical approaches and can be interpreted as a learning-based extension of many statistical approaches. To further improve the illumination estimation accuracy, we propose a novel attention mechanism for the BoCF model with two variants based on self-attention. BoCF approach and its variants achieve competitive, compared to the state of the art, results while requiring much fewer parameters on three benchmark datasets: ColorChecker RECommended, INTEL-TUT version 2, and NUS8. Firas Laakom, Nikolaos Passalis, Jenni Raitoharju, Jarno Nikkanen, Anastasios Tefas, Alexandros Iosifidis, Moncef Gabbouj |
IEEE Trans. Image Process. | 6 |
| 2020 | Deep Adaptive Input Normalization for Time Series ForecastingabstractDeep learning (DL) models can be used to tackle time series analysis tasks with great success. However, the performance of DL models can degenerate rapidly if the data are not appropriately normalized. This issue is even more apparent when DL is used for financial time series forecasting tasks, where the nonstationary and multimodal nature of the data pose significant challenges and severely affect the performance of DL models. In this brief, a simple, yet effective, neural layer that is capable of adaptively normalizing the input time series, while taking into account the distribution of the data, is proposed. The proposed layer is trained in an end-to-end fashion using backpropagation and leads to significant performance improvements compared to other evaluated normalization schemes. The proposed method differs from traditional normalization methods since it learns how to perform normalization for a given task instead of using a fixed normalization scheme. At the same time, it can be directly applied to any new time series without requiring retraining. The effectiveness of the proposed method is demonstrated using a large-scale limit order book data set, as well as a load forecasting data set. Nikolaos Passalis, Anastasios Tefas, Juho Kanniainen, Moncef Gabbouj, Alexandros Iosifidis |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2020 | Heterogeneous Multilayer Generalized Operational PerceptronabstractThe traditional multilayer perceptron (MLP) using a McCulloch-Pitts neuron model is inherently limited to a set of neuronal activities, i.e., linear weighted sum followed by nonlinear thresholding step. Previously, generalized operational perceptron (GOP) was proposed to extend the conventional perceptron model by defining a diverse set of neuronal activities to imitate a generalized model of biological neurons. Together with GOP, a progressive operational perceptron (POP) algorithm was proposed to optimize a predefined template of multiple homogeneous layers in a layerwise manner. In this paper, we propose an efficient algorithm to learn a compact, fully heterogeneous multilayer network that allows each individual neuron, regardless of the layer, to have distinct characteristics. Based on the complexity of the problem, the proposed algorithm operates in a progressive manner on a neuronal level, searching for a compact topology, not only in terms of depth but also width, i.e., the number of neurons in each layer. The proposed algorithm is shown to outperform other related learning methods in extensive experiments on several classification problems. Dat Thanh Tran, Serkan Kiranyaz, Moncef Gabbouj, Alexandros Iosifidis |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2019 | Deep Temporal Logistic Bag-of-features for Forecasting High Frequency Limit Order Book Time SeriesabstractForecasting time series has several applications in various domains. The vast amount of data that are available nowadays provide the opportunity to use powerful deep learning approaches, but at the same time pose significant challenges of high-dimensionality, velocity and variety. In this paper, a novel logistic formulation of the well-known Bag-of-Features model is proposed to tackle these challenges. The proposed method is combined with deep convolutional feature extractors and is capable of accurately modeling the temporal behavior of time series, forming powerful forecasting models that can be trained in an end-to-end fashion. The proposed method was extensively evaluated using a large-scale financial time series dataset, that consists of more than 4 million limit orders, outperforming other competitive methods. Nikolaos Passalis, Anastasios Tefas, Juho Kanniainen, Moncef Gabbouj, Alexandros Iosifidis |
ICASSP | 5 |
| 2019 | Learning to Rank: A Progressive Neural Network Learning ApproachabstractLearning to rank is an essential component in an information retrieval system. The state-of-the-art ranking systems are often based on an ensemble of classifiers, such as Random Forest or LambdaMART, which aggregates the ranking outputs produced by thousands of classifiers. The storage and computation requirement of an ensemble model is usually very high, imposing a significant operating cost to the retrieval system. To tackle this problem, we propose an algorithm that adaptively learns a single heterogeneous feedforward network architecture, composing of Generalized Operational Perceptrons, given a ranking problem. Experimental results in web search ranking and image retrieval tasks show that the proposed algorithm compares favourably to the related algorithms. Dat Thanh Tran, Alexandros Iosifidis |
ICASSP | 2 |
| 2019 | Class-Based Variational Representation Learning For Robust Image RetrievalabstractSupervised learning for Content-based Information Retrieval allows for obtaining discriminative representations that often excel within the training domain. However, recent evidence suggests that these representations can actually harm the retrieval precision for queries that do not belong to the domain of the training set compared to other, less discriminative representations. To avoid this behavior, we propose to learn discriminative representations which also encode the latent generative factors for each class. In this way, the proposed method is capable of maintaining (part of) the in-class variance, as well as being able to represent data that belong to classes that were not seen during the training by better learning the structure of the input space. The proposed method is evaluated under different in-domain and out-of-domain setups, significantly outperforming existing supervised and unsupervised representation learning approaches. Nikolaos Passalis, Anastasios Tefas, Alexandros Iosifidis, Moncef Gabbouj |
ICIP | 3 |
| 2019 | Knowledge Transfer for Face Verification Using Heterogeneous Generalized Operational PerceptronsabstractFace verification is a prominent biometric technique for identity authentication that has been used extensively in several security applications. In practice, face verification is often performed along with other visual surveillance tasks in the computing device. Thus, the ability to share the computation and reuse the information already extracted for other analysis tasks can greatly help reduce the computation load on the devices. In this study, we propose to utilize the knowledge transfer approach for the face verification problem by building a heterogeneous neural network architecture of Generalized Operational Perceptrons on top of the intermediate features extracted for object recognition purpose. Experimental results show that using our proposed approach, a face verification system can be incorporated into an existing visual analysis system with less additional memory and computational cost, compared to other similar approaches. Dat Thanh Tran, Serkan Kiranyaz, Moncef Gabbouj, Alexandros Iosifidis |
ICIP | 4 |
| 2019 | PyGOP: A Python library for Generalized Operational Perceptron algorithms
Dat Thanh Tran, Serkan Kiranyaz, Moncef Gabbouj, Alexandros Iosifidis |
Knowl. Based Syst. | 4 |
| 2019 | Temporal Attention-Augmented Bilinear Network for Financial Time-Series Data AnalysisabstractFinancial time-series forecasting has long been a challenging problem because of the inherently noisy and stochastic nature of the market. In the high-frequency trading, forecasting for trading purposes is even a more challenging task, since an automated inference system is required to be both accurate and fast. In this paper, we propose a neural network layer architecture that incorporates the idea of bilinear projection as well as an attention mechanism that enables the layer to detect and focus on crucial temporal information. The resulting network is highly interpretable, given its ability to highlight the importance and contribution of each temporal instance, thus allowing further analysis on the time instances of interest. Our experiments in a large-scale limit order book data set show that a two-hidden-layer network utilizing our proposed layer outperforms by a large margin all existing state-of-the-art results coming from much deeper architectures while requiring far fewer computations. Dat Thanh Tran, Alexandros Iosifidis, Juho Kanniainen, Moncef Gabbouj |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2018 | Acceleration Approaches for Big Data AnalysisabstractThe massive size of data that needs to be processed by Machine Learning models nowadays sets new challenges related to their computational complexity and memory footprint. These challenges span all processing steps involved in the application of the related models, i.e., from the fundamental processing steps needed to evaluate distances of vectors, to the optimization of large-scale systems, e.g. for non-linear regression using kernels, or the speed up of deep learning models formed by billions of parameters. In order to address these challenges, new approximate solutions have been recently proposed based on matrix/tensor decompositions, randomization and quantization strategies. This paper provides a comprehensive review of the related methodologies and discusses their connections. Anton Muravev, Dat Thanh Tran, Moncef Gabbouj, Alexandros Iosifidis, Serkan Kiranyaz |
ICIP | 4 |
| 2018 | Weighted Linear Discriminant Analysis Based on Class Saliency InformationabstractIn this paper, we propose a new variant of Linear Discriminant Analysis to overcome underlying drawbacks of traditional LDA and other LDA variants targeting problems involving imbalanced classes. Traditional LDA sets assumptions related to Gaussian class distribution and neglects influence of outlier classes, that might hurt in performance. We exploit intuitions coming from a probabilistic interpretation of visual saliency estimation in order to define saliency of a class in multi-class setting. Such information is then used to redefine the between-class and within-class scatters in a more robust manner. Compared to traditional LDA and other weight-based LDA variants, the proposed method has shown certain improvements on facial image classification problems in publicly available datasets. Lei Xu 0036, Alexandros Iosifidis, Moncef Gabbouj |
ICIP | 2 |
| 2018 | Subspace Support Vector Data DescriptionabstractThis paper proposes a novel method for solving one-class classification problems. The proposed approach, namely Subspace Support Vector Data Description, maps the data to a subspace that is optimized for one-class classification. In that feature space, the optimal hypersphere enclosing the target class is then determined. The method iteratively optimizes the data mapping along with data description in order to define a compact class representation in a low-dimensional feature space. We provide both linear and non-linear mappings for the proposed method. Experiments on 14 publicly available datasets indicate that the proposed Subspace Support Vector Data Description provides better performance compared to baselines and other recently proposed one-class classification methods. Fahad Sohrab, Jenni Raitoharju, Moncef Gabbouj, Alexandros Iosifidis |
ICPR | 4 |
| 2018 | Semi-supervised subclass support vector data description for image and video classification
Vasileios Mygdalis, Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas |
Neurocomputing | 2 |
| 2018 | Benchmark database for fine-grained image classification of benthic macroinvertebrates
Jenni Raitoharju, Ekaterina Riabchenko, Iftikhar Ahmad 0001, Alexandros Iosifidis, Moncef Gabbouj, Serkan Kiranyaz, Ville Tirronen, Johanna Ärje, Salme Kärkkäinen, Kristian Meissner |
Image Vis. Comput. | 4 |
| 2018 | Improving efficiency in convolutional neural networks with multilinear filtersabstractThe excellent performance of deep neural networks has enabled us to solve several automatization problems, opening an era of autonomous devices. However, current deep net architectures are heavy with millions of parameters and require billions of floating point operations. Several works have been developed to compress a pre-trained deep network to reduce memory footprint and, possibly, computation. Instead of compressing a pre-trained network, in this work, we propose a generic neural network layer structure employing multilinear projection as the primary feature extractor. The proposed architecture requires several times less memory as compared to the traditional Convolutional Neural Networks (CNN), while inherits the similar design principles of a CNN. In addition, the proposed architecture is equipped with two computation schemes that enable computation reduction or scalability. Experimental results show the effectiveness of our compact projection that outperforms traditional CNN, while requiring far fewer parameters. Dat Thanh Tran, Alexandros Iosifidis, Moncef Gabbouj |
Neural Networks | 2 |
| 2018 | Probabilistic saliency estimation
Çaglar Aytekin, Alexandros Iosifidis, Moncef Gabbouj |
Pattern Recognit. | 2 |
| 2018 | Generalized Multi-View Embedding for Visual Recognition and Cross-Modal RetrievalabstractIn this paper, the problem of multi-view embedding from different visual cues and modalities is considered. We propose a unified solution for subspace learning methods using the Rayleigh quotient, which is extensible for multiple views, supervised learning, and nonlinear embeddings. Numerous methods including canonical correlation analysis, partial least square regression, and linear discriminant analysis are studied using specific intrinsic and penalty graphs within the same framework. Nonlinear extensions based on kernels and (deep) neural networks are derived, achieving better performance than the linear ones. Moreover, a novel multi-view modular discriminant analysis is proposed by taking the view difference into consideration. We demonstrate the effectiveness of the proposed multi-view embedding methods on visual object recognition and cross-modal image retrieval, and obtain superior results in both applications compared to related methods. Guanqun Cao, Alexandros Iosifidis, Ke Chen 0004, Moncef Gabbouj |
IEEE Trans. Cybern. | 2 |
| 2017 | Class-specific kernel discriminant analysis based on Cholesky decompositionabstractIn this paper we describe a method for nonlinear class-specific discriminant learning that is based on Cholesky Decomposition. We show that the optimization problem solved in Class-Specific Kernel Discriminant Analysis is equivalent to that of Low-Rank Kernel Regression using training data independent target vectors. This connection allows us to devise a new Class-Specific Kernel Discriminant Analysis method that can be trained by exploiting fast linear system approaches, like the Cholesky decomposition. We verify our analysis in publicly available verification problems designed for human action recognition. Alexandros Iosifidis, Moncef Gabbouj |
IJCNN | 1 |
| 2017 | Generalized model of biological neural networks: Progressive operational perceptronsabstractTraditional Artificial Neural Networks (ANNs) such as Multi-Layer Perceptrons (MLPs) and Radial Basis Functions (RBFs) were designed to simulate biological neural networks; however, they are based only loosely on biology and only provide a crude model. This in turn yields well-known limitations and drawbacks on the performance and robustness. In this paper we shall address them by introducing a novel feed-forward ANN model, Generalized Operational Perceptrons (GOPs) that consist of neurons with distinct (non-)linear operators to achieve a generalized model of the biological neurons and ultimately a superior diversity. We modified the conventional back-propagation (BP) to train GOPs and furthermore, proposed Progressive Operational Perceptrons (POPs) to achieve self-organized and depth-adaptive GOPs according to the learning problem. The most crucial property of the POPs is their ability to simultaneously search for the optimal operator set and train each layer individually. The final POP is, therefore, formed layer by layer and this ability enables POPs with minimal network depth to attack the most challenging learning problems that cannot be learned by conventional ANNs even with a deeper and significantly complex configuration. Serkan Kiranyaz, Turker Ince, Alexandros Iosifidis, Moncef Gabbouj |
IJCNN | 3 |
| 2017 | The effect of automated taxa identification errors on biological indices
Johanna Ärje, Salme Kärkkäinen, Kristian Meissner, Alexandros Iosifidis, Turker Ince, Moncef Gabbouj, Serkan Kiranyaz |
Expert Syst. Appl. | 4 |
| 2017 | On the comparison of random and Hebbian weights for the training of single-hidden layer feedforward neural networks
Kaveh Samiee, Alexandros Iosifidis, Moncef Gabbouj |
Expert Syst. Appl. | 2 |
| 2017 | Approximate kernel extreme learning machine for large scale data classification
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas |
Neurocomputing | 1 |
| 2017 | Progressive Operational Perceptrons
Serkan Kiranyaz, Turker Ince, Alexandros Iosifidis, Moncef Gabbouj |
Neurocomputing | 3 |
| 2017 | CNN-based edge filtering for object proposals
Muhammad-Adeel Waris, Alexandros Iosifidis, Moncef Gabbouj |
Neurocomputing | 2 |
| 2017 | One-Class Classification Based on Extreme Learning and Geometric Class Information
Alexandros Iosifidis, Vasileios Mygdalis, Anastasios Tefas, Ioannis Pitas |
Neural Process. Lett. | 1 |
| 2017 | Learning graph affinities for spectral graph-based salient object detection
Çaglar Aytekin, Alexandros Iosifidis, Serkan Kiranyaz, Moncef Gabbouj |
Pattern Recognit. | 2 |
| 2017 | Multilinear class-specific discriminant analysisabstractThere has been a great effort to transfer linear discriminant techniques that operate on vector data to high-order data, generally referred to as Multilinear Discriminant Analysis (MDA) techniques. Many existing works focus on maximizing the inter-class variances to intra-class variances defined on tensor data representations. However, there has not been any attempt to employ class-specific discrimination criteria for the tensor data. In this paper, we propose a multilinear subspace learning technique suitable for applications requiring class-specific tensor models. The method maximizes the discrimination of each individual class in the feature space while retains the spatial structure of the input. We evaluate the efficiency of the proposed method on two problems, i.e. facial image analysis and stock price prediction based on limit order book data. Dat Thanh Tran, Moncef Gabbouj, Alexandros Iosifidis |
Pattern Recognit. Lett. | 3 |
| 2017 | Big Media Data Analysis
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas, Moncef Gabbouj |
Signal Process. Image Commun. | 1 |
| 2017 | Multi-View Nonparametric Discriminant Analysis for Image Retrieval and RecognitionabstractA novel multi-view nonparametric discriminant analysis method is proposed for the application of cross-modal image retrieval and zero-shot recognition. We exploit the class boundary structure and discrepancy information of the available views in order to formulate an optimization criterion, which is automatically adjusted to the multi-view class structures. The proposed method allows for multiple projection directions, by relaxing the Gaussian distribution assumption of related methods. The experiments demonstrate that the proposed method can achieve superior results comparing to several existing methods. Guanqun Cao, Alexandros Iosifidis, Moncef Gabbouj |
IEEE Signal Process. Lett. | 2 |
| 2017 | Class-Specific Kernel Discriminant Analysis Revisited: Further Analysis and ExtensionsabstractIn this paper, we revisit class-specific kernel discriminant analysis (KDA) formulation, which has been applied in various problems, such as human face verification and human action recognition. We show that the original optimization problem solved for the determination of class-specific discriminant projections is equivalent to a low-rank kernel regression (LRKR) problem using training data-independent target vectors. In addition, we show that the regularized version of class-specific KDA is equivalent to a regularized LRKR problem, exploiting the same targets. This analysis allows us to devise a novel fast solution. Furthermore, we devise novel incremental, approximate and deep (hierarchical) variants. The proposed methods are tested in human facial image and action video verification problems, where their effectiveness and efficiency is shown. Alexandros Iosifidis, Moncef Gabbouj |
IEEE Trans. Cybern. | 1 |
| 2016 | Supervised subspace learning based on deep randomized networksabstractIn this paper, we propose a supervised subspace learning method that exploits the rich representation power of deep feedforward networks. In order to derive a fast, yet efficient, learning scheme we employ deep randomized neural networks that have been recently shown to provide good compromise between training speed and performance. For optimally determining the learnt subspace, we formulate a regression problem where we employ target vectors designed to encode both the labeling information available for the training data and geometric properties of the training data, when represented in the feature space determined by the network's last hidden layer outputs. We experimentally show that the proposed approach is able to outperform deep randomized neural networks trained by using the standard network target vectors. Alexandros Iosifidis, Moncef Gabbouj |
ICASSP | 1 |
| 2016 | Combining multi-class maximum margin classification with linear discriminant analysis for human action recognitionabstractIn this paper, a new multi-class classification method is proposed and evaluated in the problem of human action recognition in unconstrained environments. The proposed method exploits both the maximum margin property of multi-class Support Vector Machines and Linear Discriminant Analysis-based discrimination. Experiments indicate that by exploiting such discriminant information in a multi-class maximum margin framework, classification performance can be enhanced, leading to state-of-the-art performance in human action recognition. Alexandros Iosifidis, Moncef Gabbouj |
ICIP | 1 |
| 2016 | One class classification applied in facial image analysisabstractIn this paper, we apply One-Class Classification methods in facial image analysis problems. We consider the cases where the available training data information originates from one class, or one of the available classes is of high importance. We propose a novel extension of the One-Class Extreme Learning Machines algorithm aiming at minimizing both the training error and the data dispersion and consider solutions that generate decision functions in the ELM space, as well as in ELM spaces of arbitrary dimensionality. We evaluate the performance in publicly available datasets. The proposed method compares favourably to other state-of-the-art choices. Vasileios Mygdalis, Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas |
ICIP | 2 |
| 2016 | Salient object segmentation based on linearly combined affinity graphsabstractIn this paper, we propose a graph affinity learning method for a recently proposed graph-based salient object detection method, namely Extended Quantum Cuts (EQCut). We exploit the fact that the output of EQCut is differentiable with respect to graph affinities, in order to optimize linear combination coefficients and parameters of several differentiable affinity functions by applying error backpropagation. We show that the learnt linear combination of affinities improves the performance over the baseline method and achieves comparable (or even better) performance when compared to the state-of-the-art salient object segmentation methods. Çaglar Aytekin, Alexandros Iosifidis, Serkan Kiranyaz, Moncef Gabbouj |
ICPR | 2 |
| 2016 | Learned vs. engineered features for fine-grained classification of aquatic macroinvertebratesabstractAquatic macroinvertebrate biomonitoring is an efficient way of assessment of slow and subtle anthropogenic changes and their effect on water quality. It is imperative to have reliable identification and counts of the various taxa occurring in samples as these form the basis for the quality indices used to infer the ecological status of the aquatic ecosystem. In this paper, we try to close the gap between human taxa identification accuracy (typically 90-95% on 30-40 classes of macroinvertebrates) and results of automatic fine-grained classification by introducing a novel technique based on Convolutional Neural Networks (CNN). CNN learns optimal features for macroinvertebrate classification and achieves near human accuracy when tested on 29 macroinvertebrate classes. Moreover, we perform comparative evaluation of the learned features against the hand-crafted features, which have been commonly used in classical approaches, and confirm superiority of the learned deep features over the engineered ones. Ekaterina Riabchenko, Kristian Meissner, Iftikhar Ahmad 0001, Alexandros Iosifidis, Ville Tirronen, Moncef Gabbouj, Serkan Kiranyaz |
ICPR | 4 |
| 2016 | Object proposals using CNN-based edge filteringabstractWith the success of deep learning in the last few years, the object detection community shifted from processing on exhaustive sliding windows to smaller set of object proposals using more powerful and deep visual representations. Object proposals increase the accuracy and speed up detection process by reducing the search space. In this paper we propose a novel idea of filtering irrelevant edges using semantic image filtering and true objectness learnt within convolutional layers of CNN. Our approach localizes well proposals by producing highly accurate bounding boxes and reduces the number of proposals. The greatest benefit of our approach is that it can be integrated into any existing method exploiting edge-based objectness to achieve consistently high recall across various intersection over union thresholds. Unlike other supervised methods, our approach does not require bounding box annotations for training. Experiments on PASCAL VOC 2007 dataset demonstrate that our approach improves the state-of-the-art model with a significant margin. Muhammad-Adeel Waris, Alexandros Iosifidis, Moncef Gabbouj |
ICPR | 2 |
| 2016 | Laplacian one class extreme learning machines for human action recognitionabstractA novel OCC method for human action recognition namely the Laplacian One Class Extreme Learning Machines is presented. The proposed method exploits local geometric data information within the OC-ELM optimization process. It is shown that emphasizing on preserving the local geometry of the data leads to a regularized solution, which models the target class more efficiently than the standard OC-ELM algorithm. The proposed method is extended to operate in feature spaces determined by the network hidden layer outputs, as well as in ELM spaces of arbitrary dimensions. Its superior performance against other OCC options is consistent among five publicly available human action recognition datasets. Vasileios Mygdalis, Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas |
MMSP | 2 |
| 2016 | Hierarchical class-specific kernel discriminant analysis for face verificationabstractIn this paper, a new method for nonlinear class-specific data projection is proposed for verification problems. We apply a hierarchical process formed by multiple nonlinear class-specific data projection layers in order to determine data representations in multiple subspaces enhancing class discrimination. We evaluate the proposed method on four publicly available facial image datasets and compare its performance with related methods and show its effectiveness. Alexandros Iosifidis, Moncef Gabbouj |
VCIP | 1 |
| 2016 | Exploiting stereoscopic disparity for augmenting human activity recognition performance
Ioannis Mademlis, Alexandros Iosifidis, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas |
Multim. Tools Appl. | 2 |
| 2016 | Multi-class Support Vector Machine classifiers using intrinsic and penalty graphs
Alexandros Iosifidis, Moncef Gabbouj |
Pattern Recognit. | 1 |
| 2016 | Nyström-based approximate kernel subspace learning
Alexandros Iosifidis, Moncef Gabbouj |
Pattern Recognit. | 1 |
| 2016 | Graph Embedded One-Class Classifiers for media data classification
Vasileios Mygdalis, Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas |
Pattern Recognit. | 2 |
| 2016 | Graph Embedded Extreme Learning MachineabstractIn this paper, we propose a novel extension of the extreme learning machine (ELM) algorithm for single-hidden layer feedforward neural network training that is able to incorporate subspace learning (SL) criteria on the optimization process followed for the calculation of the network's output weights. The proposed graph embedded ELM (GEELM) algorithm is able to naturally exploit both intrinsic and penalty SL criteria that have been (or will be) designed under the graph embedding framework. In addition, we extend the proposed GEELM algorithm in order to be able to exploit SL criteria in arbitrary (even infinite) dimensional ELM spaces. We evaluate the proposed approach on eight standard classification problems and nine publicly available datasets designed for three problems related to human behavior analysis, i.e., the recognition of human face, facial expression, and activity. Experimental results denote the effectiveness of the proposed approach, since it outperforms other ELM-based classification schemes in all the cases. Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas |
IEEE Trans. Cybern. | 1 |
| 2016 | Scaling Up Class-Specific Kernel Discriminant Analysis for Large-Scale Face VerificationabstractIn this paper, a novel approximate solution of the criterion used in non-linear class-specific discriminant subspace learning is proposed. We build on the class-specific kernel spectral regression method, which is a two-step process formed by an eigenanalysis step and a kernel regression step. Based on the structure of the intra-class and out-of-class scatter matrices, we provide a fast solution for the first step. For the second step, we propose the use of approximate kernel space definitions. We analytically show that the adoption of randomized and class-specific kernels has the effect of regularization and Nyström-based approximation, respectively. We evaluate the proposed approach in face verification problems and compare it with the existing approaches. Experimental results show the effectiveness and efficiency of the proposed approximate class-specific kernel spectral regression method, since it can provide satisfactory performance and scale well with the size of the data. Alexandros Iosifidis, Moncef Gabbouj |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2016 | Visual Voice Activity Detection in the WildabstractThe visual voice activity detection (V-VAD) problem in unconstrained environments is investigated in this paper. A novel method for V-VAD in the wild, exploiting local shape and motion information appearing at spatiotemporal locations of interest for facial video segment description and the bag of words model for facial video segment representation, is proposed. Facial video segment classification is subsequently performed using the state-of-the-art classification algorithms. Experimental results on one publicly available V-VAD dataset denote the effectiveness of the proposed method, since it achieves better generalization performance in unseen users, when compared to the recently proposed state-of-the-art methods. Additional results on a new unconstrained dataset provide evidence that the proposed method can be effective even in such cases in which any other existing method fails. Fotini Patrona, Alexandros Iosifidis, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas |
IEEE Trans. Multim. | 2 |
| 2015 | Enhancing class discrimination in Kernel Discriminant AnalysisabstractIn this paper, we propose an optimization scheme aiming at optimal nonlinear data projection, in terms of Fisher ratio maximization. To this end, we formulate an iterative optimization scheme consisting of two processing steps: optimal data projection calculation and optimal class representation determination. Compared to the standard approach employing the class mean vectors for class representation, the proposed optimization scheme increases class discrimination in the reduced-dimensionality feature space. We evaluate the proposed method in standard classification problems, as well as on the classification of human actions and face, and show that it is able to achieve better generalization performance, when compared to the standard approach. Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas |
ICASSP | 1 |
| 2015 | Exploiting subclass information in one-class support vector machine for video summarizationabstractIn this paper, we propose a method for video summarization based on human activity description. We formulate this problem as the one of automatic video segment selection based on a learning process that employs salient video segment paradigms. For this one-class classification problem, we introduce a novel variant of the One-Class Support Vector Machine (OC-SVM) classifier that exploits subclass information in the OC-SVM optimization problem, in order to jointly minimize the data dispersion within each subclass and determine the optimal decision function. We evaluate the proposed approach in three Hollywood movies, where the performance of the proposed SOC-SVM algorithm is compared with that of the OC-SVM. Experimental results denote that the proposed approach is able to outperform OC-SVM-based video segment selection. Vasileios Mygdalis, Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas |
ICASSP | 2 |
| 2015 | Merging linear discriminant analysis with Bag of Words model for human action recognitionabstractIn this paper we propose a novel method for human action recognition, that unifies discriminative Bag of Words (BoW)-based video representation and discriminant subspace learning. An iterative optimization scheme is proposed for sequential discriminant BoWs-based action representation and code-book adaptation based on action discrimination in a reduced dimensionality feature space where action classes are better discriminated. Experiments on four publicly available action recognition data sets demonstrate that the proposed unified approach increases the discriminative ability of the obtained video representation, providing enhanced action classification performance. Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas |
ICIP | 1 |
| 2015 | Large-scale nonlinear facial image classification based on approximate kernel Extreme Learning MachineabstractIn this paper, we propose a scheme that can be used in large-scale nonlinear facial image classification problems. An approximate solution of the kernel Extreme Learning Machine classifier is formulated and evaluated. Experiments on two publicly available facial image datasets using two popular facial image representations illustrate the effectiveness and efficiency of the proposed approach. The proposed Approximate Kernel Extreme Learning Machine classifier is able to scale well in both time and memory, while achieving good generalization performance. Specifically, it is shown that it outperforms the standard ELM approach for the same time and memory requirements. Compared to the original kernel ELM approach, it achieves similar (or better) performance, while scaling well in both time and memory with respect to the training set cardinality. Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas |
ICIP | 1 |
| 2015 | Visual voice activity detection based on spatiotemporal information and bag of wordsabstractA novel method for Visual Voice Activity Detection (V-VAD) that exploits local shape and motion information appearing at spatiotemporal locations of interest for facial region video description and the Bag of Words (BoW) model for facial region video representation is proposed in this paper. Facial region video classification is subsequently performed based on Single-hidden Layer Feedforward Neural (SLFN) network trained by applying the recently proposed kernel Extreme Learning Machine (kELM) algorithm on training facial videos depicting talking and non-talking persons. Experimental results on two publicly available V-VAD data sets, denote the effectiveness of the proposed method, since better generalization performance in unseen users is achieved, compared to recently proposed state-of-the-art methods. Fotini Patrona, Alexandros Iosifidis, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas |
ICIP | 2 |
| 2015 | Extreme learning machine based supervised subspace learning
Alexandros Iosifidis |
Neurocomputing | 1 |
| 2015 | Distance-based human action recognition using optimized class representations
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas |
Neurocomputing | 1 |
| 2015 | DropELM: Fast neural network regularization with Dropout and DropConnect
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas |
Neurocomputing | 1 |
| 2015 | On the kernel Extreme Learning Machine speedup
Alexandros Iosifidis, Moncef Gabbouj |
Pattern Recognit. Lett. | 1 |
| 2015 | On the kernel Extreme Learning Machine classifier
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas |
Pattern Recognit. Lett. | 1 |
| 2015 | Sparse extreme learning machine classifier exploiting intrinsic graphs
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas |
Pattern Recognit. Lett. | 1 |
| 2015 | Class-Specific Reference Discriminant Analysis With Application in Human Behavior AnalysisabstractIn this paper, a novel nonlinear subspace learning technique for class-specific data representation is proposed. A novel data representation is obtained by applying nonlinear class-specific data projection to a discriminant feature space, where the data belonging to the class under consideration are enforced to be close to their class representation, while the data belonging to the remaining classes are enforced to be as far as possible from it. A class is represented by an optimized class vector, enhancing class discrimination in the resulting feature space. An iterative optimization scheme is proposed to this end, where both the optimal nonlinear data projection and the optimal class representation are determined in each optimization step. The proposed approach is tested on three problems relating to human behavior analysis: Face recognition, facial expression recognition, and human action recognition. Experimental results denote the effectiveness of the proposed approach, since the proposed class-specific reference discriminant analysis outperforms kernel discriminant analysis, kernel spectral regression, and class-specific kernel discriminant analysis, as well as support vector machine-based classification, in most cases. Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas |
IEEE Trans. Hum. Mach. Syst. | 1 |
| 2014 | Stereoscopic video description for human action recognitionabstractIn this paper, a stereoscopic video description method is proposed that indirectly incorporates scene geometry information derived from stereo disparity, through the manipulation of video interest points. This approach is flexible and able to cooperate with any monocular low-level feature descriptor. The method is evaluated on the problem of recognizing complex human actions in natural settings, using a publicly available action recognition database of unconstrained stereoscopic 3D videos, coming from Hollywood movies. It is compared both against competing depth-aware approaches and a state-of-the-art monocular algorithm. Experimental results denote that the proposed approach outperforms them and achieves state-of-the-art performance. Ioannis Mademlis, Alexandros Iosifidis, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas |
CIMSIVP | 2 |
| 2014 | Minimum Variance Extreme Learning Machine for human action recognitionabstractIn this paper we propose an algorithm for Single-hidden Layer Feedforward Neural networks training. Based on the observation that the learning process of such networks can be considered to be a non-linear mapping of the training data to a high-dimensional feature space, followed by a data projection process to a low-dimensional space where classification is performed by a linear classifier, we extend the Extreme Learning Machine (ELM) algorithm in order to exploit the training data dispersion in its optimization process. The proposed Minimum Variance Extreme Learning Machine classifier is evaluated in human action recognition, where we compare its performance with that of other ELM-based classifiers, as well as the kernel Support Vector Machine classifier. Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas |
ICASSP | 1 |
| 2014 | Human action recognition based on bag of features and multi-view neural networksabstractIn this paper, we employ Single-hidden Layer Feedforward Neural networks in order to perform human action recognition based on multiple action representations. In order to determine both optimized network and action representation combination weights, we propose an optimization process that jointly minimizes the overall network training error and the within-class variance of the training data in the corresponding hidden layer spaces. The proposed approach has been evaluated by using the state-of-the-art Bag of Features-based action video representation on three publicly available action recognition databases, where it outperforms two commonly used video representation combination approaches, as well as the best single-descriptor classification outcome. Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas |
ICIP | 1 |
| 2014 | Semi-supervised Classification of Human Actions Based on Neural NetworksabstractIn this paper, we propose a novel algorithm for Single-hidden Layer Feed forward Neural networks training which is able to exploit information coming from both labeled and unlabeled data for semi-supervised action classification. We extend the Extreme Learning Machine algorithm by incorporating appropriate regularization terms describing geometric properties and discrimination criteria of the training data representation in the ELM space to this end. The proposed algorithm is evaluated on human action recognition, where its performance is compared with that of other (semi-)supervised classification schemes. Experimental results on two publicly available action recognition databases denote its effectiveness. Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas |
ICPR | 1 |
| 2014 | Regularized extreme learning machine for multi-view semi-supervised action recognition
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas |
Neurocomputing | 1 |
| 2014 | Kernel Reference Discriminant Analysis
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas |
Pattern Recognit. Lett. | 1 |
| 2014 | Discriminant Bag of Words based representation for human action recognition
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas |
Pattern Recognit. Lett. | 1 |
| 2013 | Neural Networks for Digital Media Analysis and Description
Anastasios Tefas, Alexandros Iosifidis, Ioannis Pitas |
EANN (1) | 2 |
| 2013 | Active classification for human action recognitionabstractIn this paper, we propose a novel classification method involving two processing steps. Given a test sample, the training data residing to its neighborhood are determined. Classification is performed by a Single-hidden Layer Feedforward Neural network exploiting labeling information of the training data appearing in the test sample neighborhood and using the rest training data as unlabeled. By following this approach, the proposed classification method focuses the classification problem on the training data that are more similar to the test sample under consideration and exploits information concerning to the training set structure. Compared to both static classification exploiting all the available training data and dynamic classification involving data selection for classification, the proposed active classification method provides enhanced classification performance in two publicly available action recognition databases. Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas |
ICIP | 1 |
| 2013 | Learning sparse representations for view-independent human action recognition based on fuzzy distances
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas |
Neurocomputing | 1 |
| 2013 | Dynamic action recognition based on dynemes and Extreme Learning Machine
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas |
Pattern Recognit. Lett. | 1 |
| 2013 | Multi-view action recognition based on action volumes, fuzzy distances and cluster discriminant analysis
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas |
Signal Process. | 1 |
| 2013 | Minimum Class Variance Extreme Learning Machine for Human Action RecognitionabstractIn this paper, we propose a novel method aiming at view-independent human action recognition. Action description is based on local shape and motion information appearing at spatiotemporal locations of interest in a video. Action representation involves fuzzy vector quantization, while action classification is performed by a feedforward neural network. A novel classification algorithm, called minimum class variance extreme learning machine, is proposed in order to enhance the action classification performance. The proposed method can successfully operate in situations that may appear in real application scenarios, since it does not set any assumption concerning the visual scene background and the camera view angle. Experimental results on five publicly available databases, aiming at different application scenarios, denote the effectiveness of both the adopted action recognition approach and the proposed minimum class variance extreme learning machine algorithm. Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2013 | Multidimensional Sequence Classification Based on Fuzzy Distances and Discriminant AnalysisabstractIn this paper, we present a novel method aiming at multidimensional sequence classification. We propose a novel sequence representation, based on its fuzzy distances from optimal representative signal instances, called statemes. We also propose a novel modified clustering discriminant analysis algorithm minimizing the adopted criterion with respect to both the data projection matrix and the class representation, leading to the optimal discriminant sequence class representation in a low-dimensional space, respectively. Based on this representation, simple classification algorithms, such as the nearest subclass centroid, provide high classification accuracy. A three step iterative optimization procedure for choosing statemes, optimal discriminant subspace and optimal sequence class representation in the final decision space is proposed. The classification procedure is fast and accurate. The proposed method has been tested on a wide variety of multidimensional sequence classification problems, including handwritten character recognition, time series classification and human activity recognition, providing very satisfactory classification results. Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2013 | On the Optimal Class Representation in Linear Discriminant AnalysisabstractLinear discriminant analysis (LDA) is a widely used technique for supervised feature extraction and dimensionality reduction. LDA determines an optimal discriminant space for linear data projection based on certain assumptions, e.g., on using normal distributions for each class and employing class representation by the mean class vectors. However, there might be other vectors that can represent each class, to increase class discrimination. In this brief, we propose an optimization scheme aiming at the optimal class representation, in terms of Fisher ratio maximization, for LDA-based data projection. Compared with the standard LDA approach, the proposed optimization scheme increases class discrimination in the reduced dimensionality space and achieves higher classification rates in publicly available data sets. Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2012 | Eating and drinking activity recognition based on discriminant analysis of fuzzy distances and activity volumesabstractEating and drinking activity recognition can be considered a solitary research field in activity recognition area. The development of an application capable to identify human eating and drinking activity can be really useful in a smart home environment targeting to extend independent living of older persons in the early stages of dementia. In this paper a novel method aiming at eating and drinking activity recognition is presented. Activities are considered as a sequence of human body poses forming 3D volumes, in which the third dimension refers to time. Fuzzy Vector Quantization is performed to associate the 3D volume representation of an activity video with 3D volume prototypes and Linear Discriminant Analysis is used to map activity representations in a low dimensional discriminant feature space. In this space a simple Nearest Centroid classification procedure leads to very satisfactory classification results. Alexandros Iosifidis, Ermioni Marami, Anastasios Tefas, Ioannis Pitas |
ICASSP | 1 |
| 2012 | Discriminant action representation for view-invariant person identificationabstractIn this paper we propose a novel person identification method exploiting human motion information. Persons are described by using their poses during action execution. Identification process involves Fuzzy Vector Quantization and Discriminant Learning. In the case of multiple cameras used in the identification phase, single-view identification results combination is achieved by employing a Bayesian combination strategy. The proposed identification approach does not set the assumptions of known action class and number of capturing cameras in the identification phase. Experimental results on two publicly available video databases denote the effectiveness of the proposed approach. Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas |
ICIP | 1 |
| 2012 | Neural representation and learning for multi-view human action recognitionabstractIn this paper we propose a novel method aiming at view-independent multi-view action recognition. Instead of combining the information provided by all the cameras forming the camera setup, for action representation and classification, we perform single-view action representation and classification to all the available videos depicting the person under consideration independently. Action representation involves a self organizing neural network training followed by fuzzy vector quantization. Action classification is performed by a feedforward neural network which is trained for view-invariant action recognition. Multiple action classification results combination based on Bayesian learning, in the recognition phase, results to high action recognition accuracy. The performance of the proposed action recognition method is evaluated on two publicly available databases, aiming at different application scenarios. Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas |
IJCNN | 1 |
| 2012 | Multi-view human movement recognition based on fuzzy distances and linear discriminant analysis
Alexandros Iosifidis, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas |
Comput. Vis. Image Underst. | 1 |
| 2012 | Activity-Based Person Identification Using Fuzzy Representation and Discriminant LearningabstractIn this paper, a novel view invariant person identification method based on human activity information is proposed. Unlike most methods proposed in the literature, in which “walk” (i.e., gait) is assumed to be the only activity exploited for person identification, we incorporate several activities in order to identify a person. A multicamera setup is used to capture the human body from different viewing angles. Fuzzy vector quantization and linear discriminant analysis are exploited in order to provide a discriminant activity representation. Person identification, activity recognition, and viewing angle specification results are obtained for all the available cameras independently. By properly combining these results, a view-invariant activity-independent person identification method is obtained. The proposed approach has been tested in challenging problem setups, simulating real application situations. Experimental results are very promising. Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2012 | View-Invariant Action Recognition Based on Artificial Neural NetworksabstractIn this paper, a novel view invariant action recognition method based on neural network representation and recognition is proposed. The novel representation of action videos is based on learning spatially related human body posture prototypes using self organizing maps. Fuzzy distances from human body posture prototypes are used to produce a time invariant action representation. Multilayer perceptrons are used for action classification. The algorithm is trained using data from a multi-camera setup. An arbitrary number of cameras can be used in order to recognize actions using a Bayesian framework. The proposed method can also be applied to videos depicting interactions between humans, without any modification. The use of information captured from different viewing angles leads to high classification performance. The proposed method is the first one that has been tested in challenging experimental setups, a fact that denotes its effectiveness to deal with most of the open issues in action recognition. Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2010 | Movement recognition exploiting multi-view informationabstractIn this paper a novel view-invariant movement recognition method is presented. A multi-camera setup is used to capture the movement from different observation angles. Identification of the position of each camera with respect to the subject's body is achieved by a procedure based on morphological operations and the proportions of the human body. Binary body masks from frames of all cameras, consistently arranged through the previous procedure, are concatenated to produce the so-called multi-view binary mask. These masks are rescaled and vectorized to create feature vectors in the input space. Fuzzy vector quantization is performed to associate input feature vectors with movement representations and linear discriminant analysis is used to map movements in a low dimensionality discriminant feature space. Experimental results show that the method can achieve very satisfactory recognition rates. Alexandros Iosifidis, Nikos Nikolaidis 0001, Ioannis Pitas |
MMSP | 1 |