EDBT 2026 Demo / reviewers in the wild / expert
Decebal Constantin Mocanu
dblp:133/7764
· DBLP profile ↗
50ranked-venue papers
9as first author
27since 2021 · last 2025
0000-0002-5636-7683ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 34 · 4 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 first-authorComputer networks · 2 · 1 first-authorSystems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Dynamic Sparse Training versus Dense Training: The Unexpected Winner in Image Corruption RobustnessabstractIt is generally perceived that Dynamic Sparse Training opens the door to a new era of scalability and efficiency for artificial neural networks at, perhaps, some costs in accuracy performance for the classification task. At the same time, Dense Training is widely accepted as being the "de facto" approach to train artificial neural networks if one would like to maximize their robustness against image corruption. In this paper, we question this general practice. Consequently, \textit{we claim that}, contrary to what is commonly thought, the Dynamic Sparse Training methods can consistently outperform Dense Training in terms of robustness accuracy, particularly if the efficiency aspect is not considered as a main objective (i.e., sparsity levels between 10\% and up to 50\%), without adding (or even reducing) resource cost. We validate our claim on two types of data, images and videos, using several traditional and modern deep learning architectures for computer vision and three widely studied Dynamic Sparse Training algorithms. Our findings reveal a new yet-unknown benefit of Dynamic Sparse Training and open new possibilities in improving deep learning robustness beyond the current state of the art. Boqian Wu, Qiao Xiao, Shunxin Wang, Nicola Strisciuglio, Mykola Pechenizkiy, Maurice van Keulen, Decebal Constantin Mocanu, Elena Mocanu |
ICLR | 7 |
| 2025 | Boosting Robustness in Preference-Based Reinforcement Learning with Dynamic Sparsity
Calarina Muslimani, Bram Grooten, Deepak Ranganatha Sastry Mamillapalli, Mykola Pechenizkiy, Decebal Constantin Mocanu, Matthew E. Taylor |
AAMAS | 5 |
| 2024 | Supervised Feature Selection via Ensemble Gradient Information from Sparse Neural NetworksabstractFeature selection algorithms aim to select a subset of informative features from a dataset to reduce the data dimensionality, consequently saving resource consumption and improving the model’s performance and interpretability. In recent years, feature selection based on neural networks has become a new trend, demonstrating superiority over traditional feature selection methods. However, most existing methods use dense neural networks to detect informative features, which requires significant computational and memory overhead. In this paper, taking inspiration from the successful application of local sensitivity analysis on neural networks, we propose a novel resource-efficient supervised feature selection algorithm based on sparse multi-layer perceptron called “GradEnFS". By utilizing the gradient information of various sparse models from different training iterations, our method successfully detects the informative feature subset. We performed extensive experiments on nine classification datasets spanning various domains to evaluate the effectiveness of our method. The results demonstrate that our proposed approach outperforms the state-of-the-art methods in terms of selecting informative features while saving resource consumption substantially. Moreover, we show that using a sparse neural network for feature selection not only alleviates resource consumption but also has a significant advantage over other methods when performing feature selection on noisy datasets. Kaiting Liu, Zahra Atashgahi, Ghada Sokar, Mykola Pechenizkiy, Decebal Constantin Mocanu |
AISTATS | 5 |
| 2024 | Are Sparse Neural Networks Better Hard Sample Learners?
Qiao Xiao, Boqian Wu, Lu Yin 0006, Christopher Neil Gadzinski, Tianjin Huang, Mykola Pechenizkiy, Decebal Constantin Mocanu |
BMVC | 7 |
| 2024 | Unveiling the Power of Sparse Neural Networks for Feature SelectionabstractSparse Neural Networks (SNNs) have emerged as powerful tools for efficient feature selection. Leveraging the dynamic sparse training (DST) algorithms within SNNs has demonstrated promising feature selection capabilities while drastically reducing computational overheads. Despite these advancements, several critical aspects remain insufficiently explored for feature selection. Questions persist regarding the choice of the DST algorithm for network training, the choice of metric for ranking features/neurons, and the comparative performance of these methods across diverse datasets when compared to dense networks. This paper addresses these gaps by presenting a comprehensive systematic analysis of feature selection with sparse neural networks. Moreover, we introduce a novel metric considering sparse neural network characteristics, which is designed to quantify feature importance within the context of SNNs. Our findings show that feature selection with SNNs trained with DST algorithms can achieve, on average, more than 50% memory and 55% FLOPs reduction compared to the dense networks, while outperforming them in terms of the quality of the selected features. Our code and the supplementary material are available on GitHub (https://github.com/zahraatashgahi/Neuron-Attribution). Zahra Atashgahi, Tennison Liu, Mykola Pechenizkiy, Raymond N. J. Veldhuis, Decebal Constantin Mocanu, Mihaela van der Schaar |
ECAI | 5 |
| 2024 | Dynamic Data Pruning for Automatic Speech RecognitionabstractThe recent success of Automatic Speech Recognition (ASR) is largely attributed to the ever-growing amount of training data. However, this trend has made model training prohibitively costly and imposed computational demands. While data pruning has been proposed to mitigate this issue by identifying a small subset of relevant data, its application in ASR has been barely explored, and existing works often entail significant overhead to achieve meaningful results. To fill this gap, this paper presents the first investigation of dynamic data pruning for ASR, finding that we can reach the full-data performance by dynamically selecting 70% of data. Furthermore, we introduce Dynamic Data Pruning for ASR (DDP-ASR), which offers several fine-grained pruning granularities specifically tailored for speech-related datasets, going beyond the conventional pruning of entire time sequences. Our intensive experiments show that DDP-ASR can save up to 1.6x training time with negligible performance loss. Qiao Xiao, Pingchuan Ma 0001, Adriana Fernandez-Lopez, Boqian Wu, Lu Yin 0006, Stavros Petridis, Mykola Pechenizkiy, Maja Pantic, Decebal Constantin Mocanu, Shiwei Liu 0003 |
INTERSPEECH | 9 |
| 2024 | E2ENet: Dynamic Sparse Feature Fusion for Accurate and Efficient 3D Medical Image SegmentationabstractDeep neural networks have evolved as the leading approach in 3D medical image segmentation due to their outstanding performance. However, the ever-increasing model size and computational cost of deep neural networks have become the primary barriers to deploying them on real-world, resource-limited hardware. To achieve both segmentation accuracy and efficiency, we propose a 3D medical image segmentation model called Efficient to Efficient Network (E2ENet), which incorporates two parametrically and computationally efficient designs. i. Dynamic sparse feature fusion (DSFF) mechanism: it adaptively learns to fuse informative multi-scale features while reducing redundancy. ii. Restricted depth-shift in 3D convolution: it leverages the 3D spatial information while keeping the model and computational complexity as 2D-based methods. We conduct extensive experiments on AMOS, Brain Tumor Segmentation and BTCV Challenge, demonstrating that E2ENet consistently achieves a superior trade-off between accuracy and efficiency than prior arts across various resource constraints. %In particular, with a single model and single scale, E2ENet achieves comparable accuracy on the large-scale challenge AMOS-CT, while saving over 69% parameter count and 27% FLOPs in the inference phase, compared with the previous
best-performing method. Our code has been made available at: https://github.com/boqian333/E2ENet-Medical. Boqian Wu, Qiao Xiao, Shiwei Liu 0003, Lu Yin 0006, Mykola Pechenizkiy, Decebal Constantin Mocanu, Maurice van Keulen, Elena Mocanu |
NeurIPS | 6 |
| 2024 | Adaptive Sparsity Level During Training for Efficient Time Series Forecasting with Transformers
Zahra Atashgahi, Mykola Pechenizkiy, Raymond N. J. Veldhuis, Decebal Constantin Mocanu |
ECML/PKDD (1) | 4 |
| 2023 | More ConvNets in the 2020s: Scaling up Kernels Beyond 51x51 using Sparsity
Shiwei Liu 0003, Tianlong Chen 0001, Xiaohan Chen 0001, Xuxi Chen, Qiao Xiao, Boqian Wu, Tommi Kärkkäinen, Mykola Pechenizkiy, Decebal Constantin Mocanu, Zhangyang Wang |
ICLR | 9 |
| 2023 | Fantastic Weights and How to Find Them: Where to Prune in Dynamic Sparse TrainingabstractDynamic Sparse Training (DST) is a rapidly evolving area of research that seeks to optimize the sparse initialization of a neural network by adapting its topology during training. It has been shown that under specific conditions, DST is able to outperform dense models. The key components of this framework are the pruning and growing criteria, which are repeatedly applied during the training process to adjust the network’s sparse connectivity. While the growing criterion's impact on DST performance is relatively well studied, the influence of the pruning criterion remains overlooked. To address this issue, we design and perform an extensive empirical analysis of various pruning criteria to better understand their impact on the dynamics of DST solutions. Surprisingly, we find that most of the studied methods yield similar results. The differences become more significant in the low-density regime, where the best performance is predominantly given by the simplest technique: magnitude-based pruning. Aleksandra Nowak 0001, Bram Grooten, Decebal Constantin Mocanu, Jacek Tabor |
NeurIPS | 3 |
| 2023 | Hu-bot: promoting the cooperation between humans and mobile robotsabstractAbstract This paper investigates human–robot collaboration in a novel setup: a human helps a mobile robot that can move and navigate freely in an environment. Specifically, the human helps by remotely taking over control during the learning of a task. The task is to find and collect several items in a walled arena, and Reinforcement Learning is used to seek a suitable controller. If the human observes undesired robot behavior, they can directly issue commands for the wheels through a game joystick. Experiments in a simulator showed that human assistance improved robot behavior efficacy by 30% and efficiency by 12%. The best policies were also tested in real life, using physical robots. Hardware experiments showed no significant difference concerning the simulations, providing empirical validation of our approach in practice. Karine Miras, Decebal Constantin Mocanu, A. E. Eiben |
Neural Comput. Appl. | 2 |
| 2022 | Deep Ensembling with No Overhead for either Training or Testing: The All-Round Blessings of Dynamic Sparsity
Shiwei Liu 0003, Tianlong Chen 0001, Zahra Atashgahi, Xiaohan Chen 0001, Ghada Sokar, Elena Mocanu, Mykola Pechenizkiy, Zhangyang Wang, Decebal Constantin Mocanu |
ICLR | 9 |
| 2022 | The Unreasonable Effectiveness of Random Pruning: Return of the Most Naive Baseline for Sparse Training
Shiwei Liu 0003, Tianlong Chen 0001, Xiaohan Chen 0001, Li Shen 0008, Decebal Constantin Mocanu, Zhangyang Wang, Mykola Pechenizkiy |
ICLR | 5 |
| 2022 | Dynamic Sparse Training for Deep Reinforcement LearningabstractDeep reinforcement learning (DRL) agents are trained through trial-and-error interactions with the environment. This leads to a long training time for dense neural networks to achieve good performance. Hence, prohibitive computation and memory resources are consumed. Recently, learning efficient DRL agents has received increasing attention. Yet, current methods focus on accelerating inference time. In this paper, we introduce for the first time a dynamic sparse training approach for deep reinforcement learning to accelerate the training process. The proposed approach trains a sparse neural network from scratch and dynamically adapts its topology to the changing data distribution during training. Experiments on continuous control tasks show that our dynamic sparse agents achieve higher performance than the equivalent dense methods, reduce the parameter count and floating-point operations (FLOPs) by 50%, and have a faster learning speed that enables reaching the performance of dense agents with 40−50% reduction in the training steps. Ghada Sokar, Elena Mocanu, Decebal Constantin Mocanu, Mykola Pechenizkiy, Peter Stone 0001 |
IJCAI | 3 |
| 2022 | Where to Pay Attention in Sparse Training for Feature Selection?abstractA new line of research for feature selection based on neural networks has recently emerged. Despite its superiority to classical methods, it requires many training iterations to converge and detect the informative features. For datasets with a large number of samples or a very high dimensional feature space, the computational time becomes prohibitively long. In this paper, we present a new efficient unsupervised method for feature selection based on sparse autoencoders. In particular, we propose a new sparse training algorithm that optimizes a model's sparse topology during training to quickly pay attention to informative features. The attention-based adaptation of the sparse topology enables fast detection of informative features after a few training iterations. We performed extensive experiments on 10 datasets of different types, including image, speech, text, artificial, and biological. They cover a wide range of characteristics, such as low and high-dimensional feature spaces, as well as few and large training samples. Our proposed approach outperforms the state-of-the-art methods in terms of the selection of informative features while reducing training iterations and computational costs substantially. Moreover, the experiments show the robustness of our method in extremely noisy environments. Ghada Sokar, Zahra Atashgahi, Mykola Pechenizkiy, Decebal Constantin Mocanu |
NeurIPS | 4 |
| 2022 | Dynamic Sparse Network for Time Series Classification: Learning What to "See"abstractThe receptive field (RF), which determines the region of time series to be “seen” and used, is critical to improve the performance for time series classification (TSC). However, the variation of signal scales across and within time series data, makes it challenging to decide on proper RF sizes for TSC. In this paper, we propose a dynamic sparse network (DSN) with sparse connections for TSC, which can learn to cover various RF without cumbersome hyper-parameters tuning. The kernels in each sparse layer are sparse and can be explored under the constraint regions by dynamic sparse training, which makes it possible to reduce the resource cost. The experimental results show that the proposed DSN model can achieve state-of-art performance on both univariate and multivariate TSC datasets with less than 50% computational cost compared with recent baseline methods, opening the path towards more accurate resource-aware methods for time series analyses. Our code is publicly available at: https://github.com/QiaoXiao7282/DSN. Qiao Xiao, Boqian Wu, Yu Zhang 0006, Shiwei Liu 0003, Mykola Pechenizkiy, Elena Mocanu, Decebal Constantin Mocanu |
NeurIPS | 7 |
| 2022 | Avoiding Forgetting and Allowing Forward Transfer in Continual Learning via Sparse Networks
Ghada Sokar, Decebal Constantin Mocanu, Mykola Pechenizkiy |
ECML/PKDD (3) | 2 |
| 2022 | A brain-inspired algorithm for training highly sparse neural networksabstractAbstract Sparse neural networks attract increasing interest as they exhibit comparable performance to their dense counterparts while being computationally efficient. Pruning the dense neural networks is among the most widely used methods to obtain a sparse neural network. Driven by the high training cost of such methods that can be unaffordable for a low-resource device, training sparse neural networks sparsely from scratch has recently gained attention. However, existing sparse training algorithms suffer from various issues, including poor performance in high sparsity scenarios, computing dense gradient information during training, or pure random topology search. In this paper, inspired by the evolution of the biological brain and the Hebbian learning theory, we present a new sparse training approach that evolves sparse neural networks according to the behavior of neurons in the network. Concretely, by exploiting the cosine similarity metric to measure the importance of the connections, our proposed method, “Cosine similarity-based and random topology exploration (CTRE)”, evolves the topology of sparse neural networks by adding the most important connections to the network without calculating dense gradient in the backward. We carried out different experiments on eight datasets, including tabular, image, and text datasets, and demonstrate that our proposed method outperforms several state-of-the-art sparse training algorithms in extremely sparse neural networks by a large gap. The implementation code is available on Github. Zahra Atashgahi, Joost Pieterse, Shiwei Liu 0003, Decebal Constantin Mocanu, Raymond N. J. Veldhuis, Mykola Pechenizkiy |
Mach. Learn. | 4 |
| 2022 | Quick and robust feature selection: the strength of energy-efficient sparse training for autoencodersabstractAbstract Major complications arise from the recent increase in the amount of high-dimensional data, including high computational costs and memory requirements. Feature selection, which identifies the most relevant and informative attributes of a dataset, has been introduced as a solution to this problem. Most of the existing feature selection methods are computationally inefficient; inefficient algorithms lead to high energy consumption, which is not desirable for devices with limited computational and energy resources. In this paper, a novel and flexible method for unsupervised feature selection is proposed. This method, named QuickSelection (The code is available at: https://github.com/zahraatashgahi/QuickSelection), introduces the strength of the neuron in sparse neural networks as a criterion to measure the feature importance. This criterion, blended with sparsely connected denoising autoencoders trained with the sparse evolutionary training procedure, derives the importance of all input features simultaneously. We implement QuickSelection in a purely sparse manner as opposed to the typical approach of using a binary mask over connections to simulate sparsity. It results in a considerable speed increase and memory reduction. When tested on several benchmark datasets, including five low-dimensional and three high-dimensional datasets, the proposed method is able to achieve the best trade-off of classification and clustering accuracy, running time, and maximum memory usage, among widely used approaches for feature selection. Besides, our proposed method requires the least amount of energy among the state-of-the-art autoencoder-based feature selection methods. Zahra Atashgahi, Ghada Sokar, Tim van der Lee, Elena Mocanu, Decebal Constantin Mocanu, Raymond N. J. Veldhuis, Mykola Pechenizkiy |
Mach. Learn. | 5 |
| 2022 | Situation-Aware Drivable Space Estimation for Automated DrivingabstractAn automated vehicle (AV) must always have a correct representation of the drivable space to position itself accurately and operate safely. To determine the drivable space, current research focuses on single sources of information, either using pre-computed high-definition maps, or mapping the environment online with sensors such as LiDARs or cameras. However, each of these information sources can fail, some are too costly, and maps could be outdated. In this work a new method for situation-aware drivable space (SDS) estimation combining multiple information sources is proposed, which is also suitable for AVs equipped with inexpensive sensors. Depending on the situation, semantic information of sensed objects is combined with domain knowledge to estimate the drivability of the space surrounding each object (e.g. traffic light, another vehicle). These estimates are modeled as probabilistic graphs to account for the uncertainty of information sources, and an optimal spatial configuration of their elements is determined via graph-based simultaneous localization and mapping (SLAM). To investigate the robustness of SDS towards potentially unreliable sensors and maps, it has been tested in a simulation environment and real world data. Results on different use cases (e.g. straight roads, curved roads, and intersections) show considerable robustness towards unreliable inputs, and the recovered drivable space allows for accurate in-lane localization of the AV even in extreme cases where no prior knowledge of the road network is available. Manuel Muñoz Sánchez, Denis Pogosov, Emilia Silvas, Decebal Constantin Mocanu, Jos Elfring, René van de Molengraft |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2021 | Selfish Sparse RNN TrainingabstractSparse neural networks have been widely applied to reduce the computational demands of training and deploying over-parameterized deep neural networks. For inference acceleration, methods that discover a sparse network from a pre-trained dense network (dense-to-sparse training) work effectively. Recently, dynamic sparse training (DST) has been proposed to train sparse neural networks without pre-training a dense model (sparse-to-sparse training), so that the training process can also be accelerated. However, previous sparse-to-sparse methods mainly focus on Multilayer Perceptron Networks (MLPs) and Convolutional Neural Networks (CNNs), failing to match the performance of dense-to-sparse methods in the Recurrent Neural Networks (RNNs) setting. In this paper, we propose an approach to train intrinsically sparse RNNs with a fixed parameter count in one single run, without compromising performance. During training, we allow RNN layers to have a non-uniform redistribution across cell gates for better regularization. Further, we propose SNT-ASGD, a novel variant of the averaged stochastic gradient optimizer, which significantly improves the performance of all sparse training methods for RNNs. Using these strategies, we achieve state-of-the-art sparse training results, better than the dense-to-sparse methods, with various types of RNNs on Penn TreeBank and Wikitext-2 datasets. Our codes are available at https://github.com/Shiweiliuiiiiiii/Selfish-RNN. Shiwei Liu 0003, Decebal Constantin Mocanu, Yulong Pei, Mykola Pechenizkiy |
ICML | 2 |
| 2021 | Do We Actually Need Dense Over-Parameterization? In-Time Over-Parameterization in Sparse TrainingabstractIn this paper, we introduce a new perspective on training deep neural networks capable of state-of-the-art performance without the need for the expensive over-parameterization by proposing the concept of In-Time Over-Parameterization (ITOP) in sparse training. By starting from a random sparse network and continuously exploring sparse connectivities during training, we can perform an Over-Parameterization over the course of training, closing the gap in the expressibility between sparse training and dense training. We further use ITOP to understand the underlying mechanism of Dynamic Sparse Training (DST) and discover that the benefits of DST come from its ability to consider across time all possible parameters when searching for the optimal sparse connectivity. As long as sufficient parameters have been reliably explored, DST can outperform the dense neural network by a large margin. We present a series of experiments to support our conjecture and achieve the state-of-the-art sparse training performance with ResNet-50 on ImageNet. More impressively, ITOP achieves dominant performance over the overparameterization-based sparse methods at extreme sparsities. When trained with ResNet-34 on CIFAR-100, ITOP can match the performance of the dense model at an extreme sparsity 98%. Shiwei Liu 0003, Lu Yin 0006, Decebal Constantin Mocanu, Mykola Pechenizkiy |
ICML | 3 |
| 2021 | Sparse Training via Boosting Pruning Plasticity with NeuroregenerationabstractWorks on lottery ticket hypothesis (LTH) and single-shot network pruning (SNIP) have raised a lot of attention currently on post-training pruning (iterative magnitude pruning), and before-training pruning (pruning at initialization). The former method suffers from an extremely large computation cost and the latter usually struggles with insufficient performance. In comparison, during-training pruning, a class of pruning methods that simultaneously enjoys the training/inference efficiency and the comparable performance, temporarily, has been less explored. To better understand during-training pruning, we quantitatively study the effect of pruning throughout training from the perspective of pruning plasticity (the ability of the pruned networks to recover the original performance). Pruning plasticity can help explain several other empirical observations about neural network pruning in literature. We further find that pruning plasticity can be substantially improved by injecting a brain-inspired mechanism called neuroregeneration, i.e., to regenerate the same number of connections as pruned. We design a novel gradual magnitude pruning (GMP) method, named gradual pruning with zero-cost neuroregeneration (GraNet), that advances state of the art. Perhaps most impressively, its sparse-to-sparse version for the first time boosts the sparse-to-sparse training performance over various dense-to-sparse methods with ResNet-50 on ImageNet without extending the training time. We release all codes in https://github.com/Shiweiliuiiiiiii/GraNet. Shiwei Liu 0003, Tianlong Chen 0001, Xiaohan Chen 0001, Zahra Atashgahi, Lu Yin 0006, Huanyu Kou, Li Shen 0008, Mykola Pechenizkiy, Zhangyang Wang, Decebal Constantin Mocanu |
NeurIPS | 10 |
| 2021 | Evolving Plasticity for Autonomous Learning under Changing Environmental ConditionsabstractA fundamental aspect of learning in biological neural networks is the plasticity property which allows them to modify their configurations during their lifetime. Hebbian learning is a biologically plausible mechanism for modeling the plasticity property in artificial neural networks (ANNs), based on the local interactions of neurons. However, the emergence of a coherent global learning behavior from local Hebbian plasticity rules is not very well understood. The goal of this work is to discover interpretable local Hebbian learning rules that can provide autonomous global learning. To achieve this, we use a discrete representation to encode the learning rules in a finite search space. These rules are then used to perform synaptic changes, based on the local interactions of the neurons. We employ genetic algorithms to optimize these rules to allow learning on two separate tasks (a foraging and a prey-predator scenario) in online lifetime learning settings. The resulting evolved rules converged into a set of well-defined interpretable types, that are thoroughly discussed. Notably, the performance of these rules, while adapting the ANNs during the learning tasks, is comparable to that of offline learning methods such as hill climbing. Anil Yaman, Giovanni Iacca, Decebal Constantin Mocanu, Matt Coler, George Fletcher 0001, Mykola Pechenizkiy |
Evol. Comput. | 3 |
| 2021 | SpaceNet: Make Free Space for Continual LearningabstractThe continual learning (CL) paradigm aims to enable neural networks to learn tasks continually in a sequential fashion. The fundamental challenge in this learning paradigm is catastrophic forgetting previously learned tasks when the model is optimized for a new task, especially when their data is not accessible. Current architectural-based methods aim at alleviating the catastrophic forgetting problem but at the expense of expanding the capacity of the model. Regularization-based methods maintain a fixed model capacity; however, previous studies showed the huge performance degradation of these methods when the task identity is not available during inference (e.g. class incremental learning scenario). In this work, we propose a novel architectural-based method referred as SpaceNet1 for class incremental learning scenario where we utilize the available fixed capacity of the model intelligently. SpaceNet trains sparse deep neural networks from scratch in an adaptive way that compresses the sparse connections of each task in a compact number of neurons. The adaptive training of the sparse connections results in sparse representations that reduce the interference between the tasks. Experimental results show the robustness of our proposed method against catastrophic forgetting old tasks and the efficiency of SpaceNet in utilizing the available capacity of the model, leaving space for more tasks to be learned. In particular, when SpaceNet is tested on the well-known benchmarks for CL: split MNIST, split Fashion-MNIST, CIFAR-10/100, and iCIFAR100, it outperforms regularization-based methods by a big performance gap. Moreover, it achieves better performance than architectural-based methods without model expansion and achieves comparable results with rehearsal-based methods, while offering a huge memory reduction. Ghada Sokar, Decebal Constantin Mocanu, Mykola Pechenizkiy |
Neurocomputing | 2 |
| 2021 | Sparse evolutionary deep learning with over one million artificial neurons on commodity hardwareabstractAbstract Artificial neural networks (ANNs) have emerged as hot topics in the research community. Despite the success of ANNs, it is challenging to train and deploy modern ANNs on commodity hardware due to the ever-increasing model size and the unprecedented growth in the data volumes. Particularly for microarray data, the very high dimensionality and the small number of samples make it difficult for machine learning techniques to handle. Furthermore, specialized hardware such as graphics processing unit (GPU) is expensive. Sparse neural networks are the leading approaches to address these challenges. However, off-the-shelf sparsity-inducing techniques either operate from a pretrained model or enforce the sparse structure via binary masks. The training efficiency of sparse neural networks cannot be obtained practically. In this paper, we introduce a technique allowing us to train truly sparse neural networks with fixed parameter count throughout training. Our experimental results demonstrate that our method can be applied directly to handle high-dimensional data, while achieving higher accuracy than the traditional two-phase approaches. Moreover, we have been able to create truly sparse multilayer perceptron models with over one million neurons and to train them on a typical laptop without GPU ( https://github.com/dcmocanu/sparse-evolutionary-artificial-neural-networks/tree/master/SET-MLP-Sparse-Python-Data-Structures ), this being way beyond what is possible with any state-of-the-art technique. Shiwei Liu 0003, Decebal Constantin Mocanu, Amarsagar Reddy Ramapuram Matavalam, Yulong Pei, Mykola Pechenizkiy |
Neural Comput. Appl. | 2 |
| 2021 | Efficient and effective training of sparse recurrent neural networksabstractAbstract Recurrent neural networks (RNNs) have achieved state-of-the-art performances on various applications. However, RNNs are prone to be memory-bandwidth limited in practical applications and need both long periods of training and inference time. The aforementioned problems are at odds with training and deploying RNNs on resource-limited devices where the memory and floating-point operations (FLOPs) budget are strictly constrained. To address this problem, conventional model compression techniques usually focus on reducing inference costs, operating on a costly pre-trained model. Recently, dynamic sparse training has been proposed to accelerate the training process by directly training sparse neural networks from scratch. However, previous sparse training techniques are mainly designed for convolutional neural networks and multi-layer perceptron. In this paper, we introduce a method to train intrinsically sparse RNN models with a fixed number of parameters and floating-point operations (FLOPs) during training. We demonstrate state-of-the-art sparse performance with long short-term memory and recurrent highway networks on widely used tasks, language modeling, and text classification. We simply use the results to advocate that, contrary to the general belief that training a sparse neural network from scratch leads to worse performance than dense networks, sparse training with adaptive connectivity can usually achieve better performance than dense models for RNNs. Shiwei Liu 0003, Iftitahu Ni'mah, Vlado Menkovski, Decebal Constantin Mocanu, Mykola Pechenizkiy |
Neural Comput. Appl. | 4 |
| 2020 | Topological Insights into Sparse Neural Networks
Shiwei Liu 0003, Tim van der Lee, Anil Yaman, Zahra Atashgahi, Davide Ferraro, Ghada Sokar, Mykola Pechenizkiy, Decebal Constantin Mocanu |
ECML/PKDD (3) | 8 |
| 2020 | A Hybrid Framework Combining Vehicle System Knowledge with Machine Learning Methods for Improved Highway Trajectory PredictionabstractVehicle-to-vehicle communication is a solution to improve the quality of on-road traveling in terms of throughput, safety, efficiency and comfort. However, road users that do not communicate their planned activities can create dangerous situations, so prediction models are needed to foresee and anticipate their motions in the drivable space. Various prediction methods exist, either physics-based, data-based or hybrids, but they all make conservative assumptions about others' intentions, or they are developed using unrealistic data, and it is unclear how they perform for trajectory prediction. In this work, we introduce and demonstrate an optimal hybrid framework that overcomes these limitations, by combining the predictions of several physics-based and data-based models. Using on-road measured data we show that this novel framework outperforms the individual models in both longitudinal and lateral position predictions. We also discuss the required prediction boundaries from a safety perspective and estimate the accuracies of the models in relation to automated vehicle functions. The results achieved by this method will enable increased safety, comfort and even more proactive reactions of the automated vehicles. Manuel Muñoz Sánchez, Emilia Silvas, Denis Pogosov, Decebal Constantin Mocanu |
SMC | 4 |
| 2019 | Learning with delayed synaptic plasticityabstractThe plasticity property of biological neural networks allows them to perform learning and optimize their behavior by changing their configuration. Inspired by biology, plasticity can be modeled in artificial neural networks by using Hebbian learning rules, i.e. rules that update synapses based on the neuron activations and reinforcement signals. However, the distal reward problem arises when the reinforcement signals are not available immediately after each network output to associate the neuron activations that contributed to receiving the reinforcement signal. In this work, we extend Hebbian plasticity rules to allow learning in distal reward cases. We propose the use of neuron activation traces (NATs) to provide additional data storage in each synapse to keep track of the activation of the neurons. Delayed reinforcement signals are provided after each episode relative to the networks' performance during the previous episode. We employ genetic algorithms to evolve delayed synaptic plasticity (DSP) rules and perform synaptic updates based on NATs and delayed reinforcement signals. We compare DSP with an analogous hill climbing algorithm that does not incorporate domain knowledge introduced with the NATs, and show that the synaptic updates performed by the DSP rules demonstrate more effective training performance relative to the HC algorithm. Anil Yaman, Giovanni Iacca, Decebal Constantin Mocanu, George Fletcher 0001, Mykola Pechenizkiy |
GECCO | 3 |
| 2018 | Limited evaluation cooperative co-evolutionary differential evolution for large-scale neuroevolutionabstractMany real-world control and classification tasks involve a large number of features. When artificial neural networks (ANNs) are used for modeling these tasks, the network architectures tend to be large. Neuroevolution is an effective approach for optimizing ANNs; however, there are two bottlenecks that make their application challenging in case of high-dimensional networks using direct encoding. First, classic evolutionary algorithms tend not to scale well for searching large parameter spaces; second, the network evaluation over a large number of training instances is in general time-consuming. In this work, we propose an approach called the Limited Evaluation Cooperative Co-evolutionary Differential Evolution algorithm (LECCDE) to optimize high-dimensional ANNs. Anil Yaman, Decebal Constantin Mocanu, Giovanni Iacca, George Fletcher 0001, Mykola Pechenizkiy |
GECCO | 2 |
| 2017 | Unsupervised deep learning for real-time assessment of video streaming servicesabstractEvaluating quality of experience in video streaming services requires a quality metric that works in real time and for a broad range of video types and network conditions. This means that, subjective video quality assessment studies, or complex objective video quality assessment metrics, which would be best suited from the accuracy perspective, cannot be used for this tasks (due to their high requirements in terms of time and complexity, in addition to their lack of scalability). In this paper we propose a light-weight No Reference (NR) method that, by means of unsupervised machine learning techniques and measurements on the client side is able to assess quality in real-time, accurately and in an adaptable and scalable manner. Our method makes use of the excellent density estimation capabilities of the unsupervised deep learning techniques, the restricted Boltzmann machines, and light-weight video features computed just on the impaired video to provide a delta of quality degradation. We have tested our approach in two network impaired video sets, the LIMP and the ReTRiEVED video quality databases, benchmarking the results of our method against the well-known full reference metric VQM. We have obtained levels of accuracy of at least 85% in both datasets using all possible cases. Maria Torres Vega, Decebal Constantin Mocanu, Antonio Liotta |
Multim. Tools Appl. | 2 |
| 2017 | Estimating 3D trajectories from 2D projections via disjunctive factored four-way conditional restricted Boltzmann machines
Decebal Constantin Mocanu, Haitham Bou-Ammar, Luis Puig, Eric Eaton, Antonio Liotta |
Pattern Recognit. | 1 |
| 2017 | Predictive no-reference assessment of video quality
Maria Torres Vega, Decebal Constantin Mocanu, Stavros Stavrou, Antonio Liotta |
Signal Process. Image Commun. | 2 |
| 2017 | Deep Learning for Quality Assessment in Live Video StreamingabstractVideo content providers put stringent requirements on the quality assessment methods realized on their services. They need to be accurate, real-time, adaptable to new content, and scalable as the video set grows. In this letter, we introduce a novel automated and computationally efficient video assessment method. It enables accurate real-time (online) analysis of delivered quality in an adaptable and scalable manner. Offline deep unsupervised learning processes are employed at the server side and inexpensive no-reference measurements at the client side. This provides both real-time assessment and performance comparable to the full reference counterpart, while maintaining its no-reference characteristics. We tested our approach on the LIMP Video Quality Database (an extensive packet loss impaired video set) obtaining a correlation between 78% and 91% to the FR benchmark (the video quality metric). Due to its unsupervised learning essence, our method is flexible and dynamically adaptable to new content and scalable with the number of videos. Maria Torres Vega, Decebal Constantin Mocanu, Jeroen Famaey, Stavros Stavrou, Antonio Liotta |
IEEE Signal Process. Lett. | 2 |
| 2016 | On the Synergy of Network Science and Artificial Intelligence
Decebal Constantin Mocanu |
IJCAI | 1 |
| 2016 | A Regression Method for real-time video quality evaluation
Maria Torres Vega, Decebal Constantin Mocanu, Antonio Liotta |
MoMM | 2 |
| 2016 | Big IoT data mining for real-time energy disaggregation in buildingsabstractIn the smart grid context, the identification and prediction of building energy flexibility is a challenging open question, thus paving the way for new optimized behaviors from the demand side. At the same time, the latest smart meters developments allow us to monitor in real-time the power consumption level of the home appliances, aiming at a very accurate energy disaggregation. However, due to practical constraints is infeasible in the near future to attach smart meter devices on all home appliances, which is the problem addressed herein. We propose a hybrid approach, which combines sparse smart meters with machine learning methods. Using a subset of buildings equipped with subset of smart meters we can create a database on which we train two deep learning models, i.e. Factored Four-Way Conditional Restricted Boltzmann Machines (FFW-CRBMs) and Disjunctive FFW-CRBM. We show how our method may be used to accurately predict and identify the energy flexibility of buildings unequipped with smart meters, starting from their aggregated energy values. The proposed approach was validated on a real database, namely the Reference Energy Disaggregation Dataset. The results show that for the flexibility prediction problem solved here, Disjunctive FFW-CRBM outperforms the FFW-CRBMs approach, where for classification task their capabilities are comparable. Decebal Constantin Mocanu, Elena Mocanu, Phuong H. Nguyen, Madeleine Gibescu, Antonio Liotta |
SMC | 1 |
| 2016 | A topological insight into restricted Boltzmann machinesabstractRestricted Boltzmann Machines (RBMs) and models derived from them have been successfully used as basic building blocks in deep artificial neural networks for automatic features extraction, unsupervised weights initialization, but also as density estimators. Thus, their generative and discriminative capabilities, but also their computational time are instrumental to a wide range of applications. Our main contribution is to look at RBMs from a topological perspective, bringing insights from network science. Firstly, here we show that RBMs and Gaussian RBMs (GRBMs) are bipartite graphs which naturally have a small-world topology. Secondly, we demonstrate both on synthetic and real-world datasets that by constraining RBMs and GRBMs to a scale-free topology (while still considering local neighborhoods and data distribution), we reduce the number of weights that need to be computed by a few orders of magnitude, at virtually no loss in generative performance. Thirdly, we show that, for a fixed number of weights, our proposed sparse models (which by design have a higher number of hidden neurons) achieve better generative capabilities than standard fully connected RBMs and GRBMs (which by design have a smaller number of hidden neurons), at no additional computational costs. Decebal Constantin Mocanu, Elena Mocanu, Phuong H. Nguyen, Madeleine Gibescu, Antonio Liotta |
Mach. Learn. | 1 |
| 2015 | Reduced reference image quality assessment via Boltzmann MachinesabstractMonitoring and controlling the user's perceived quality, in modern video services is a challenging proposition, mainly due to the limitations of current Image Quality Assessment (IQA) algorithms. Subjective Quality of Experience (QoE) is widely used to get a right impression, but unfortunately this can not be used in real world scenarios. In general, objective QoE algorithms represent a good substitution for the subjective ones, and they are split in three main directions: Full Reference (FR), Reduced Reference (RR), and No Reference (NR). From these three, the RR IQA approach offers a practical solution to assess the quality of an impaired image due to the fact that just a small amount of information is needed from the original image. At the same time, keeping in mind that we need automated QoE algorithms which are context independent, in this paper we introduce a novel stochastic RR IQA metric to assess the quality of an image based on Deep Learning, namely Restricted Boltzmann Machine Similarity Measure (RBMSim). RBMSim was evaluated on two benchmarked image databases with subjective studies, against objective IQA algorithms. The results show that its performance is comparable, or even better in some cases, with widely known FR IQA methods. Decebal Constantin Mocanu, Georgios Exarchakos, Haitham Bou-Ammar, Antonio Liotta |
IM | 1 |
| 2015 | Cognitive streaming on android devicesabstractAs the number of mobile devices increases, so do the complexity of wireless networks and the user's requirements. This tendency makes necessary for Multimedia Services to take the needed actions to adapt to the upcoming technology. A prominent example of this type of services is HTTP Adaptive Video Streaming Applications. In this research, we have studied how the latest HTTP Adaptive Streaming techniques, mainly developed for standard computers, could be adapted and used in mobile wireless devices. Furthermore, inspired by these solutions, which usually make use of Reinforcement Learning (RL) algorithms to find the suitable streaming rate, we have conceived a novel smart video player client in Java for Android platform using the Dynamic Adaptive Streaming over HTTP (DASH) protocol. We have assessed the performance of our proposed solution in a self-developed wireless test-bed under different network conditions. Thus, we have seen that by including in the reward function contributions regarding the download speed of the video segments, especially needed due to the fluctuating nature of the wireless networks, and the segments already buffered, improves drastically the overall performance of the video client. Besides that, we have discovered that, in a cognitive adaptive approach, bandwidth constraints affect the user's experience more substantially, while impairments such as packet loss can be prevented. Maria Torres Vega, Decebal Constantin Mocanu, Rosario Barresi, Giancarlo Fortino, Antonio Liotta |
IM | 2 |
| 2015 | Accuracy of No-Reference Quality Metrics in Network-impaired Video StreamsabstractThe Video Quality Metric (VQM) is nowadays one of the most used objective methods to assess video quality, thanks to its high correlation with both the human visual system (HVS) and subjective methods. VQM is, however, not viable in real-time deployments such as mobile streaming, not only due to its high computational demands but, specifically, because it is a Full-Reference (FR) metric, which requires as input both the original video and its impaired counterpart. On the other hand, No-Reference (NR) objective algorithms operate directly on the impaired video and are considerably faster, but loose out when it comes to accuracy. In this research, we assess a range of NR metrics, alongside a lightweight FR metric, using VQM as benchmark. Our study covers a range of methods, a diverse set of video types and encoding conditions, and a range of network impairment test-cases. We show the extent by which packet loss affects different video types, correlating the accuracy of NR metrics to the FR benchmark. Our study helps identifying the conditions under which simple metrics may be used effectively and indicates an avenue to control the quality of streaming systems in line with human perception. Maria Torres Vega, Vittorio Sguazzo, Decebal Constantin Mocanu, Antonio Liotta |
MoMM | 3 |
| 2015 | Factored four way conditional restricted Boltzmann machines for activity recognition
Decebal Constantin Mocanu, Haitham Bou-Ammar, Dietwig Lowet, Kurt Driessens, Antonio Liotta, Gerhard Weiss 0001, Karl Tuyls |
Pattern Recognit. Lett. | 1 |
| 2014 | Deep learning for objective quality assessment of 3D imagesabstractImproving the users' Quality of Experience (QoE) in modern 3D Multimedia Systems is a challenging proposition, mainly due to our limited knowledge of 3D image Quality Assessment algorithms. While subjective QoE methods would better reflect the nature of human perception, these are not suitable in real-time automation cases. In this paper we tackle this issue from a new angle, using deep learning to make predictions on the user's QoE rather than trying to measure it through deterministic algorithms. We benchmark our method, dubbed Quality of Experience for 3D images through Factored Third Order Restricted Boltzmann Machine (Q3D-RBM), with subjective QoE methods, to determine its accuracy for different types of 3D images. The outcome is a Reduced Reference QoE assessment process for automatic image assessment and has significant potential to be extended to work on 3D video assessment. Decebal Constantin Mocanu, Georgios Exarchakos, Antonio Liotta |
ICIP | 1 |
| 2014 | When does lower bitrate give higher quality in modern video services?abstractDue to the difficulties on approximating the human perception with algorithms, increasing the users Quality of Experience (QoE) in modern video services is a challenging task. But more than that, prior to estimating QoE, it is important to know how different types of network impairments actually affect the video quality. This paper takes a closer look at the relation between the network quality of service (QoS) and the video QoE degradation. Using a sophisticated network emulation environment, we benchmark a range of video types and video quality levels under controlled network conditions. Our analysis shows that, along with a number of expected situations come also some counterintuitive QoS-to-QoE conditions. We discuss ways in which a better understanding of the mutual influence between networks and video streams could lead to more efficient utilization of the Internet. Decebal Constantin Mocanu, Antonio Liotta, Arianna Ricci, Maria Torres Vega, Georgios Exarchakos |
NOMS | 1 |
| 2014 | Node centrality awareness via swarming effectsabstractCentralization is a weakness in large scale dynamic topologies and, thus, collaboratively electing at runtime the most impactful (central) nodes is necessary to ensure reliability. However, little has been achieved in measuring the centrality of nodes in an accurate, fast, decentralized and with low overhead method. This paper proposes a swarm-inspired approach (DANIS) to detect the nodes that would most impact the network connectivity if removed. The idea lies on the trivial fact that the more accessible a node is, the more resources per time unit it loses. Experiments on random, scale-free and small-world graph topologies indicate that DANIS achieves higher accuracy, faster convergence and fewer communication overhead compared to other methods. Decebal Constantin Mocanu, Georgios Exarchakos, Antonio Liotta |
SMC | 1 |
| 2014 | Inexpensive user tracking using Boltzmann MachinesabstractInexpensive user tracking is an important problem in various application domains such as healthcare, human-computer interaction, energy savings, safety, robotics, security and so on. Yet, it cannot be easily solved due to its probabilistic nature, high level of abstraction and uncertainties, on the one side, and to the limitations of our current technologies and learning algorithms, on the other side. In this paper, we tackle this problem by using the Multi-integrated Sensor Technology, which comes at a low price. At the same time, we are aiming to address the lightweight learning requirements by investigating Factored Conditional Restricted Boltzmann Machines (FCRBMs), a form of Deep Learning, that has proven to be an efficient and effective machine learning framework. However, due to their construction properties, the conventional FCRBMs are only capable of performing predictions but are not capable of making classification. Herein, we are proposing extended FCRBMs (eFCRBMs), which incorporate a novel classification scheme, to solve this problem. Experiments performed on both artificially generated as well as real-world data demonstrate the effectiveness and efficiency of the proposed technique. We show that eFCRBMs outperform popular approaches including Support Vector Machines, Naive Bayes, AdaBoost, and Gaussian Mixture Models. Elena Mocanu, Decebal Constantin Mocanu, Haitham Bou-Ammar, Zoran Zivkovic, Antonio Liotta, Evgueni N. Smirnov |
SMC | 2 |
| 2013 | Predicting Battery Depletion of Neighboring Wireless Sensor Nodes
Roshan Kotian, Georgios Exarchakos, Decebal Constantin Mocanu, Antonio Liotta |
ICA3PP (2) | 3 |
| 2013 | Instantaneous Video Quality Assessment for lightweight devicesabstractMonitoring and controlling the user's Quality of Experience (QoE) in modern video services is a challenging proposition, mainly due to the limitations of current video quality assessment algorithms. While subjective QoE methods would better reflect the nature of human perception, these are not suitable in real-time automation cases. On the other hand, the existing objective algorithms are either too complex or too inaccurate, particularly in the context of lightweight devices such as camera sensors or smart phones. This paper introduces a novel objective QoE algorithm, Instantaneous Video Quality Assessment (IVQA), that is comparably as accurate as the most heavyweight algorithm available in the literature but can also be run in real-time. This approach is tested against a selection of ten objective metrics and benchmarked with a subjective user dataset. Antonio Liotta, Decebal Constantin Mocanu, Vlado Menkovski, Luciana Cagnetta, Georgios Exarchakos |
MoMM | 2 |
| 2013 | Automatically Mapped Transfer between Reinforcement Learning Tasks via Three-Way Restricted Boltzmann Machines
Haitham Bou-Ammar, Decebal Constantin Mocanu, Matthew E. Taylor, Kurt Driessens, Karl Tuyls, Gerhard Weiss 0001 |
ECML/PKDD (2) | 2 |