Aurelio Uncini

dblp:34/3010 · DBLP profile ↗
← Back
103ranked-venue papers
4as first author
21since 2021 · last 2027
0000-0002-5793-0917ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 53 · 3 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 38 · 1 first-author · 6 since 2021Systems, architecture and hardware · 8Databases, data management, data science and information retrieval · 3Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2027 Physics-informed adaptive filtering for acoustic echo cancellation
abstract
This paper introduces a physics-informed adaptive filtering framework for acoustic echo cancellation (AEC). Unlike conventional adaptive algorithms that rely solely on data-driven error minimization, the proposed method incorporates physically motivated priors derived from acoustic wave propagation and room impulse response structure. The echo path estimation problem is formulated as a composite stochastic optimization task, where the instantaneous squared error is regularized by constraints encoding causality, exponential energy decay, time-weighted sparsity of early reflections, spectral smoothness, and slow temporal variation of the acoustic path. The resulting Physics-Informed Normalized Least-Mean-Squares (PI-NLMS) algorithm performs stochastic gradient descent on the regularized cost while enforcing hard causality through projection. The proposed formulation restricts adaptation to a physically plausible echo-path manifold, improving conditioning and reducing variance without substantially increasing computational complexity. Theoretical analysis establishes mean convergence conditions and characterizes the bias-variance trade-off introduced by structured regularization. Simulation results under stationary and time-varying echo paths demonstrate faster convergence, improved steady-state misalignment, and enhanced echo return loss enhancement (ERLE) compared to conventional NLMS and sparsity-aware baselines.
Michele Scarpiniti, Danilo Comminiello, Aurelio Uncini
Signal Process.3
2025 FoleyGRAM: Video-to-Audio Generation with GRAM-Aligned Multimodal Encoders
abstract
In this work, we present FoleyGRAM, a novel approach to video-to-audio generation that emphasizes semantic conditioning through the use of aligned multimodal encoders. Building on prior advancements in video-to-audio generation, FoleyGRAM leverages the Gramian Representation Alignment Measure (GRAM) to align embeddings across video, text, and audio modalities, enabling precise semantic control over the audio generation process. The core of FoleyGRAM is a diffusion-based audio synthesis model conditioned on GRAM-aligned embeddings and waveform envelopes, ensuring both semantic richness and temporal alignment with the corresponding input video. We evaluate FoleyGRAM on the Greatest Hits dataset, a standard benchmark for video-to-audio models. Our experiments demonstrate that aligning multimodal encoders using GRAM enhances the system’s ability to semantically align generated audio with video content, advancing the state of the art in video-to-audio synthesis.
Riccardo F. Gramaccioni, Christian Marinoni, Eleonora Grassucci, Giordano Cicchetti, Aurelio Uncini, Danilo Comminiello
IJCNN5
2025 EfficientAudioNet: Enhancing Environmental Sound Classification through Data Fusion of Multiple Audio Representations
abstract
Environmental Sound Classification (ESC) is becoming an ever increasingly important application in different scenarios, such as smart cities, autonomous systems, safety, and industrial monitoring. Traditional methods for ESC mainly rely on features extracted from a single-representation, usually spectrograms or MFCCs. However, while deep learning-based CNN models have demonstrated excellent performance, they still suffer from certain limitations due to the reliance on a single feature representation. In this regard, this work exploits a multi-representation strategy by fusing five kinds of audio features, namely: spectrograms, phasograms, scalograms, wavelet phasograms, and MFCC-grams. Each representation captures different properties of the audio. These representations are combined in a structured manner by investigating three fusion strategies: early, intermediate, and late fusion using a novel model based on the EfficientNet, named EfficientAudioNet. The proposed strategies are evaluated on four benchmark datasets: a Construction Site machinery sounds dataset, the ESC-10 and ESC-50 environmental sound datasets, and the UrbanSound8K dataset. Experimental results demonstrate that the multi-representation fusion, specially the early fusion, significantly enhances the classification performance. Overall, the proposed approach overcomes state-of-the-art accuracy on all the tested datasets.
Michele Scarpiniti, Saud Hussain, Wangyi Pu, Aurelio Uncini, Yong-Cheol Lee
IJCNN4
2025 Quaternion Wavelet-Conditioned Diffusion Models for Image Super-Resolution
abstract
Image Super-Resolution is a fundamental problem in computer vision with broad applications spacing from medical imaging to satellite analysis. The ability to reconstruct high-resolution images from low-resolution inputs is crucial for enhancing downstream tasks such as object detection and segmentation. While deep learning has significantly advanced SR, achieving high-quality reconstructions with fine-grained details and realistic textures remains challenging, particularly at high upscaling factors. Recent approaches leveraging diffusion models have demonstrated promising results, yet they often struggle to balance perceptual quality with structural fidelity. In this work, we introduce ResQu a novel SR framework that integrates a quaternion wavelet preprocessing framework with latent diffusion models, incorporating a new quaternion wavelet- and time-aware encoder. Unlike prior methods that simply apply wavelet transforms within diffusion models, our approach enhances the conditioning process by exploiting quaternion wavelet embeddings, which are dynamically integrated at different stages of denoising. Furthermore, we also leverage the generative priors of foundation models such as Stable Diffusion. Extensive experiments on domain-specific datasets demonstrate that our method achieves outstanding SR results, outperforming in many cases existing approaches in perceptual quality and standard evaluation metrics. The code will be available after the revision process.
Luigi Sigillo, Christian Bianchi, Aurelio Uncini, Danilo Comminiello
IJCNN3
2025 Generalizing medical image representations via quaternion wavelet networks
abstract
Neural network generalizability is becoming a broad research field due to the increasing availability of datasets from different sources and for various tasks. This issue is even wider when processing medical data, where a lack of methodological standards causes large variations being provided by different imaging centers or acquired with various devices and cofactors. To overcome these limitations, we introduce a novel, generalizable, data- and task-agnostic framework able to extract salient features from medical images. The proposed quaternion wavelet network (QUAVE) can be easily integrated with any pre-existing medical image analysis or synthesis task, and it can be involved with real, quaternion, or hypercomplex-valued models, generalizing their adoption to single-channel data. QUAVE first extracts different sub-bands through the quaternion wavelet transform, resulting in both low-frequency/approximation bands and high-frequency/fine-grained features. Then, it weighs the most representative set of sub-bands to be involved as input to any other neural model for image processing, replacing standard data samples. We conduct an extensive experimental evaluation comprising different datasets, diverse image analysis, and synthesis tasks including reconstruction, segmentation, and modality translation. We also evaluate QUAVE in combination with both real and quaternion-valued models. Results demonstrate the effectiveness and the generalizability of the proposed framework that improves network performance while being flexible to be adopted in manifold scenarios and robust to domain shifts. The full code is available at: https://github.com/ispamm/QWT.
Luigi Sigillo, Eleonora Grassucci, Aurelio Uncini, Danilo Comminiello
Neurocomputing3
2024 Efficient Functional Link Adaptive Filters Based On Nearest Kronecker Product Decomposition
abstract
Functional link adaptive filters (FLAFs) utilize expansion blocks to nonlinearly augment the input signal to a higher dimensional space, after which an adaptive weight algorithm is applied. These filters are useful for nonlinear system identification tasks, as they can update a large number of coefficients to effectively model the nonlinear system, even when the degree of nonlinearity is not comprehended in advance. However, in many cases, not all of the weights in the nonlinear and linear filter will significantly contribute to the identified model. This paper introduces a novel class of FLAFs based on the nearest Kronecker product (NKP) decomposition. Utilizing the inherent low-rank nature of weight vectors in many scenarios, our approach aims to improve convergence performance and tracking capabilities compared to traditional FLAFs. Additionally, we address noise mitigation challenges, particularly in nonlinear acoustic echo cancellation scenarios. By incorporating NKP decomposition, our proposed FLAFs offer promising solutions for enhancing adaptability and performance in nonlinear system identification, making them valuable tools in practical applications.
Alireza Nezamdoust, Mario Huemer, Aurelio Uncini, Danilo Comminiello
ICASSP3
2024 A Meta-Learning Approach for Training Explainable Graph Neural Networks
abstract
In this article, we investigate the degree of explainability of graph neural networks (GNNs). The existing explainers work by finding global/local subgraphs to explain a prediction, but they are applied after a GNN has already been trained. Here, we propose a meta-explainer for improving the level of explainability of a GNN directly at training time, by steering the optimization procedure toward minima that allow post hoc explainers to achieve better results, without sacrificing the overall accuracy of GNN. Our framework (called MATE, MetA-Train to Explain) jointly trains a model to solve the original task, e.g., node classification, and to provide easily processable outputs for downstream algorithms that explain the model's decisions in a human-friendly way. In particular, we meta-train the model's parameters to quickly minimize the error of an instance-level GNNExplainer trained on-the-fly on randomly sampled nodes. The final internal representation relies on a set of features that can be "better" understood by an explanation algorithm, e.g., another instance of GNNExplainer. Our model-agnostic approach can improve the explanations produced for different GNN architectures and use any instance-based explainer to drive this process. Experiments on synthetic and real-world datasets for node and graph classification show that we can produce models that are consistently easier to explain by different algorithms. Furthermore, this increase in explainability comes at no cost to the accuracy of the model.
Indro Spinelli, Simone Scardapane, Aurelio Uncini
IEEE Trans. Neural Networks Learn. Syst.3
2023 Overview of the L3DAS23 Challenge on Audio-Visual Extended Reality
abstract
The primary goal of the L3DAS23 Signal Processing Grand Challenge at ICASSP 2023 is to promote and support collaborative research on machine learning for 3D audio signal processing, with a specific emphasis on 3D speech enhancement and 3D Sound Event Localization and Detection in Extended Reality applications. As part of our latest competition, we provide a brand-new dataset, which maintains the same general characteristics of the L3DAS21 and L3DAS22 datasets, but with first-order Ambisonics recordings from multiple reverberant simulated environments. Moreover, we start exploring an audio-visual scenario by providing images of these environments, as perceived by the different microphone positions and orientations. We also propose updated baseline models for both tasks that can now support audio-image couples as input and a supporting API to replicate our results. Finally, we present the results of the participants. Further details about the challenge are available at www.l3das.com/icassp2023.
Christian Marinoni, Riccardo F. Gramaccioni, Changan Chen, Aurelio Uncini, Danilo Comminiello
ICASSP4
2023 Continual learning with invertible generative models
Jary Pomponi, Simone Scardapane, Aurelio Uncini
Neural Networks3
2023 Dual quaternion ambisonics array for six-degree-of-freedom acoustic representation
Eleonora Grassucci, Gioia Mancini, Christian Brignone, Aurelio Uncini, Danilo Comminiello
Pattern Recognit. Lett.4
2023 GROUSE: A Task and Model Agnostic Wavelet- Driven Framework for Medical Imaging
abstract
In recent years, deep learning has permeated the field of medical image analysis gaining increasing attention from clinicians. However, medical images always require specific preprocessing that often includes downscaling due to computational constraints. This may cause a crucial loss of information magnified by the fact that the region of interest is usually a tiny portion of the image. To overcome these limitations, we propose GROUSE, a novel and generalizable framework that produces salient features from medical images bygroupingandselectingfrequency sub-bands that provide approximations and fine-grained details useful for building a more complete input representation. The framework provides the most enlightening set of bands by learning their statistical dependency to avoid redundancy and by scoring their informativeness to provide meaningful data. This set of representative features can be fed as input to any neural model, replacing the conventional image input. Our method is task- and model-agnostic, thus it can be generalized to any medical image benchmark, as we extensively demonstrate with different tasks, datasets, and model domains. We show that the proposed framework enhances model performance in every test we conduct without requiring ad-hoc preprocessing or network adjustments.
Eleonora Grassucci, Luigi Sigillo, Aurelio Uncini, Danilo Comminiello
IEEE Signal Process. Lett.3
2023 A New Class of Efficient Adaptive Filters for Online Nonlinear Modeling
abstract
Nonlinear models are known to provide excellent performance in real-world applications that often operate in nonideal conditions. However, such applications often require online processing to be performed with limited computational resources. To address this problem, we propose a new class of efficient nonlinear models for online applications. The proposed algorithms are based on linear-in-the-parameters (LIPs) nonlinear filters using functional link expansions. In order to make this class of functional link adaptive filters (FLAFs) efficient, we propose low-complexity expansions and frequency-domain adaptation of the parameters. Among this family of algorithms, we also define the partitioned-block frequency-domain FLAF (FD-FLAF), whose implementation is particularly suitable for online nonlinear modeling problems. We assess and compare FD-FLAFs with different expansions providing the best possible tradeoff between performance and computational complexity. Experimental results prove that the proposed algorithms can be considered as an efficient and effective solution for online applications, such as the acoustic echo cancellation, even in the presence of adverse nonlinear conditions and with limited availability of computational resources.
Danilo Comminiello, Alireza Nezamdoust, Simone Scardapane, Michele Scarpiniti, Amir Hussain 0001, Aurelio Uncini
IEEE Trans. Syst. Man Cybern. Syst.6
2022 L3DAS22 Challenge: Learning 3D Audio Sources in a Real Office Environment
abstract
The L3DAS22 Challenge is aimed at encouraging the development of machine learning strategies for 3D speech enhancement and 3D sound localization and detection in office-like environments. This challenge improves and extends the tasks of the L3DAS21 edition1. We generated a new dataset, which maintains the same general characteristics of L3DAS21 datasets, but with an extended number of data points and adding constrains that improve the baseline model’s efficiency and overcome the major difficulties encountered by the participants of the previous challenge. We updated the baseline model of Task 1, using the architecture that ranked first in the previous challenge edition. We wrote a new supporting API, improving its clarity and ease-of-use. In the end, we present and discuss the results submitted by all participants. L3DAS22 Challenge website: www.l3das.com/icassp2022.
Eric Guizzo, Christian Marinoni, Marco Pennese, Xinlei Ren, Xiguang Zheng, Bruno S. Masiero, Aurelio Uncini, Danilo Comminiello
ICASSP8
2022 Hypercomplex Image- to- Image Translation
abstract
Image-to-image translation (I2I) aims at transferring the content representation from an input domain to an output one, bouncing along different target domains. Recent I2I generative models, which gain outstanding results in this task, comprise a set of diverse deep networks each with tens of million parameters. Moreover, images are usually three-dimensional being composed of RGB channels and common neural models do not take dimensions correlation into account, losing beneficial information. In this paper, we propose to leverage hypercomplex algebra properties to define lightweight I2I generative models capable of preserving pre-existing relations among image dimensions, thus exploiting additional input information. On manifold I2I benchmarks, we show how the proposed Quaternion StarGANv2 and parameterized hypercomplex StarGANv2 (PHStarGANv2) reduce parameters and storage memory amount while ensuring high domain translation performance and good image quality as measured by FID and LPIPS scores. Full code is available at https://github.com/ispamm/HI2I.
Eleonora Grassucci, Luigi Sigillo, Aurelio Uncini, Danilo Comminiello
IJCNN3
2022 Pixle: a fast and effective black-box attack based on rearranging pixels
abstract
Recent research has found that neural networks are vulnerable to several types of adversarial attacks, where the input samples are modified in such a way that the model produces a wrong prediction that misclassifies the adversarial sample. In this paper we focus on black-box adversarial attacks, that can be performed without knowing the inner structure of the attacked model, nor the training procedure, and we propose a novel attack that is capable of correctly attacking a high percentage of samples by rearranging a small number of pixels within the attacked image. We demonstrate that our attack works on a large number of datasets and models, that it requires a small number of iterations, and that the distance between the original sample and the adversarial one is negligible to the human eye.
Jary Pomponi, Simone Scardapane, Aurelio Uncini
IJCNN3
2022 CoVal-SGAN: A Complex-Valued Spectral GAN architecture for the effective audio data augmentation in construction sites
abstract
Generative audio data augmentation in a construction site is one of challenging research areas due to the high dissimilarity between work sounds of involved machines and equipment. However, it becomes necessary since the availability of audio data of critical work classes is often rare. Motivated by these considerations and demands, in this paper, we propose a complex-valued GAN architecture working with the audio spectrogram, named CoVal-SGAN, for an effective augmentation of audio data. Specifically, the proposed CoVal-SGAN exploits both the magnitude and phase information to improve the quality of the artificially generated audio signals and increase the overall performance of the underlying classifier. Numerical results, performed on the data recorded in real-world construction sites, along with the comparisons with available state-of-the-art approaches, show the effectiveness of the proposed idea by obtaining an improved accuracy.
Michele Scarpiniti, Cristiano Mauri, Danilo Comminiello, Aurelio Uncini, Yong-Cheol Lee
IJCNN4
2021 A Quaternion-Valued Variational Autoencoder
abstract
Deep probabilistic generative models have achieved incredible success in many fields of application. Among such models, variational autoencoders (VAEs) have proved their ability in modeling a generative process by learning a latent representation of the input. In this paper, we propose a novel VAE defined in the quaternion domain, which exploits the properties of quaternion algebra to improve performance while significantly reducing the number of parameters required by the network. The success of the proposed quaternion VAE with respect to traditional VAEs relies on the ability to leverage the internal relations between quaternion-valued input features and on the properties of second-order statistics which allow to define the latent variables in the augmented quaternion domain. In order to show the advantages due to such properties, we define a plain convolutional VAE in the quaternion domain and we evaluate its performance with respect to its real-valued counterpart on the CelebA face dataset.
Eleonora Grassucci, Danilo Comminiello, Aurelio Uncini
ICASSP3
2021 Deep Belief Network based audio classification for construction sites monitoring
Michele Scarpiniti, Francesco Colasante, Simone Di Tanna, Marco Ciancia, Yong-Cheol Lee, Aurelio Uncini
Expert Syst. Appl.6
2021 Bayesian Neural Networks with Maximum Mean Discrepancy regularization
Jary Pomponi, Simone Scardapane, Aurelio Uncini
Neurocomputing3
2021 Structured Ensembles: An approach to reduce the memory footprint of ensemble methods
abstract
In this paper, we propose a novel ensembling technique for deep neural networks, which is able to drastically reduce the required memory compared to alternative approaches. In particular, we propose to extract multiple sub-networks from a single, untrained neural network by solving an end-to-end optimization task combining differentiable scaling over the original architecture, with multiple regularization terms favouring the diversity of the ensemble. Since our proposal aims to detect and extract sub-structures, we call it Structured Ensemble. On a large experimental evaluation, we show that our method can achieve higher or comparable accuracy to competing methods while requiring significantly less storage. In addition, we evaluate our ensembles in terms of predictive calibration and uncertainty, showing they compare favourably with the state-of-the-art. Finally, we draw a link with the continual learning literature, and we propose a modification of our framework to handle continuous streams of tasks with a sub-linear memory cost. We compare with a number of alternative strategies to mitigate catastrophic forgetting, highlighting advantages in terms of average accuracy and memory.
Jary Pomponi, Simone Scardapane, Aurelio Uncini
Neural Networks3
2021 Adaptive Propagation Graph Convolutional Network
abstract
Graph convolutional networks (GCNs) are a family of neural network models that perform inference on graph data by interleaving vertexwise operations and message-passing exchanges across nodes. Concerning the latter, two key questions arise: 1) how to design a differentiable exchange protocol (e.g., a one-hop Laplacian smoothing in the original GCN) and 2) how to characterize the tradeoff in complexity with respect to the local updates. In this brief, we show that the state-of-the-art results can be achieved by adapting the number of communication steps independently at every node. In particular, we endow each node with a halting unit (inspired by Graves' adaptive computation time [1]) that after every exchange decides whether to continue communicating or not. We show that the proposed adaptive propagation GCN (AP-GCN) achieves superior or similar results to the best proposed models so far on a number of benchmarks while requiring a small overhead in terms of additional parameters. We also investigate a regularization term to enforce an explicit tradeoff between communication and accuracy. The code for the AP-GCN experiments is released as an open-source library.
Indro Spinelli, Simone Scardapane, Aurelio Uncini
IEEE Trans. Neural Networks Learn. Syst.3
2020 Differentiable Branching In Deep Networks for Fast Inference
abstract
In this paper, we consider the design of deep neural networks augmented with multiple auxiliary classifiers departing from the main (backbone) network. These classifiers can be used to perform early-exit from the network at various layers, making them convenient for energy-constrained applications such as IoT, embedded devices, or Fog computing. However, designing an optimized early-exit strategy is a difficult task, generally requiring a large amount of manual fine-tuning. In this paper, we propose a way to jointly optimize this strategy together with the branches, providing an end-to-end trainable algorithm for this emerging class of neural networks. We achieve this by replacing the original output of the branches with a 'soft', differentiable approximation. In addition, we also propose a regularization approach to trade-off the computational efficiency of the early-exit strategy with respect to the overall classification accuracy. We evaluate our proposed design approach on a set of image classification benchmarks, showing significant gains in accuracy and inference time.
Simone Scardapane, Danilo Comminiello, Michele Scarpiniti, Enzo Baccarelli, Aurelio Uncini
ICASSP5
2020 Efficient continual learning in neural networks with embedding regularization
Jary Pomponi, Simone Scardapane, Vincenzo Lomonaco, Aurelio Uncini
Neurocomputing4
2020 Optimized training and scalable implementation of Conditional Deep Neural Networks with early exits for Fog-supported IoT applications
Enzo Baccarelli, Simone Scardapane, Michele Scarpiniti, Alireza Momenzadeh, Aurelio Uncini
Inf. Sci.5
2020 Missing data imputation with adversarially-trained graph convolutional networks
Indro Spinelli, Simone Scardapane, Aurelio Uncini
Neural Networks3
2019 Quaternion Convolutional Neural Networks for Detection and Localization of 3D Sound Events
abstract
Learning from data in the quaternion domain enables us to exploit internal dependencies of 4D signals and treating them as a single entity. One of the models that perfectly suits with quaternion-valued data processing is represented by 3D acoustic signals in their spherical harmonics decomposition. In this paper, we address the problem of localizing and detecting sound events in the spatial sound field by using quaternion-valued data processing. In particular, we consider the spherical harmonic components of the signals captured by a first-order ambisonic microphone and process them by using a quaternion convolutional neural network. Experimental results show that the proposed approach exploits the correlated nature of the ambisonic signals, thus improving accuracy results in 3D sound event detection and localization.
Danilo Comminiello, Marco Lella, Simone Scardapane, Aurelio Uncini
ICASSP4
2019 Frequency-domain Adaptive Filtering: from Real to Hypercomplex Signal Processing
abstract
Frequency-domain adaptive filters (FDAFs) have been widely used over the years, but they are still matter of research due to their powerful capabilities that differentiate them from the whole family of time-domain adaptive filters. This paper aims at providing an overview on FDAFs through a unifying framework that can be used for the derivation of the most popular algorithms of the FDAF family and enables the processing of a wide variety of signals, from real-valued ones to complex- and hypercomplex-valued signals. In particular, we focus on a recent class of FDAFs in the quaternion domain and we show how to derive it from the described framework. Moreover, we evaluate the application of the derived quaternion FDAF to the processing of 3D audio signals. Experimental results show the effectiveness of the proposed adaptive filter in estimating the inverse of a multidimensional acoustic impulse response.
Danilo Comminiello, Michele Scarpiniti, Raffaele Parisi, Aurelio Uncini
ICASSP4
2019 Widely Linear Kernels for Complex-valued Kernel Activation Functions
abstract
Complex-valued neural networks (CVNNs) have been shown to be powerful nonlinear approximators when the input data can be properly modeled in the complex domain. One of the major challenges in scaling up CVNNs in practice is the design of complex activation functions. Recently, we proposed a novel framework for learning these activation functions neuron-wise in a data-dependent fashion, based on a cheap one-dimensional kernel expansion and the idea of kernel activation functions (KAFs). In this paper we argue that, despite its flexibility, this framework is still limited in the class of functions that can be modeled in the complex domain. We leverage the idea of widely linear complex kernels to extend the formulation, allowing for a richer expressiveness without an increase in the number of adaptable parameters. We test the resulting model on a set of complex-valued image classification benchmarks. Experimental results show that the resulting CVNNs can achieve higher accuracy while at the same time converging faster.
Simone Scardapane, Steven Van Vaerenbergh, Danilo Comminiello, Aurelio Uncini
ICASSP4
2019 Kafnets: Kernel-based non-parametric activation functions for neural networks
Simone Scardapane, Steven Van Vaerenbergh, Simone Totaro, Aurelio Uncini
Neural Networks4
2018 Sparse functional link adaptive filter using an ℓ1-norm regularization
abstract
Linear-in-the-parameters nonlinear adaptive filters often show some sparse behavior due to the fact that not all the coefficients are equally useful for the modeling of any nonlinearity. Recently, proportionate algorithms have been proposed to leverage sparsity behaviors in nonlinear filtering. In this paper, we deal with this problem by introducing a proportionate adaptive algorithm based on an ℓ1-norm penalty of the cost function, which regularizes the solution, to be used for a class of nonlinear filters based on functional links. The proposed algorithm stresses the difference between useful and useless functional links for the purpose of nonlinear modeling. Experimental results clearly show faster convergence performance with respect to the standard (i.e., non-regularized) version of the algorithm.
Danilo Comminiello, Michele Scarpiniti, Simone Scardapane, Aurelio Uncini
ISCAS4
2018 Bayesian Random Vector Functional-Link Networks for Robust Data Modeling
abstract
Random vector functional-link (RVFL) networks are randomized multilayer perceptrons with a single hidden layer and a linear output layer, which can be trained by solving a linear modeling problem. In particular, they are generally trained using a closed-form solution of the (regularized) least-squares approach. This paper introduces several alternative strategies for performing full Bayesian inference (BI) of RVFL networks. Distinct from standard or classical approaches, our proposed Bayesian training algorithms allow to derive an entire probability distribution over the optimal output weights of the network, instead of a single pointwise estimate according to some given criterion (e.g., least-squares). This provides several known advantages, including the possibility of introducing additional prior knowledge in the training process, the availability of an uncertainty measure during the test phase, and the capability of automatically inferring hyper-parameters from given data. In this paper, two BI algorithms for regression are first proposed that, under some practical assumptions, can be implemented by a simple iterative process with closed-form computations. Simulation results show that one of the proposed algorithms, Bayesian RVFL, is able to outperform standard training algorithms for RVFL networks with a proper regularization factor selected carefully via a line search procedure. A general strategy based on variational inference is also presented, with an application to data modeling problems with noisy outputs or outliers. As we discuss in this paper, using recent advances in automatic differentiation this strategy can be applied to a wide range of additional situations in an immediate fashion.
Simone Scardapane, Dianhui Wang 0001, Aurelio Uncini
IEEE Trans. Cybern.3
2018 Energy performance of heuristics and meta-heuristics for real-time joint resource scaling and consolidation in virtualized networked data centers
Michele Scarpiniti, Enzo Baccarelli, Paola Gabriela Vinueza Naranjo, Aurelio Uncini
J. Supercomput.4
2017 On the use of deep recurrent neural networks for detecting audio spoofing attacks
abstract
Biometric security systems based on predefined speech sentences are extremely common nowadays, particularly in low-cost applications where the simplicity of the hardware involved is a great advantage. Audio spoofing verification is the problem of detecting whether a speech segment acquired from such a system is genuine, or whether it was synthesized or modified by a computer in order to make it sound like an authorized person. Developing countermeasures for spoofing attacks is clearly essential for having effective biometric and security systems based on audio features, all the more significant due to recent advances in generative machine learning. Nonetheless, the problem is complicated by the possible lack of knowledge on the technique(s) used to put forward the attack, so that anti-spoofing systems should be able to withstand also spoofing attacks that were not considered explicitly in the training stage. In this paper, we analyze the use of deep recurrent networks applied to this task, i.e. networks made by the successive combination of multiple feedforward and recurrent layers. These networks are routinely used in speech recognition and language identification but, to the best of our knowledge, they were never considered for this specific problem. We evaluate several architectures on the dataset released for the ASVspoof 2015 challenge last year. We show that, by working with very standard feature extraction routines and with a minimum amount of fine-tuning, the networks can already reach very promising error rates, comparable to state-of-the-art approaches, paving the way to further investigations on the problem using deep RNN models.
Simone Scardapane, Lucas Stoffl, Florian Röhrbein, Aurelio Uncini
IJCNN4
2017 Group sparse regularization for deep neural networks
Simone Scardapane, Danilo Comminiello, Amir Hussain 0001, Aurelio Uncini
Neurocomputing4
2017 Combined nonlinear filtering architectures involving sparse functional link adaptive filters
Danilo Comminiello, Michele Scarpiniti, Luis Antonio Azpicueta-Ruiz, Jerónimo Arenas-García, Aurelio Uncini
Signal Process.5
2017 Frequency domain quaternion adaptive filters: Algorithms and convergence performance
Francesca Ortolani, Danilo Comminiello, Michele Scarpiniti, Aurelio Uncini
Signal Process.4
2017 Fully Decentralized Semi-supervised Learning via Privacy-preserving Matrix Completion
abstract
Distributed learning refers to the problem of inferring a function when the training data are distributed among different nodes. While significant work has been done in the contexts of supervised and unsupervised learning, the intermediate case of Semi-supervised learning in the distributed setting has received less attention. In this paper, we propose an algorithm for this class of problems, by extending the framework of manifold regularization. The main component of the proposed algorithm consists of a fully distributed computation of the adjacency matrix of the training patterns. To this end, we propose a novel algorithm for low-rank distributed matrix completion, based on the framework of diffusion adaptation. Overall, the distributed Semi-supervised algorithm is efficient and scalable, and it can preserve privacy by the inclusion of flexible privacy-preserving mechanisms for similarity computation. The experimental results and comparison on a wide range of standard Semi-supervised benchmarks validate our proposal.
Roberto Fierimonte, Simone Scardapane, Aurelio Uncini, Massimo Panella
IEEE Trans. Neural Networks Learn. Syst.3
2016 Design of hybrid nonlinear spline adaptive filters for active noise control
abstract
In this paper, we focus on the problem of removing noise in the acoustic domain. To this end, we introduce a class of hybrid nonlinear spline filters, which are designed as a cascade of an adaptive spline function and a single layer adaptive nonlinear network. The adaptive nonlinear networks employed in this work are the functional link network and the even mirror Fourier nonlinear network. Suitable update rules, which not only update the adaptive weights of the nonlinear networks, but also introduce adaptability in the developed spline function are derived. The proposed nonlinear filters have been successfully applied to nonlinear system identification as well as nonlinear active noise control. The new filters have been shown to outperform other popular nonlinear filters.
Vinal Patel, Danilo Comminiello, Michele Scarpiniti, Nithin V. George, Aurelio Uncini
IJCNN5
2016 Distributed spectral clustering based on Euclidean distance matrix completion
abstract
In this paper, we consider the problem of distributed spectral clustering, wherein the data to be clustered is (horizontally) partitioned over a set of interconnected agents with limited connectivity. In order to solve it, we consider the equivalent problem of reconstructing the Euclidean distance matrix of pairwise distances among the joint set of datapoints. This is obtained in a fully decentralized fashion, making use of an innovative distributed gradient-based procedure, where at every agent we interleave gradient steps on a low-rank factorization of the distance matrix, with local averaging steps considering all its neighbors' current estimates. The procedure can be applied to any spectral clustering algorithm, including normalized and unnormalized variations, for multiple choices of the underlying Laplacian matrix. Experimental evaluations demonstrate that the solution is competitive with a fully centralized solver, where data is collected beforehand on a (virtual) coordinating agent.
Simone Scardapane, Rosa Altilio, Massimo Panella, Aurelio Uncini
IJCNN4
2016 A semi-supervised random vector functional-link network based on the transductive framework
Simone Scardapane, Danilo Comminiello, Michele Scarpiniti, Aurelio Uncini
Inf. Sci.4
2016 Distributed semi-supervised support vector machines
Simone Scardapane, Roberto Fierimonte, Paolo Di Lorenzo, Massimo Panella, Aurelio Uncini
Neural Networks5
2015 Functional link expansions for nonlinear modeling of audio and speech signals
abstract
Nonlinear distortions pose a serious problem for the quality preservation of audio and speech signals. To address this problem, such signals are processed by nonlinear models. Functional link adaptive filter (FLAF) is a linear-in-the-parameter nonlinear model, whose nonlinear transformation of the input is characterized by a basis function expansion, satisfying the universal approximation properties. Since the expansion type affects the nonlinear modeling according to the nature of the input signal, in this paper we investigate the FLAF modeling performance involving the most popular functional expansions when audio and speech signals are processed. A comprehensive analysis is conducted to provide the best suitable solution for the processing of nonlinear signals. Experimental results are assessed also in terms of signal quality and intelligibility.
Danilo Comminiello, Simone Scardapane, Michele Scarpiniti, Raffaele Parisi, Aurelio Uncini
IJCNN5
2015 An interactive optimization procedure for stereophonic acoustic echo cancellation systems
abstract
Acoustic echo cancellers are used in teleconferencing systems in order to reduce undesired echoes due to coupling between microphones and loudspeakers. Stereophonic systems provide more realistic experience than single-channel systems, since listeners have spatial information that helps to identify the speaker position. Assuming this scenario, a suitable choice for the system parameters becomes essential to improve the audio reproduction quality. Error-driven optimization strategies are usually used to obtain an optimal system configuration but there is no relationship with the quality desired by the user. In this paper, an interactive evolutionary algorithm is adopted for a stereophonic acoustic echo cancellation system in order to meet subjective specifications in the optimization stage. In this way, the optimal system configuration is derived according to a user-driven approach in order to satisfy the quality requirements demanded by users availing such stereophonic systems. Experimental results prove the effectiveness of the proposed interactive architecture according to both objective and subjective measures.
Laura Romoli, Stefania Cecchi, Francesco Piazza, Danilo Comminiello, Michele Scarpiniti, Aurelio Uncini
IJCNN6
2015 Distributed music classification using Random Vector Functional-Link nets
abstract
In this paper, we investigate the problem of music classification when training data is distributed throughout a network of interconnected agents (e.g. computers, or mobile devices), and it is available in a sequential stream. Under the considered setting, the task is for all the nodes, after receiving any new chunk of training data, to agree on a single classifier in a decentralized fashion, without reliance on a master node. In particular, in this paper we propose a fully decentralized, sequential learning algorithm for a class of neural networks known as Random Vector Functional-Link nets. The proposed algorithm does not require the presence of a single coordinating agent, and it is formulated exclusively in term of local exchanges between neighboring nodes, thus making it useful in a wide range of realistic situations. Experimental simulations on four music classification benchmarks show that the algorithm has comparable performance with respect to a centralized solution, where a single agent collects all the local data from every node and subsequently updates the model.
Simone Scardapane, Roberto Fierimonte, Dianhui Wang 0001, Massimo Panella, Aurelio Uncini
IJCNN5
2015 Distributed learning for Random Vector Functional-Link networks
Simone Scardapane, Dianhui Wang 0001, Massimo Panella, Aurelio Uncini
Inf. Sci.4
2015 Prediction of telephone calls load using Echo State Network with exogenous variables
Filippo Maria Bianchi, Simone Scardapane, Aurelio Uncini, Antonello Rizzi, Alireza Sadeghian
Neural Networks3
2015 Improving nonlinear modeling capabilities of functional link adaptive filters
Danilo Comminiello, Michele Scarpiniti, Simone Scardapane, Raffaele Parisi, Aurelio Uncini
Neural Networks5
2015 Nonlinear system identification using IIR Spline Adaptive Filters
Michele Scarpiniti, Danilo Comminiello, Raffaele Parisi, Aurelio Uncini
Signal Process.4
2015 Intelligent Acoustic Interfaces With Multisensor Acquisition for Immersive Reproduction
abstract
Immersive speech communication systems have been gaining increasing attention due to their ability to reproduce enhanced acoustic images, and thus achieving good performance in terms of sound quality and accuracy. In this context , a fundamental role is played by intelligent acoustic interfaces (IAIs), which aim at acquiring and/or reproducing desired acoustic information with enhanced perception. The recent widespread availability of multimedia devices, equipped with different kind of sensors, has broadened the range of data processing methods, thus giving a chance for developing advanced IAIs. In this paper, we propose an immersive communication system composed of two IAIs: the first one exploits microphones and cameras, together with a signal processing system, to reduce unwanted noise and enhance the speech quality of the desired information in the transmitting room; the second one is an advanced reproduction system based on a loudspeaker array and on an effective wave field synthesis technique capable of reproducing the spatial perception of the desired speech source in the receiving room. The whole system has been assessed in simulated and real immersive communication scenarios: objective and subjective evaluations have been shown the effectiveness of the proposed system.
Danilo Comminiello, Stefania Cecchi, Michele Scarpiniti, Michele Gasparini, Laura Romoli, Francesco Piazza, Aurelio Uncini
IEEE Trans. Multim.7
2015 Online Sequential Extreme Learning Machine With Kernels
abstract
The extreme learning machine (ELM) was recently proposed as a unifying framework for different families of learning algorithms. The classical ELM model consists of a linear combination of a fixed number of nonlinear expansions of the input vector. Learning in ELM is hence equivalent to finding the optimal weights that minimize the error on a dataset. The update works in batch mode, either with explicit feature mappings or with implicit mappings defined by kernels. Although an online version has been proposed for the former, no work has been done up to this point for the latter, and whether an efficient learning algorithm for online kernel-based ELM exists remains an open problem. By explicating some connections between nonlinear adaptive filtering and ELM theory, in this brief, we present an algorithm for this task. In particular, we propose a straightforward extension of the well-known kernel recursive least-squares, belonging to the kernel adaptive filtering (KAF) family, to the ELM framework. We call the resulting algorithm the kernel online sequential ELM (KOS-ELM). Moreover, we consider two different criteria used in the KAF field to obtain sparse filters and extend them to our context. We show that KOS-ELM, with their integration, can result in a highly efficient algorithm, both in terms of obtained generalization error and training time. Empirical evaluations demonstrate interesting results on some benchmarking datasets.
Simone Scardapane, Danilo Comminiello, Michele Scarpiniti, Aurelio Uncini
IEEE Trans. Neural Networks Learn. Syst.4
2014 GP-based kernel evolution for L2-Regularization Networks
abstract
In kernel-based learning methods, a crucial design parameter is given by the choice of the kernel function to be used. Although there is, in theory, an infinite range of potential candidates, a handful of kernels covers the majority of actual applications. Partly, this is due to the difficulty of choosing an optimal kernel function in absence of a-priori information. In this respect, Genetic Programming (GP) techniques have shown interesting capabilities of learning non-trivial kernel functions that outperform commonly used ones. However, experiments have been restricted to the use of Support Vector Machines (SVMs), and have not addressed some problems that are specific to GP implementations, such as diversity maintenance. In these respects, the aim of this paper is twofold. First, we present a customized GP-based kernel search method that we apply using an L2-Regularization Network as the base learning algorithm. Second, we investigate the problem of diversity maintenance in the context of kernel evolution, and test an adaptive criterion for maintaining it in our algorithm. For the former point, experiments show a gain in accuracy for our method against fine-tuned standard kernels. For the latter, we show that diversity is decreasing critically fast during the GP iterations, but this decrease does not seems to affect performance of the algorithm.
Simone Scardapane, Danilo Comminiello, Michele Scarpiniti, Aurelio Uncini
IEEE Congress on Evolutionary Computation4
2014 An interpretable graph-based image classifier
abstract
The generalization capability is usually recognized as the most desired feature of data-driven learning systems, such as classifiers. However, in many practical applications obtaining human-understandable information, relevant to the problem at hand, from the classidication model can be equally important. In this paper we propose a classification system able to fulfill these two requirements simultaneously for a generic image classification task. As a first preprocessing step, an input image to the classifier is represented by a labeled graph, relying on a segmentation algorithm. The graph is conceived to represent visual and topological information of the relevant segments of the image. Then, the graph is classified by a suited inductive inference engine. In the learning procedure all the training set images are represented by graphs, feeding a state-of-the-art classification system working on structured domains. The synthesis procedure consists in extracting characterizing subgraphs from the training set, which are used to embed the graphs into a vector space, enabling thus the applicability of well-known classifiers for feature-based patterns. Such characterizing subgraphs, which are derived in an unsupervised fashion, are interpretable by suitable field experts, allowing a semantic analysis of the discovered classification rules for the given problem at hand. The system is optimized with a genetic algorithm, which tunes the system parameters according to a cross-validation scheme. We show the validity of the approach by performing experiments considering some image classification problems derived from an on-line repository.
Filippo Maria Bianchi, Simone Scardapane, Lorenzo Livi, Aurelio Uncini, Antonello Rizzi
IJCNN4
2014 Advanced intelligent acoustic interfaces for multichannel audio reproduction
abstract
Nowadays, there is a large interest towards multimedia audio systems as a consequence of the development of advanced digital signal processing techniques. In particular, immersive speech communication system has gaining increasing attention since they allow to reproduce realistic acoustic image, and thus achieving good performance in terms of sound quality and accuracy. In this scenario a fundamental role is played by intelligent acoustic interfaces which aim at acquiring audio information, processing it, and returning the processed information to the audio rendering system. In this paper, an effective intelligent acoustic interface composed of a microphone array and a signal processing system capable to enhance the intelligibility of the desired information of the transmitting room is proposed, combined with an advanced reproduction system based on an efficient application of a wave field synthesis technique capable to reproduce an immersive scenario in the receiving room. The whole system has been assessed within a speech communication application involving a moving desired source in a real scenario: objective and subjective evaluation has been reported in order to show the overall system performance.
Danilo Comminiello, Stefania Cecchi, Michele Gasparini, Michele Scarpiniti, Aurelio Uncini, Francesco Piazza
IJCNN5
2014 An effective criterion for pruning reservoir's connections in Echo State Networks
abstract
Echo State Networks (ESNs) were introduced to simplify the design and training of Recurrent Neural Networks (RNNs), by explicitly subdividing the recurrent part of the network, the reservoir, from the non-recurrent part. A standard practice in this context is the random initialization of the reservoir, subject to few loose constraints. Although this results in a simple-to-solve optimization problem, it is in general suboptimal, and several additional criteria have been devised to improve its design. In this paper we provide an effective algorithm for removing redundant connections inside the reservoir during training. The algorithm is based on the correlation of the states of the nodes, hence it depends only on the input signal, it is efficient to implement, and it is also local. By applying it, we can obtain an optimally sparse reservoir in a robust way. We present the performance of our algorithm on two synthetic datasets, which show its effectiveness in terms of better generalization and lower computational complexity of the resulting ESN. This behavior is also investigated for increasing levels of memory and non-linearity required by the task.
Simone Scardapane, Gabriele Nocco, Danilo Comminiello, Michele Scarpiniti, Aurelio Uncini
IJCNN5
2014 Hammerstein uniform cubic spline adaptive filters: Learning and convergence properties
Michele Scarpiniti, Danilo Comminiello, Raffaele Parisi, Aurelio Uncini
Signal Process.4
2014 Nonlinear Acoustic Echo Cancellation Based on Sparse Functional Link Representations
abstract
Recently, a new class of nonlinear adaptive filtering architectures has been introduced based on the functional link adaptive filter (FLAF) model. Here we focus specifically on the split FLAF (SFLAF) architecture, which separates the adaptation of linear and nonlinear coefficients using two different adaptive filters in parallel. This property makes the SFLAF a well-suited method for problems like nonlinear acoustic echo cancellation (NAEC), in which the separation of filtering tasks brings some performance improvement. Although flexibility is one of the main features of the SFLAF, some problem may occur when the nonlinearity degree of the input signal is not known a priori. This implies a non-optimal choice of the number of coefficients to be adapted in the nonlinear path of the SFLAF. In order to tackle this problem, we propose a proportionate FLAF (PFLAF), which is based on sparse representations of functional links, thus giving less importance to those coefficients that do not actively contribute to the nonlinear modeling. Experimental results show that the proposed PFLAF achieves performance improvement with respect to the SFLAF in several nonlinear scenarios.
Danilo Comminiello, Michele Scarpiniti, Luis Antonio Azpicueta-Ruiz, Jerónimo Arenas-García, Aurelio Uncini
IEEE ACM Trans. Audio Speech Lang. Process.5
2013 Combined adaptive beamforming schemes for nonstationary interfering noise reduction
Danilo Comminiello, Michele Scarpiniti, Raffaele Parisi, Aurelio Uncini
Signal Process.4
2013 Nonlinear spline adaptive filtering
Michele Scarpiniti, Danilo Comminiello, Raffaele Parisi, Aurelio Uncini
Signal Process.4
2013 Functional Link Adaptive Filters for Nonlinear Acoustic Echo Cancellation
abstract
This paper introduces a new class of nonlinear adaptive filters, whose structure is based on Hammerstein model. Such filters derive from the functional link adaptive filter (FLAF) model, defined by a nonlinear input expansion, which enhances the representation of the input signal through a projection in a higher dimensional space, and a subsequent adaptive filtering. In particular, two robust FLAF-based architectures are proposed and designed ad hoc to tackle nonlinearities in acoustic echo cancellation (AEC). The simplest architecture is the split FLAF, which separates the adaptation of linear and nonlinear elements using two different adaptive filters in parallel. In this way, the architecture can accomplish distinctly at best the linear and the nonlinear modeling. Moreover, in order to give robustness against different degrees of nonlinearity, a collaborative FLAF is proposed based on the adaptive combination of filters. Such architecture allows to achieve the best performance regardless of the nonlinearity degree in the echo path. Experimental results show the effectiveness of the proposed FLAF-based architectures in nonlinear AEC scenarios, thus resulting an important solution to the modeling of nonlinear acoustic channels.
Danilo Comminiello, Michele Scarpiniti, Luis Antonio Azpicueta-Ruiz, Jerónimo Arenas-García, Aurelio Uncini
IEEE Trans. Speech Audio Process.5
2012 Cepstrum Prefiltering for Binaural Source Localization in Reverberant Environments
abstract
Binaural sound source localization can be performed by imitation of the fundamental mechanisms of the human auditory system, which is based on the integrated effects of ear, pinnae, head and torso. In particular, two physical cues can be exploited, i.e. the Interaural Time Difference (ITD) and the Interaural Level Difference (ILD). It is known that joint use of ITD and ILD provides good source azimuth estimations. In many practical situations binaural localization has to be performed in closed environments, where the presence of reverberation degrades the performance of available position estimators. In this paper a possible solution to this difficult problem is introduced. The proposed solution is based on proper use of cepstral prefiltering prior to source localization by ITD and ILD. It is shown that cepstrum can help in reducing the effects of reverberation, thus yielding better location estimates.
Raffaele Parisi, Flavia Camoes, Michele Scarpiniti, Aurelio Uncini
IEEE Signal Process. Lett.4
2010 A novel affine projection algorithm for superdirective microphone array beamforming
abstract
This paper describes a new adaptive algorithm and assesses its effectiveness within speech enhancement applications. The proposed variable step size block exact APA (VSS-BEAPA) filtering is based on the affine projection algorithm (APA) and introduces a block processing with a variable step size that allows to consider under-modeling scenarios. The algorithm shows improved convergence performance and computational efficiency and its robustness is proved in a typical context of a hands-free teleconferencing application in a noisy environment. The experiments show that a microphone array system joined with VSS-BEAPA filtering is capable of both decreasing the noise level and enhancing the speech signal quality.
Danilo Comminiello, Michele Scarpiniti, Raffaele Parisi, Aurelio Uncini
ISCAS4
2010 Improved TDOA disambiguation techniques for sound source localization in reverberant environments
abstract
Given a single sound source in a non reverberant environment, an estimate of the Time Difference Of Arrival (TDOA) between microphones can be obtained by observing the time value at which the cross correlation of the two microphone signals displays a maximum. In the presence of reverberation, however, the cross correlation function displays a great number of local maxima caused by reflected signals blending unpredictably. This results in TDOA estimation ambiguity, making correct source localization impossible. The aim of this work is to present an improved disambiguation algorithm and compare it to existing algorithms in terms of estimation accuracy and processing time.
Cecilia Maria Zannini, Albenzio Cirillo, Raffaele Parisi, Aurelio Uncini
ISCAS4
2008 Sound mapping in reverberant rooms by a robust direct method
abstract
Direct methods estimate the position of an acoustic source by sampling the environment through a set of properly placed microphones. SRP-Phat is probably the most popular direct method. It is based on computation of the generalized cross- correlation (GCC) of signals on a grid of preselected points. Anyway, in the presence of reverberation, the functional employed by SRP-Phat can be very irregular from point to point, thus making the source localization a difficult task. In this paper, a new functional is presented that regularizes the SRP- Phat approach and makes it more efficient the use of optimization algorithms to further refine the source position estimation. After a brief introduction, the proposed approach is described and compared to SRP-Phat on simulated and real data at different reverberation levels.
Albenzio Cirillo, Raffaele Parisi, Aurelio Uncini
ICASSP3
2008 Model order selection for estimation of Common Acoustical Poles
abstract
Common acoustical poles provide a convenient way of modeling different Room Transfer Functions relative to the same environment, by retaining the resonant properties of the room itself under the form of common auto-regressive coefficients; different methods have been proposed to estimate the latter, but all need an order to be priorly established. Although, for a few room geometries, a theoretical order can be computed, actual usable values (always much smaller) may depend on the particular chosen estimation configuration, and no specific method is available for their determination. In this work, a novel empirical technique to select an appropriate model order is presented. This technique is tested in a simulated environment and the results are presented, compared to the theoretical values.
Gabriele Bunkheila, Raffaele Parisi, Aurelio Uncini
ISCAS3
2008 Genre classification of compressed audio data
abstract
This paper deals with the musical genre classification problem, starting from a set of features extracted directly from MPEG-1 layer III compressed audio data. The automatic classification of compressed audio signals into a short hierarchy of musical genres is explored. More specifically, three feature sets for representing timbre, rhythmic content and energy content are proposed for a four leafs tree genre hierarchy. The adopted set of features are computed from the spectral information available in the MPEG decoding stage. The performance and relative importance of the proposed approach is investigated by training a classification model using the audio collections proposed in musical genre contests. We also used an optimization strategy based on genetic algorithms. The results are comparable to those obtained by PCM-based musical genre classification systems.
Antonello Rizzi, Nicola Maurizio Buccino, Massimo Panella, Aurelio Uncini
MMSP4
2008 Flexible Nonlinear Blind Signal Separation in the Complex Domain
abstract
This paper introduces an Independent Component Analysis (ICA) approach to the separation of nonlinear mixtures in the complex domain. Source separation is performed by a complex INFOMAX approach. The neural network which realizes the separation employs the so called "Mirror Model" and is based on adaptive activation functions, whose shape is properly modified during learning. Nonlinear functions involved in the processing of complex signals are realized by pairs of spline neurons called "splitting functions", working on the real and the imaginary part of the signal respectively. Theoretical proof of existence and uniqueness of the solution under proper assumptions is also provided. In particular a simple adaptation algorithm is derived and some experimental results that demonstrate the effectiveness of the proposed solution are shown.
Daniele Vigliano, Michele Scarpiniti, Raffaele Parisi, Aurelio Uncini
Int. J. Neural Syst.4
2008 Generalized splitting functions for blind separation of complex signals
Michele Scarpiniti, Daniele Vigliano, Raffaele Parisi, Aurelio Uncini
Neurocomputing4
2007 Source Localization in Reverberant Environments by Consistent Peak Selection
abstract
Acoustic source localization in the presence of reverberation is a difficult task. Conventional approaches, based on time delay estimation performed by generalized cross correlation (GCC) on a set of microphone pairs, followed by geometric triangulation, are often unsatisfactory. Prefiltering is usually adopted to reduce the spurious peaks due to reflections. In this work an alternative strategy is proposed, based on the concept that secondary peaks of the GCCs can be crucial in order to correctly locate the source. More specifically, an iterative weighting procedure is introduced, based on the rationale that peaks corresponding to the actual source position should be consistently weighted. The position estimate is then refined by use of an effective and fast clustering technique. Experimental results on simulated data demonstrate the effectiveness of the proposed solution.
Raffaele Parisi, Albenzio Cirillo, Massimo Panella, Aurelio Uncini
ICASSP (1)4
2006 Particle swarm localization of acoustic sources in the presence of reverberation
abstract
In this work a novel approach to acoustic source localization in reverberant environments is introduced. The proposed method effectively employs particle filtering (PF) and particle swarm optimization (PSO) in an integrated framework. Source localization is viewed as a global minimization problem where the solution is searched by properly exploiting competition and cooperation among the individuals of a population. Experiments on simulated data show that the proposed strategy is fast and robust with respect to the presence of reverberation
Raffaele Parisi, P. Croene, Aurelio Uncini
ISCAS3
2006 An IIR architecture for BSS in strong nonlinear convolutive environments
abstract
This paper introduces an IIR flexible ICA approach to the problem of blind source separation in convolutive nonlinear environments. The proposed algorithm performs separation in the presence of convolutive mixing of post nonlinear convolutive mixtures (CPNL-C), it is based on redundancy reduction and realizes source separation by minimizing the output mutual information. Experimental results are described to show the effectiveness of the described technique
Daniele Vigliano, Raffaele Parisi, Aurelio Uncini
ISCAS3
2005 An information theoretic approach to a novel nonlinear independent component analysis paradigm
Daniele Vigliano, Raffaele Parisi, Aurelio Uncini
Signal Process.3
2004 A novel recurrent network for independent component analysis of post nonlinear convolutive mixtures
abstract
The paper introduces a novel independent component analysis approach to the separation of nonlinear convolutive mixtures. In particular, convolutive mixing of post nonlinear mixtures is considered. Source separation is performed by a new efficient recurrent network, which is able to ensure faster training with respect to currently available feedforward architectures, with lower computational costs. The proposed architecture makes proper use of flexible spline neurons for on-line estimation of the score function. Experimental results are described to demonstrate the effectiveness of the proposed technique.
Daniele Vigliano, Raffaele Parisi, Aurelio Uncini
ICASSP (5)3
2004 Regularising neural networks using flexible multivariate activation function
Mirko Solazzi, Aurelio Uncini
Neural Networks2
2003 Special issue on evolving solution with neural networks
Alessandra Fanni, Aurelio Uncini
Neurocomputing2
2003 Audio signal processing by neural networks
Aurelio Uncini
Neurocomputing1
2003 Blind signal processing by complex domain adaptive spline neural networks
abstract
In this paper, neural networks based on an adaptive nonlinear function suitable for both blind complex time domain signal separation and blind frequency domain signal deconvolution, are presented. This activation function, whose shape is modified during learning, is based on a couple of spline functions, one for the real and one for the imaginary part of the input. The shape control points are adaptively changed using gradient-based techniques. B-splines are used, because they allow to impose only simple constraints on the control parameters in order to ensure a monotonously increasing characteristic. This new adaptive function is then applied to the outputs of a one-layer neural network in order to separate complex signals from mixtures by maximizing the entropy of the function outputs. We derive a simple form of the adaptation algorithm and present some experimental results that demonstrate the effectiveness of the proposed method.
Aurelio Uncini, Francesco Piazza
IEEE Trans. Neural Networks1
2002 Blind source separation of convolutive nonlinear mixtures by flexible spline nonlinear functions
abstract
In this paper a nonlinear deconvolving system, based on the use of the recently introduced flexible activation function whose control points are adaptively changed, is proposed.
Fabio Milani, Mirko Solazzi, Aurelio Uncini
ICASSP3
2002 Subband neural networks prediction for on-line audio signal recovery
abstract
In this paper, a subbands multirate architecture is presented for audio signal recovery. Audio signal recovery is a common problem in digital music signal restoration field, because of corrupted samples that must be replaced. The subband approach allows for the reconstruction of a long audio data sequence from forward-backward predicted samples. In order to improve prediction performances, neural networks with spline flexible activation function are used as narrow subband nonlinear forward-backward predictors. Previous neural-networks approaches involved a long training process. Due to the small networks needed for each subband and to the spline adaptive activation functions that speed-up the convergence time and improve the generalization performances, the proposed signal recovery scheme works in online (or in continuous learning) mode as a simple nonlinear adaptive filter. Experimental results show the mean square reconstruction error and maximum error obtained with increasing gap length, from 200 to 5000 samples for different musical genres. A subjective performances analysis is also reported. The method gives good results for the reconstruction of over 100 ms of audio signal with low audible effects in overall quality and outperforms the previous approaches.
Gianandrea Cocchi, Aurelio Uncini
IEEE Trans. Neural Networks2
2001 Subbands audio signal recovering using neural nonlinear prediction
abstract
Audio signal recovery is a common problem in the digital audio restoration field, because of corrupted samples that must be replaced. In this paper a subband architecture is presented for audio signal recovery, using neural nonlinear prediction based on adaptive spline neural networks. The experimental results show the mean square reconstruction error, and maximum error obtained with increasing gap length, from 200 to 5000 samples. The method gives good results allowing the reconstruction of over 100 ms of signal with low audible effects in overall quality.
Gianandrea Cocchi, Aurelio Uncini
ICASSP2
2001 Nonlinear blind source separation by spline neural networks
abstract
In this paper a new neural network model for blind demixing of nonlinear mixtures is proposed. We address the use of the adaptive spline neural network recently introduced for supervised and unsupervised neural networks. These networks are built using neurons with flexible B-spline activation functions and in order to separate signals from mixtures, a gradient-ascending algorithm which maximizes the outputs entropy is derived. In particular a suitable architecture composed by two layers of flexible nonlinear functions for the separation of nonlinear mixtures is proposed. Some experimental results that demonstrate the effectiveness of the proposed neural architecture are presented.
Mirko Solazzi, Francesco Piazza, Aurelio Uncini
ICASSP3
2001 Complex discriminative learning Bayesian neural equalizer
Mirko Solazzi, Aurelio Uncini, Elio D. Di Claudio, Raffaele Parisi
Signal Process.2
2000 Neural equalizer with adaptive multidimensional spline activation functions
abstract
This paper presents a new neural architecture suitable for digital signal processing application. The architecture, based on adaptable multidimensional activation functions, allows one to collect information from the previous network layer in aggregate form. In other words the number of network connections (structural complexity) can be very low respect to the problem complexity. This fact, as experimentally demonstrated in the paper, improve the network generalization capabilities and speed up the convergence of the learning process. A specific learning algorithm is derived and experimental results, on channel equalization, demonstrate the effectiveness of the proposed architecture.
Mirko Solazzi, Aurelio Uncini, Francesco Piazza
ICASSP2
2000 Low Complexity Adaptive Non-Linear Function for Blind Signal Separation
abstract
An adaptive nonlinear function for blind signal separation is presented. It is based on a spline approximation whose control points are adaptively changed using information maximization techniques. The monotonously increasing characteristic is obtained using suitable B-spline functions imposing simple constraints on its control points. In particular, the problem of adaptively maximizing the entropy of the output is considered in the context of blind separation of independent sources. We derive a simple form of the learning algorithm which allows us not only to adapt the separation matrix coefficients but also the shape of the nonlinear functions. A comparison with the mixture-of-densities approach is also presented on some experimental data that demonstrates the effectiveness and efficiency of the proposed method.
Andrea Pierani, Francesco Piazza, Mirko Solazzi, Aurelio Uncini
IJCNN (3)4
2000 Artificial Neural Networks with Adaptive Multidimensional Spline Activation Functions
abstract
This work concerns a new kind of neural structure that involves a multidimensional adaptive activation function. The proposed architecture, based on multidimensional cubic spline, allows to collect information from the previous network layer in aggregate form. In other words the number of network connections (structural complexity) can be very low respect to the problem complexity. This fact, as experimentally demonstrated in the paper, improve the network generalization capabilities and speed up the convergence of the learning process. A specific learning algorithm is derived and experimental results demonstrate the effectiveness of the proposed architecture.
Mirko Solazzi, Aurelio Uncini
IJCNN (3)2
2000 A Signal-Flow-Graph Approach to On-line Gradient Calculation
abstract
A large class of nonlinear dynamic adaptive systems such as dynamic recurrent neural networks can be effectively represented by signal flow graphs (SFGs). By this method, complex systems are described as a general connection of many simple components, each of them implementing a simple one-input, one-output transformation, as in an electrical circuit. Even if graph representations are popular in the neural network community, they are often used for qualitative description rather than for rigorous representation and computational purposes. In this article, a method for both on-line and batch-backward gradient computation of a system output or cost function with respect to system parameters is derived by the SFG representation theory and its known properties. The system can be any causal, in general nonlinear and time-variant, dynamic system represented by an SFG, in particular any feedforward, time-delay, or recurrent neural network. In this work, we use discrete-time notation, but the same theory holds for the continuous-time case. The gradient is obtained in a straightforward way by the analysis of two SFGs, the original one and its adjoint (obtained from the first by simple transformations), without the complex chain rule expansions of derivatives usually employed. This method can be used for sensitivity analysis and for learning both off-line and on-line. On-line learning is particularly important since it is required by many real applications, such as digital signal processing, system identification and control, channel equalization, and predistortion.
Paolo Campolucci, Aurelio Uncini, Francesco Piazza
Neural Comput.2
1999 Frequency recovery of narrow-band speech using adaptive spline neural networks
abstract
A new system for speech quality enhancement (SQE) is presented. A SQE system attempts to recover the high and low frequencies from a narrow-band speech signal, usually working as a post-processor at the receiver side of a transmission system. The new system operates directly in the frequency domain using complex-valued neural networks. In order to reduce the computational burden and improve the generalization capabilities, a new architecture based on a previously introduced neural network, called the adaptive spline neural network (ASNN), is employed. Experimental results demonstrate the effectiveness of the proposed method.
Aurelio Uncini, Frencesco Gobbi, Francesco Piazza
ICASSP1
1999 On-line learning algorithms for locally recurrent neural networks
abstract
This paper focuses on on-line learning procedures for locally recurrent neural networks with emphasis on multilayer perceptron (MLP) with infinite impulse response (IIR) synapses and its variations which include generalized output and activation feedback multilayer networks (MLN's). We propose a new gradient-based procedure called recursive backpropagation (RBP) whose on-line version, causal recursive backpropagation (CRBP), presents some advantages with respect to the other on-line training methods. The new CRBP algorithm includes as particular cases backpropagation (BP), temporal backpropagation (TBP), backpropagation for sequences (BPS), Back-Tsoi algorithm among others, thereby providing a unifying view on gradient calculation techniques for recurrent networks with local feedback. The only learning method that has been proposed for locally recurrent networks with no architectural restriction is the one by Back and Tsoi. The proposed algorithm has better stability and higher speed of convergence with respect to the Back-Tsoi algorithm, which is supported by the theoretical development and confirmed by simulations. The computational complexity of the CRBP is comparable with that of the Back-Tsoi algorithm, e.g., less that a factor of 1.5 for usual architectures and parameter settings. The superior performance of the new algorithm, however, easily justifies this small increase in computational burden. In addition, the general paradigms of truncated BPTT and RTRL are applied to networks with local feedback and compared with the new CRBP method. The simulations show that CRBP exhibits similar performances and the detailed analysis of complexity reveals that CRBP is much simpler and easier to implement, e.g., CRBP is local in space and in time while RTRL is not local in space.
Paolo Campolucci, Aurelio Uncini, Francesco Piazza, Bhaskar D. Rao
IEEE Trans. Neural Networks2
1999 Multilayer feedforward networks with adaptive spline activation function
abstract
In this paper, a new adaptive spline activation function neural network (ASNN) is presented. Due to the ASNN's high representation capabilities, networks with a small number of interconnections can be trained to solve both pattern recognition and data processing real-time problems. The main idea is to use a Catmull-Rom cubic spline as the neuron's activation function, which ensures a simple structure suitable for both software and hardware implementation. Experimental results demonstrate improvements in terms of generalization capability and of learning speed in both pattern recognition and data processing tasks.
S. Guarnieri, Francesco Piazza, Aurelio Uncini
IEEE Trans. Neural Networks3
1998 Learning and Approximation Capabilities of Adaptive Spline Activation Function Neural Networks
Lorenzo Vecci, Francesco Piazza, Aurelio Uncini
Neural Networks3
1997 Application of the MEC Network to Principal Component Analysis and Source Separation
Simone G. O. Fiori, Aurelio Uncini, Francesco Piazza
ICANN2
1997 A new IIR-MLP learning algorithm for on-line signal processing
abstract
We propose a new learning algorithm for locally recurrent neural networks, called truncated recursive backpropagation which can be easily implemented on-line with good performance. Moreover it generalises the algorithm proposed by Waibel et al. (1989) for TDNN, and includes the Back and Tsoi (1991) algorithm as well as BPS and standard on-line backpropagation as particular cases. The proposed algorithm has a memory and computational complexity that can be adjusted by a careful choice of two parameters h and h' and so it is more flexible than a previous algorithm proposed by us. Although for the sake of brevity we present the new algorithm only for IIR-MLP networks, it can be applied also to any locally recurrent neural network. Some computer simulations of dynamical system identification tests, reported in literature, are also presented to assess the performance of the proposed algorithm applied to the IIR-MLP.
Paolo Campolucci, Simone G. O. Fiori, Aurelio Uncini, Francesco Piazza
ICASSP3
1997 A new unsupervised neural learning rule for orthonormal signal processing
abstract
We derive a new class of neural unsupervised learning rules which arises from the analysis of the dynamics of an abstract mechanical system. The corresponding algorithms can be used to solve several problems in the area of digital signal processing, where orthonormal matrices are involved. We present an application which deals with blind separation of sources, i.e. a new method to perform efficient independent component analysis (ICA) of random signals.
Simone G. O. Fiori, Paolo Campolucci, Aurelio Uncini, Francesco Piazza
ICASSP3
1996 Fast adaptive IIR-MLP neural networks for signal processing applications
abstract
Neural networks with internal temporal dynamics can be applied to non-linear DSP problems. The classical fully connected recurrent architectures, can be replaced by less complex neural networks, based on the well known multilayer perceptron (MLP) where the temporal dynamics is modelled by replacing each synapses either with an FIR filter or with an IIR filter. A general learning algorithm (back-propagation-through-time or BPTT) for a dynamic neural MLP has been introduced by P.J. Werbos (1990). This is a non-causal algorithm, being able to work only in batch mode, while many real problems require on-line adaptation. In this paper we show a new on-line learning algorithm for the IIR-MLP networks which is an approximation of the true BPTT and which includes, as a particular case, the already known Back and Tsoi (1991) algorithm. Several computer simulations of identification of dynamical systems are also reported.
Paolo Campolucci, Aurelio Uncini, Francesco Piazza
ICASSP2
1996 Improved power-of-two sharpening filter design by genetic algorithm
abstract
For high-speed low-complexity filter design, it is common practice to constrain the filters' coefficients to be a power-of-two or a sum of powers-of-two terms (P2), avoiding the full multiplication. Tapped interconnection of different sub-filters are sometimes used to enhance the ripple and stop-band attenuation performances. An extension of the simple cascade architectures, suitable for hardware implementation, is the polynomial sharpening technique. Firstly proposed by Kaiser-Hamming, it allows the use of identical sub-filters with a small hardware overhead. We propose a new approach to the design of P2 sharpening filters based on a specific genetic algorithm which optimizes both the FIR sub-filter and the sharpening polynomial coefficients expressed as P2 terms. This allows one to obtain better performances than the classical P2 design techniques when FIR filters with long impulse responses are involved.
Paolo Gentili, Francesco Piazza, Aurelio Uncini
ICASSP3
1995 Efficient genetic algorithm design for power-of-two FIR filters
abstract
This paper presents an efficient genetic approach to the design of digital finite impulse response (FIR) filters with coefficients constrained to be sums of power-of-two terms. To obtain such efficiency, i.e. a reduction of computational costs and an improvement in performance, a specific filter coefficient coding scheme has been studied and implemented. The resulting genetic algorithm (GA) is explained and compared experimentally with other state-of-the-art design techniques on several power-of-two FIR filter design cases. It can be seen that the proposed genetic technique is able to attain results as good as or better than the other methods. Moreover it can be easily implemented on parallel hardware.
Paolo Gentili, Francesco Piazza, Aurelio Uncini
ICASSP3
1994 Efficient DOA Estimation by a Specific SVD Algorithm
abstract
In this paper an alternative algorithm for the Singular Value Decomposition (SVD) of the data matrix used for Direction-Of-Arrival (DOA) estimation is presented. The proposed algorithm transforms the data matrix into a bi-diagonal form by a procedure based on fast Givens rotations. This procedure is specifically tailored for DOA estimation and allows a certain computational saving with respect to the well known Golub-Kahan algorithm.>
A. Massetani, Francesco Piazza, Aurelio Uncini
ISCAS3
1993 Fast neural networks without multipliers
abstract
Multilayer perceptrons (MLPs) with weight values restricted to powers of two or sums of powers of two are introduced. In a digital implementation, these neural networks do not need multipliers but only shift registers when computing in forward mode, thus saving chip area and computation time. A learning procedure, based on backpropagation, is presented for such neural networks. This learning procedure requires full real arithmetic and therefore must be performed offline. Some test cases are presented, concerning MLPs with hidden layers of different sizes, on pattern recognition problems. Such tests demonstrate the validity and the generalization capability of the method and give some insight into the behavior of the learning algorithm.
Michele Marchesi, Gianni Orlandi, Francesco Piazza, Aurelio Uncini
IEEE Trans. Neural Networks4
1991 Non linear satellite radio links equalized using blind neural networks
abstract
After an ad hoc extension of the multilayer perceptron model to complex signals, a blind neural network equalizer is proposed. Several experimental results are also reported which prove the effectiveness of this approach for equalizing a digital satellite radio channel in the presence of mild nonlinearities and intersymbol interference. Moreover, the proposed method works without any knowledge about the nonlinear nature of the channel.>
Nevio Benvenuto, Michele Marchesi, Francesco Piazza, Aurelio Uncini
ICASSP4
1991 Optimal weighted LS AR estimation in presence of impulsive noise
abstract
A procedure for assigning optimal weights to the prediction equations which are used to obtain the parameters of an autoregressive (AR) model for spectrum estimation by the least squares (LS) solution is presented. The set of weights is computed, by linear programming techniques, in order to reduce the effects of strong impulsive noise onto the AR parameter estimate. The method is particularly effective when the Gaussian white noise component is much smaller than both spikes and useful signal. In order to demonstrate the capability of the proposed approach, the results of a simple AR parameter estimation experiment are also reported.>
Elio D. Di Claudio, Gianni Orlandi, Francesco Piazza, Aurelio Uncini
ICASSP4
1990 Results on the application of simulated annealing algorithm for the design of digital filters with powers-of-two coefficients
abstract
Results on the application of the simulated annealing (SA) algorithm to the problem of finding the coefficients of a digital filter with very coarse coefficients values (namely, power-of-two) are presented. For a minimax criterion in the frequency response, the algorithm was found particularly useful for designing a cascade-form FIR (finite impulse response) filter whose performance is known to be superior to that of direct-form filters. Because the SA algorithm is a global optimization method, all stages are designed simultaneously, thus avoiding the classical iterative procedure. The final result is better performance at a higher computational cost.>
Nevio Benvenuto, Michele Marchesi, Aurelio Uncini
ICASSP3
1990 Multi-layer perceptrons with discrete weights
abstract
The feasibility of restricting the weight values in multilayer perceptrons to powers of two or sums of powers of two is studied. Multipliers could be thus replaced by shifters and adders on digital hardware, saving both time and chip area, under the assumption that the neuron activation function is computed through a lookup table (LUT) and that a LUT can be shared among many neurons. A learning procedure based on back-propagation for obtaining a neural network with such discrete weights is presented. This learning procedure requires full real arithmetic and therefore must be performed offline. It starts from a multilayer perceptron with continuous weights learned using back-propagation. Then a weight normalization is made to ensure that the whole shifting dynamics is used and to maximize the match between continuous and discrete weights of neurons sharing the same LUT. Finally, a discrete version of BP algorithm with automatic learning rate control is applied up to convergence. Some test runs on a simple pattern recognition problem show the feasibility of the approach
Michele Marchesi, Gianni Orlandi, Francesco Piazza, L. Pollonara, Aurelio Uncini
IJCNN5
1990 Improved evoked potential estimation using neural network
abstract
The possibility of using the multilayer perceptron (MLP) neural network for the processing of the evoked potentials (EPs) is analyzed. In this case, the process can be conceived as deterministic low amplitude signals (damped sine waves), corresponding to the brain's response to stimuli, embedded in strongly colored noise, the EEG background activity. Typical values of the signal-to-noise ratio are less than 0 dB. The network, used as a nonlinear filter, is trained using iteratively as the input signal one of a set of available EP ensembles and as the target signal another ensemble of the same set. Experimental results, both on synthetic and real data, show that the method provides good results with very few EP ensembles. Therefore, it allows a noteworthy reduction of the signal nonstationarity and the patient's annoyance
Aurelio Uncini, Michele Marchesi, Gianni Orlandi, Francesco Piazza
IJCNN1
1989 Finite wordlength digital filter design using an annealing algorithm
abstract
A very versatile algorithm for the design of finite-wordlength filters is described. It is a simulated annealing algorithm, derived directly from the simulated annealing algorithm for functions of continuous variables which searches over discrete values of the filter coefficients. It has very fast development time and produces filters with good performances. Its only drawback is a very long execution time. Three examples of filter design are included.>
Nevio Benvenuto, Michele Marchesi, Gianni Orlandi, Francesco Piazza, Aurelio Uncini
ICASSP5