EDBT 2026 Demo / reviewers in the wild / expert
Danilo Comminiello
dblp:33/9433
· DBLP profile ↗
58ranked-venue papers
13as first author
32since 2021 · last 2027
0000-0003-4067-4504ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 5 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 22 · 5 first-author · 10 since 2021Systems, architecture and hardware · 4 · 2 first-author · 2 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Physics-informed adaptive filtering for acoustic echo cancellationabstractThis paper introduces a physics-informed adaptive filtering framework for acoustic echo cancellation (AEC). Unlike conventional adaptive algorithms that rely solely on data-driven error minimization, the proposed method incorporates physically motivated priors derived from acoustic wave propagation and room impulse response structure. The echo path estimation problem is formulated as a composite stochastic optimization task, where the instantaneous squared error is regularized by constraints encoding causality, exponential energy decay, time-weighted sparsity of early reflections, spectral smoothness, and slow temporal variation of the acoustic path. The resulting Physics-Informed Normalized Least-Mean-Squares (PI-NLMS) algorithm performs stochastic gradient descent on the regularized cost while enforcing hard causality through projection. The proposed formulation restricts adaptation to a physically plausible echo-path manifold, improving conditioning and reducing variance without substantially increasing computational complexity. Theoretical analysis establishes mean convergence conditions and characterizes the bias-variance trade-off introduced by structured regularization. Simulation results under stationary and time-varying echo paths demonstrate faster convergence, improved steady-state misalignment, and enhanced echo return loss enhancement (ERLE) compared to conventional NLMS and sparsity-aware baselines. Michele Scarpiniti, Danilo Comminiello, Aurelio Uncini |
Signal Process. | 2 |
| 2026 | Metadata, wavelet, and time aware diffusion models for satellite image super resolutionabstractThe acquisition of high-resolution satellite imagery is often constrained by the spatial and temporal limitations of satellite sensors, as well as the high costs associated with frequent observations. These challenges hinder applications such as environmental monitoring, disaster response, and agricultural management, which require fine-grained and high-resolution data. In this paper, we propose MWT-Diff, an innovative framework for satellite image super-resolution (SR) that combines latent diffusion models with wavelet transforms to address these challenges. At the core of the framework is a novel metadata-, wavelet-, and time-aware encoder (MWT-Encoder), which generates embeddings that capture metadata attributes, multi-scale frequency information, and temporal relationships. The embedded feature representations steer the hierarchical diffusion dynamics, through which the model progressively reconstructs high-resolution satellite imagery from low-resolution inputs. This process preserves critical spatial characteristics including textural patterns, boundary discontinuities, and high-frequency spectral components essential for detailed remote sensing analysis. The comparative analysis of MWT-Diff across multiple datasets demonstrated favorable performance compared to recent approaches, as measured by standard perceptual quality metrics including FID and LPIPS. Luigi Sigillo, Renato Giamba, Danilo Comminiello |
Pattern Recognit. Lett. | 3 |
| 2025 | Guess What I Think: Streamlined EEG-to-Image Generation with Latent Diffusion ModelsabstractGenerating images from brain waves is gaining increasing attention due to its potential to advance brain-computer interface (BCI) systems by understanding how brain signals encode visual cues. Most of the literature has focused on fMRI-to-Image tasks as fMRI is characterized by high spatial resolution. However, fMRI is an expensive neuroimaging modality and does not allow for real-time BCI. On the other hand, electroencephalography (EEG) is a low-cost, non-invasive, and portable neuroimaging technique, making it an attractive option for future real-time applications. Nevertheless, EEG presents inherent challenges due to its low spatial resolution and susceptibility to noise and artifacts, which makes generating images from EEG more difficult. In this paper, we address these problems with a streamlined framework based on the ControlNet adapter for conditioning a latent diffusion model (LDM) through EEG signals. We conduct experiments and ablation studies on popular benchmarks to demonstrate that the proposed method beats other state-of-the-art models. Unlike these methods, which often require extensive preprocessing, pretraining, different losses, and captioning models, our approach is efficient and straightforward, requiring only minimal preprocessing and a few components. The code is available at https://github.com/LuigiSigillo/GWIT. Eleonora Lopez, Luigi Sigillo, Federica Colonnese, Massimo Panella, Danilo Comminiello |
ICASSP | 5 |
| 2025 | Gramian Multimodal Representation Learning and AlignmentabstractHuman perception integrates multiple modalities—such as vision, hearing, and language—into a unified understanding of the surrounding reality. While recent multimodal models have achieved significant progress by aligning pairs of modalities via contrastive learning, their solutions are unsuitable when scaling to multiple modalities. These models typically align each modality to a designated anchor without ensuring the alignment of all modalities with each other, leading to suboptimal performance in tasks requiring a joint understanding of multiple modalities. In this paper, we structurally rethink the pairwise conventional approach to multimodal learning and we present the novel Gramian Representation Alignment Measure (GRAM), which overcomes the above-mentioned limitations. GRAM learns and then aligns $n$ modalities directly in the higher-dimensional space in which modality embeddings lie by minimizing the Gramian volume of the $k$-dimensional parallelotope spanned by the modality vectors, ensuring the geometric alignment of all modalities simultaneously. GRAM can replace cosine similarity in any downstream method, holding for 2 to $n$ modalities and providing more meaningful alignment with respect to previous similarity measures. The novel GRAM-based contrastive loss function enhances the alignment of multimodal models in the higher-dimensional embedding space, leading to new state-of-the-art performance in downstream tasks such as video-audio-text retrieval and audio-video classification. Giordano Cicchetti, Eleonora Grassucci, Luigi Sigillo, Danilo Comminiello |
ICLR | 4 |
| 2025 | GATSY: Graph Attention Network for Music Artist SimilarityabstractThe artist similarity quest has become a crucial subject in social and scientific contexts, driven by the desire to enhance music discovery according to user preferences. Modern research solutions facilitate music discovery according to user tastes. However, defining similarity among artists remains challenging due to its inherently subjective nature, which can impact recommendation accuracy. This paper introduces GATSY, a novel recommendation system built upon graph attention networks and driven by a clusterized embedding of artists. The proposed framework leverages the graph topology of the input data to achieve outstanding performance results without relying heavily on hand-crafted features. This flexibility allows us to include fictitious artists within a music dataset, facilitating connections between previously unlinked artists and enabling diverse recommendations from various and heterogeneous sources. Experimental results prove the effectiveness of the proposed method with respect to state-of-the-art solutions while maintaining flexibility. The code to reproduce these experiments is available at https://github.com/difra100/GATSY-Music_Artist_Similarity. Andrea Giuseppe Di Francesco, Giuliano Giampietro, Indro Spinelli, Danilo Comminiello |
IJCNN | 4 |
| 2025 | FoleyGRAM: Video-to-Audio Generation with GRAM-Aligned Multimodal EncodersabstractIn this work, we present FoleyGRAM, a novel approach to video-to-audio generation that emphasizes semantic conditioning through the use of aligned multimodal encoders. Building on prior advancements in video-to-audio generation, FoleyGRAM leverages the Gramian Representation Alignment Measure (GRAM) to align embeddings across video, text, and audio modalities, enabling precise semantic control over the audio generation process. The core of FoleyGRAM is a diffusion-based audio synthesis model conditioned on GRAM-aligned embeddings and waveform envelopes, ensuring both semantic richness and temporal alignment with the corresponding input video. We evaluate FoleyGRAM on the Greatest Hits dataset, a standard benchmark for video-to-audio models. Our experiments demonstrate that aligning multimodal encoders using GRAM enhances the system’s ability to semantically align generated audio with video content, advancing the state of the art in video-to-audio synthesis. Riccardo F. Gramaccioni, Christian Marinoni, Eleonora Grassucci, Giordano Cicchetti, Aurelio Uncini, Danilo Comminiello |
IJCNN | 6 |
| 2025 | StereoSync: Spatially-Aware Stereo Audio Generation from Video
Christian Marinoni, Riccardo F. Gramaccioni, Kazuki Shimada, Takashi Shibuya 0001, Yuki Mitsufuji, Danilo Comminiello |
IJCNN | 6 |
| 2025 | Quaternion Wavelet-Conditioned Diffusion Models for Image Super-ResolutionabstractImage Super-Resolution is a fundamental problem in computer vision with broad applications spacing from medical imaging to satellite analysis. The ability to reconstruct high-resolution images from low-resolution inputs is crucial for enhancing downstream tasks such as object detection and segmentation. While deep learning has significantly advanced SR, achieving high-quality reconstructions with fine-grained details and realistic textures remains challenging, particularly at high upscaling factors. Recent approaches leveraging diffusion models have demonstrated promising results, yet they often struggle to balance perceptual quality with structural fidelity. In this work, we introduce ResQu a novel SR framework that integrates a quaternion wavelet preprocessing framework with latent diffusion models, incorporating a new quaternion wavelet- and time-aware encoder. Unlike prior methods that simply apply wavelet transforms within diffusion models, our approach enhances the conditioning process by exploiting quaternion wavelet embeddings, which are dynamically integrated at different stages of denoising. Furthermore, we also leverage the generative priors of foundation models such as Stable Diffusion. Extensive experiments on domain-specific datasets demonstrate that our method achieves outstanding SR results, outperforming in many cases existing approaches in perceptual quality and standard evaluation metrics. The code will be available after the revision process. Luigi Sigillo, Christian Bianchi, Aurelio Uncini, Danilo Comminiello |
IJCNN | 4 |
| 2025 | Seasonal-Trend Transformer-based Diffusion Model for Non-Intrusive Load MonitoringabstractNon-intrusive load monitoring (NILM) separates household electricity consumption into individual appliance profiles, providing a non-invasive method for energy management. Deep learning has greatly advanced the field of NILM, with Transformer-based architectures playing a relevant role. However, the generalization ability of these models to unseen houses and high-density appliance environments is undermined by the noise and volatility of the mixture of signals from multiple overlapping sources. To face these challenges, we present DifFormer-NILM, a novel framework that integrates a specialized Transformer-based architecture into a Denoising Diffusion Probabilistic Model (DDPM). By formulating energy disaggregation as a conditional generative task, this work represents an innovative method of exploiting Transformers and diffusion models to increase overall robustness. In particular, the proposed network introduces Trend and Fourier Synthetic Layers for effective seasonal-trend decomposition, which reveal crucial in capturing both periodic patterns and short-term fluctuations. Tested on real-world use cases from the UKDALE dataset, DifFormer-NILM outperforms Transformer baselines in household scenarios with a high density of appliances and when trained with limited data available. Francesco Sudoso, Redemptor Laceda Taloma, Patrizio Pisani, Danilo Comminiello |
IJCNN | 4 |
| 2025 | A Deep Cascade Framework for Non-Intrusive Power Disaggregation in Solar-Powered HouseholdsabstractInverter-Based Resources are commonly installed behind the customer meters. Thus, non-intrusive power monitoring systems must handle power signals of different natures measured at the main meter and estimate power generation to ensure the observability of the power grid. This paper proposes a non-intrusive disaggregation approach that includes photovoltaic power production with load monitoring. The approach is based on an innovative cascade learning framework that exploits the solar power estimate to simplify the load monitoring task, thereby improving the overall disaggregation performance. Compared to five state-of-the-art models, our method achieves the lowest disaggregation error on two different real-world public datasets, with improvements of 20.86% and 8.67% over the runner-up benchmark. The code to reproduce the method is available on GitHub1. Giulia Tanoni, Redemptor Laceda Taloma, Emanuele Principi, Danilo Comminiello, Stefano Squartini |
ISCAS | 4 |
| 2025 | A TRIANGLE Enables Multimodal Alignment Beyond Cosine SimilarityabstractMultimodal learning plays a pivotal role in advancing artificial intelligence systems by incorporating information from multiple modalities to build a more comprehensive representation. Despite its importance, current state-of-the-art models still suffer from severe limitations that prevent the successful development of a fully multimodal model. Such methods do not provide indicators that all the involved modalities are effectively aligned. As a result, a set of modalities may not be aligned, undermining the effectiveness of the model in downstream tasks where multiple modalities should provide additional information that the model fails to exploit.
In this paper, we present TRIANGLE: TRI-modAl Neural Geometric LEarning, the novel proposed similarity measure that is directly computed in the higher-dimensional space spanned by the modality embeddings. TRIANGLE improves the joint alignment of three modalities via a triangle‑area similarity, avoiding additional fusion layers.
When incorporated in contrastive losses replacing cosine similarity, TRIANGLE significantly boosts the performance of multimodal modeling, while yielding interpretable alignment rationales. Extensive evaluation in three-modal tasks such as video-text and audio-text retrieval or audio-video classification, demonstrates that TRIANGLE achieves state-of-the-art results across different datasets improving the performance of cosine-based methods up to 9 points of Recall@1. Giordano Cicchetti, Eleonora Grassucci, Danilo Comminiello |
NeurIPS | 3 |
| 2025 | Generalizing medical image representations via quaternion wavelet networksabstractNeural network generalizability is becoming a broad research field due to the increasing availability of datasets from different sources and for various tasks. This issue is even wider when processing medical data, where a lack of methodological standards causes large variations being provided by different imaging centers or acquired with various devices and cofactors. To overcome these limitations, we introduce a novel, generalizable, data- and task-agnostic framework able to extract salient features from medical images. The proposed quaternion wavelet network (QUAVE) can be easily integrated with any pre-existing medical image analysis or synthesis task, and it can be involved with real, quaternion, or hypercomplex-valued models, generalizing their adoption to single-channel data. QUAVE first extracts different sub-bands through the quaternion wavelet transform, resulting in both low-frequency/approximation bands and high-frequency/fine-grained features. Then, it weighs the most representative set of sub-bands to be involved as input to any other neural model for image processing, replacing standard data samples. We conduct an extensive experimental evaluation comprising different datasets, diverse image analysis, and synthesis tasks including reconstruction, segmentation, and modality translation. We also evaluate QUAVE in combination with both real and quaternion-valued models. Results demonstrate the effectiveness and the generalizability of the proposed framework that improves network performance while being flexible to be adopted in manifold scenarios and robust to domain shifts. The full code is available at: https://github.com/ispamm/QWT. Luigi Sigillo, Eleonora Grassucci, Aurelio Uncini, Danilo Comminiello |
Neurocomputing | 4 |
| 2024 | Syncfusion: Multimodal Onset-Synchronized Video-to-Audio Foley SynthesisabstractSound design involves creatively selecting, recording, and editing sound effects for various media like cinema, video games, and virtual/augmented reality. One of the most time-consuming steps when designing sound is synchronizing audio with video. In some cases, environmental recordings from video shoots are available, which can aid in the process. However, in video games and animations, no reference audio exists, requiring manual annotation of event timings from the video. We propose a system to extract repetitive actions onsets from a video, which are then used - in conjunction with audio or textual embeddings - to condition a diffusion model trained to generate a new synchronized sound effects audio track. In this way, we leave complete creative control to the sound designer while removing the burden of synchronization with video. Furthermore, editing the onset track or changing the conditioning embedding requires much less effort than editing the audio track itself, simplifying the sonification process. We provide sound examples, source code, and pretrained models to faciliate reproducibility1. Marco Comunità, Riccardo F. Gramaccioni, Emilian Postolache, Emanuele Rodolà, Danilo Comminiello, Joshua D. Reiss |
ICASSP | 5 |
| 2024 | Diffusion Models for Audio Semantic CommunicationabstractDirectly sending audio signals from a transmitter to a receiver across a noisy channel may absorb consistent bandwidth and be prone to errors when trying to recover the transmitted bits. On the contrary, the recent semantic communication approach proposes to send the semantics and then regenerate semantically consistent content at the receiver without exactly recovering the bitstream. In this paper, we propose a generative audio semantic communication framework that faces the communication problem as an inverse problem, therefore being robust to different corruptions. Our method transmits lower-dimensional representations of the audio signal and of the associated semantics to the receiver, which generates the corresponding signal with a particular focus on its meaning (i.e., the semantics) thanks to the conditional diffusion model at its core. During the generation process, the diffusion model restores the received information from multiple degradations at the same time including corruption noise and missing parts caused by the transmission over the noisy channel. We show that our framework outperforms competitors in a real-world scenario and with different channel conditions. Visit the project page to listen to samples and access code and experimental procedures: https://ispamm.github.io/diffusion-audio-semantic-communication/. Eleonora Grassucci, Christian Marinoni, Andrea Rodriguez, Danilo Comminiello |
ICASSP | 4 |
| 2024 | Enhancing Semantic Communication with Deep Generative Models: An OverviewabstractSemantic communication is poised to play a pivotal role in shaping the landscape of future AI-driven communication systems. Its challenge of extracting semantic information from the original complex content and regenerating semantically consistent data at the receiver, possibly being robust to channel corruptions, can be addressed with deep generative models. This ICASSP special session overview paper discloses the semantic communication challenges from the machine learning perspective and unveils how deep generative models will significantly enhance semantic communication frameworks in dealing with real-world complex data, extracting and exploiting semantic information, and being robust to channel corruptions. Alongside establishing this emerging field, this paper charts novel research pathways for the next generative semantic communication frameworks. Eleonora Grassucci, Yuki Mitsufuji, Danilo Comminiello |
ICASSP | 4 |
| 2024 | Efficient Functional Link Adaptive Filters Based On Nearest Kronecker Product DecompositionabstractFunctional link adaptive filters (FLAFs) utilize expansion blocks to nonlinearly augment the input signal to a higher dimensional space, after which an adaptive weight algorithm is applied. These filters are useful for nonlinear system identification tasks, as they can update a large number of coefficients to effectively model the nonlinear system, even when the degree of nonlinearity is not comprehended in advance. However, in many cases, not all of the weights in the nonlinear and linear filter will significantly contribute to the identified model. This paper introduces a novel class of FLAFs based on the nearest Kronecker product (NKP) decomposition. Utilizing the inherent low-rank nature of weight vectors in many scenarios, our approach aims to improve convergence performance and tracking capabilities compared to traditional FLAFs. Additionally, we address noise mitigation challenges, particularly in nonlinear acoustic echo cancellation scenarios. By incorporating NKP decomposition, our proposed FLAFs offer promising solutions for enhancing adaptability and performance in nonlinear system identification, making them valuable tools in practical applications. Alireza Nezamdoust, Mario Huemer, Aurelio Uncini, Danilo Comminiello |
ICASSP | 4 |
| 2024 | Detecting Audio Deepfakes: Integrating CNN and BiLSTM with Multi-Feature ConcatenationabstractAudio deepfake detection is emerging as a crucial field in digital media, as distinguishing real audio from deepfakes becomes increasingly challenging due to the advancement of deepfake technologies. These methods threaten information authenticity and pose serious security risks. Addressing this challenge, we propose a novel architecture that combines Convolutional Neural Networks (CNN) and Bidirectional Long Short-Term Memory (BiLSTM) for effective deepfake audio detection. Our approach is distinguished by the feature concatenation of a comprehensive set of acoustic features: Mel Frequency Cepstral Coefficients (MFCC), Mel spectrograms, Constant Q Cepstral Coefficients (CQCC), and Constant-Q Transform (CQT) vectors. In the proposed architecture, features processed by a CNN are concatenated into two multi-dimensional features for comprehensive analysis, then analyzed by a BiLSTM network to capture temporal dynamics and contextual dependencies in audio data. This synergistic method ensures an understanding of both spatial and sequential audio characteristics. We validate our model on the ASVSpoof 2019 and FoR datasets, using accuracy and Equal Error Rate (EER) metrics for the evaluation. Taiba Majid Wani, Syed Asif Ahmad Qadri, Danilo Comminiello, Irene Amerini |
IH&MMSec | 3 |
| 2024 | Towards Explaining Hypercomplex Neural NetworksabstractHypercomplex neural networks are gaining increasing interest in the deep learning community. The attention directed towards hypercomplex models originates from several aspects, spanning from purely theoretical and mathematical characteristics to the practical advantage of lightweight models over conventional networks, and their unique properties to capture both global and local relations. In particular, a branch of these architectures, parameterized hypercomplex neural networks (PHNNs), has also gained popularity due to their versatility across a multitude of application domains. Nonetheless, only few attempts have been made to explain or interpret their intricacies. In this paper, we propose inherently interpretable PHNNs and quaternion-like networks, thus without the need for any post-hoc method. To achieve this, we define a type of cosine-similarity transform within the parameterized hypercomplex domain. This PHB-cos transform induces weight alignment with relevant input features and allows to reduce the model into a single linear transform, rendering it directly interpretable. In this work, we start to draw insights into how this unique branch of neural models operates. We observe that hypercomplex networks exhibit a tendency to concentrate on the shape around the main object of interest, in addition to the shape of the object itself. We provide a thorough analysis, studying single neurons of different layers and comparing them against how real-valued networks learn. The code of the paper is available at https://github.com/ispamm/HxAI. Eleonora Lopez, Eleonora Grassucci, Debora Capriotti, Danilo Comminiello |
IJCNN | 4 |
| 2024 | Ship in Sight: Diffusion Models for Ship-Image Super ResolutionabstractIn recent years, remarkable advancements have been achieved in the field of image generation, primarily driven by the escalating demand for high-quality outcomes across various image generation subtasks, such as inpainting, denoising, and super resolution. A major effort is devoted to exploring the application of super-resolution techniques to enhance the quality of low-resolution images. In this context, our method explores in depth the problem of ship image super resolution, which is crucial for coastal and port surveillance. We investigate the opportunity given by the growing interest in text-to-image diffusion models, taking advantage of the prior knowledge that such foundation models have already learned. In particular, we present a diffusion-model-based architecture that leverages text conditioning during training while being class-aware, to best preserve the crucial details of the ships during the generation of the super-resoluted image. Since the specificity of this task and the scarcity availability of off-the-shelf data, we also introduce a large labeled ship dataset scraped from online ship images, mostly from ShipSpotting1website. Our method achieves more robust results than other deep learning models previously employed for super resolution, as proven by the multiple experiments performed. Moreover, we investigate how this model can benefit downstream tasks, such as classification and object detection, thus emphasizing practical implementation in a real-world scenario. Experimental results show flexibility, reliability, and impressive performance of the proposed framework over state-of-the-art methods for different tasks. The code is available at: https://github.com/LuigiSigillo/ShipinSight Luigi Sigillo, Riccardo F. Gramaccioni, Alessandro Nicolosi, Danilo Comminiello |
IJCNN | 4 |
| 2024 | Semantic Communication Challenges: Understanding Dos and Avoiding Don'tsabstractSemantic communication, emerging as a promising paradigm for data transmission, offers an innovative departure from the constraints of Shannon theory, heralding significant advancements in future communication technologies. Despite the proliferation of proposed approaches, there are still numerous challenges. In this paper, we review current semantic communication methodologies and shed light on pivotal issues and addressing certain discrepancies that exist within the field. By elucidating both dos and don'ts, we aim to provide valuable insights into the emerging landscape of semantic communication. Jinho Choi 0001, Jihong Park, Eleonora Grassucci, Danilo Comminiello |
VTC Spring | 4 |
| 2024 | Attention-map augmentation for hypercomplex breast cancer classificationabstractBreast cancer is the most widespread neoplasm among women and early detection of this disease is critical. Deep learning techniques have become of great interest to improve diagnostic performance. However, distinguishing between malignant and benign masses in whole mammograms poses a challenge, as they appear nearly identical to an untrained eye, and the region of interest (ROI) constitutes only a small fraction of the entire image. In this paper, we propose a framework, parameterized hypercomplex attention maps (PHAM), to overcome these problems. Specifically, we deploy an augmentation step based on computing attention maps. Then, the attention maps are used to condition the classification step by constructing a multi-dimensional input comprised of the original breast cancer image and the corresponding attention map. In this step, a parameterized hypercomplex neural network (PHNN) is employed to perform breast cancer classification. The framework offers two main advantages. First, attention maps provide critical information regarding the ROI and allow the neural model to concentrate on it. Second, the hypercomplex architecture has the ability to model local relations between input dimensions thanks to hypercomplex algebra rules, thus properly exploiting the information provided by the attention map. We demonstrate the efficacy of the proposed framework on both mammography images as well as histopathological ones. We surpass attention-based state-of-the-art networks and the real-valued counterpart of our approach. The code of our work is available at https://github.com/elelo22/AttentionBCS. Eleonora Lopez, Filippo Betello, Federico Carmignani, Eleonora Grassucci, Danilo Comminiello |
Pattern Recognit. Lett. | 5 |
| 2024 | PHNNs: Lightweight Neural Networks via Parameterized Hypercomplex ConvolutionsabstractHypercomplex neural networks have proven to reduce the overall number of parameters while ensuring valuable performance by leveraging the properties of Clifford algebras. Recently, hypercomplex linear layers have been further improved by involving efficient parameterized Kronecker products. In this article, we define the parameterization of hypercomplex convolutional layers and introduce the family of parameterized hypercomplex neural networks (PHNNs) that are lightweight and efficient large-scale models. Our method grasps the convolution rules and the filter organization directly from data without requiring a rigidly predefined domain structure to follow. PHNNs are flexible to operate in any user-defined or tuned domain, from 1-D to [Formula: see text] regardless of whether the algebra rules are preset. Such a malleability allows processing multidimensional inputs in their natural domain without annexing further dimensions, as done, instead, in quaternion neural networks (QNNs) for 3-D inputs like color images. As a result, the proposed family of PHNNs operates with 1/n free parameters as regards its analog in the real domain. We demonstrate the versatility of this approach to multiple domains of application by performing experiments on various image datasets and audio datasets in which our method outperforms real and quaternion-valued counterparts. Full code is available at: https://github.com/eleGAN23/HyperNets. Eleonora Grassucci, Aston Zhang, Danilo Comminiello |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Overview of the L3DAS23 Challenge on Audio-Visual Extended RealityabstractThe primary goal of the L3DAS23 Signal Processing Grand Challenge at ICASSP 2023 is to promote and support collaborative research on machine learning for 3D audio signal processing, with a specific emphasis on 3D speech enhancement and 3D Sound Event Localization and Detection in Extended Reality applications. As part of our latest competition, we provide a brand-new dataset, which maintains the same general characteristics of the L3DAS21 and L3DAS22 datasets, but with first-order Ambisonics recordings from multiple reverberant simulated environments. Moreover, we start exploring an audio-visual scenario by providing images of these environments, as perceived by the different microphone positions and orientations. We also propose updated baseline models for both tasks that can now support audio-image couples as input and a supporting API to replicate our results. Finally, we present the results of the participants. Further details about the challenge are available at www.l3das.com/icassp2023. Christian Marinoni, Riccardo F. Gramaccioni, Changan Chen, Aurelio Uncini, Danilo Comminiello |
ICASSP | 5 |
| 2023 | StawGAN: Structural-Aware Generative Adversarial Networks for Infrared Image TranslationabstractThis paper addresses the problem of translating night-time thermal infrared images, which are the most adopted image modalities to analyze night-time scenes, to daytime color images (NTIT2DC), which provide better perceptions of objects. We introduce a novel model that focuses on enhancing the quality of the target generation without merely colorizing it. The proposed structural aware (StawGAN) enables the translation of better-shaped and high-definition objects in the target domain. We test our model on aerial images of the DroneVeichle dataset containing RGB-IR paired images. The proposed approach produces a more accurate translation with respect to other state-of-the-art image translation models. The source code will be available after the revision process. Luigi Sigillo, Eleonora Grassucci, Danilo Comminiello |
ISCAS | 3 |
| 2023 | Dual quaternion ambisonics array for six-degree-of-freedom acoustic representation
Eleonora Grassucci, Gioia Mancini, Christian Brignone, Aurelio Uncini, Danilo Comminiello |
Pattern Recognit. Lett. | 5 |
| 2023 | GROUSE: A Task and Model Agnostic Wavelet- Driven Framework for Medical ImagingabstractIn recent years, deep learning has permeated the field of medical image analysis gaining increasing attention from clinicians. However, medical images always require specific preprocessing that often includes downscaling due to computational constraints. This may cause a crucial loss of information magnified by the fact that the region of interest is usually a tiny portion of the image. To overcome these limitations, we propose GROUSE, a novel and generalizable framework that produces salient features from medical images bygroupingandselectingfrequency sub-bands that provide approximations and fine-grained details useful for building a more complete input representation. The framework provides the most enlightening set of bands by learning their statistical dependency to avoid redundancy and by scoring their informativeness to provide meaningful data. This set of representative features can be fed as input to any neural model, replacing the conventional image input. Our method is task- and model-agnostic, thus it can be generalized to any medical image benchmark, as we extensively demonstrate with different tasks, datasets, and model domains. We show that the proposed framework enhances model performance in every test we conduct without requiring ad-hoc preprocessing or network adjustments. Eleonora Grassucci, Luigi Sigillo, Aurelio Uncini, Danilo Comminiello |
IEEE Signal Process. Lett. | 4 |
| 2023 | Learning Speech Emotion Representations in the Quaternion DomainabstractThe modeling of human emotion expression in speech signals is an important, yet challenging task. The high resource demand of speech emotion recognition models, combined with the general scarcity of emotion-labelled data are obstacles to the development and application of effective solutions in this field. In this paper, we present an approach to jointly circumvent these difficulties. Our method, named RH-emo, is a novel semi-supervised architecture aimed at extracting quaternion embeddings from real-valued monoaural spectrograms, enabling the use of quaternion-valued networks for speech emotion recognition tasks. RH-emo is a hybrid real/quaternion autoencoder network that consists of a real-valued encoder in parallel to a real-valued emotion classifier and a quaternion-valued decoder. On the one hand, the classifier permits to optimization of each latent axis of the embeddings for the classification of a specific emotion-related characteristic: valence, arousal, dominance, and overall emotion. On the other hand, quaternion reconstruction enables the latent dimension to develop intra-channel correlations that are required for an effective representation as a quaternion entity. We test our approach on speech emotion recognition tasks using four popular datasets: IEMOCAP, RAVDESS, EmoDB, and TESS, comparing the performance of three well-established real-valued CNN architectures (AlexNet, ResNet-50, VGG) and their quaternion-valued equivalent fed with the embeddings created with RH-emo. We obtain a consistent improvement in the test accuracy for all datasets, while drastically reducing the resources' demand of models. Moreover, we performed additional experiments and ablation studies that confirm the effectiveness of our approach. The RH-emo repository is available at:https://github.com/ispamm/rhemo. Eric Guizzo, Tillman Weyde, Simone Scardapane, Danilo Comminiello |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2023 | A New Class of Efficient Adaptive Filters for Online Nonlinear ModelingabstractNonlinear models are known to provide excellent performance in real-world applications that often operate in nonideal conditions. However, such applications often require online processing to be performed with limited computational resources. To address this problem, we propose a new class of efficient nonlinear models for online applications. The proposed algorithms are based on linear-in-the-parameters (LIPs) nonlinear filters using functional link expansions. In order to make this class of functional link adaptive filters (FLAFs) efficient, we propose low-complexity expansions and frequency-domain adaptation of the parameters. Among this family of algorithms, we also define the partitioned-block frequency-domain FLAF (FD-FLAF), whose implementation is particularly suitable for online nonlinear modeling problems. We assess and compare FD-FLAFs with different expansions providing the best possible tradeoff between performance and computational complexity. Experimental results prove that the proposed algorithms can be considered as an efficient and effective solution for online applications, such as the acoustic echo cancellation, even in the presence of adverse nonlinear conditions and with limited availability of computational resources. Danilo Comminiello, Alireza Nezamdoust, Simone Scardapane, Michele Scarpiniti, Amir Hussain 0001, Aurelio Uncini |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2022 | L3DAS22 Challenge: Learning 3D Audio Sources in a Real Office EnvironmentabstractThe L3DAS22 Challenge is aimed at encouraging the development of machine learning strategies for 3D speech enhancement and 3D sound localization and detection in office-like environments. This challenge improves and extends the tasks of the L3DAS21 edition1. We generated a new dataset, which maintains the same general characteristics of L3DAS21 datasets, but with an extended number of data points and adding constrains that improve the baseline model’s efficiency and overcome the major difficulties encountered by the participants of the previous challenge. We updated the baseline model of Task 1, using the architecture that ranked first in the previous challenge edition. We wrote a new supporting API, improving its clarity and ease-of-use. In the end, we present and discuss the results submitted by all participants. L3DAS22 Challenge website: www.l3das.com/icassp2022. Eric Guizzo, Christian Marinoni, Marco Pennese, Xinlei Ren, Xiguang Zheng, Bruno S. Masiero, Aurelio Uncini, Danilo Comminiello |
ICASSP | 9 |
| 2022 | Hypercomplex Image- to- Image TranslationabstractImage-to-image translation (I2I) aims at transferring the content representation from an input domain to an output one, bouncing along different target domains. Recent I2I generative models, which gain outstanding results in this task, comprise a set of diverse deep networks each with tens of million parameters. Moreover, images are usually three-dimensional being composed of RGB channels and common neural models do not take dimensions correlation into account, losing beneficial information. In this paper, we propose to leverage hypercomplex algebra properties to define lightweight I2I generative models capable of preserving pre-existing relations among image dimensions, thus exploiting additional input information. On manifold I2I benchmarks, we show how the proposed Quaternion StarGANv2 and parameterized hypercomplex StarGANv2 (PHStarGANv2) reduce parameters and storage memory amount while ensuring high domain translation performance and good image quality as measured by FID and LPIPS scores. Full code is available at https://github.com/ispamm/HI2I. Eleonora Grassucci, Luigi Sigillo, Aurelio Uncini, Danilo Comminiello |
IJCNN | 4 |
| 2022 | CoVal-SGAN: A Complex-Valued Spectral GAN architecture for the effective audio data augmentation in construction sitesabstractGenerative audio data augmentation in a construction site is one of challenging research areas due to the high dissimilarity between work sounds of involved machines and equipment. However, it becomes necessary since the availability of audio data of critical work classes is often rare. Motivated by these considerations and demands, in this paper, we propose a complex-valued GAN architecture working with the audio spectrogram, named CoVal-SGAN, for an effective augmentation of audio data. Specifically, the proposed CoVal-SGAN exploits both the magnitude and phase information to improve the quality of the artificially generated audio signals and increase the overall performance of the underlying classifier. Numerical results, performed on the data recorded in real-world construction sites, along with the comparisons with available state-of-the-art approaches, show the effectiveness of the proposed idea by obtaining an improved accuracy. Michele Scarpiniti, Cristiano Mauri, Danilo Comminiello, Aurelio Uncini, Yong-Cheol Lee |
IJCNN | 3 |
| 2021 | A Quaternion-Valued Variational AutoencoderabstractDeep probabilistic generative models have achieved incredible success in many fields of application. Among such models, variational autoencoders (VAEs) have proved their ability in modeling a generative process by learning a latent representation of the input. In this paper, we propose a novel VAE defined in the quaternion domain, which exploits the properties of quaternion algebra to improve performance while significantly reducing the number of parameters required by the network. The success of the proposed quaternion VAE with respect to traditional VAEs relies on the ability to leverage the internal relations between quaternion-valued input features and on the properties of second-order statistics which allow to define the latent variables in the augmented quaternion domain. In order to show the advantages due to such properties, we define a plain convolutional VAE in the quaternion domain and we evaluate its performance with respect to its real-valued counterpart on the CelebA face dataset. Eleonora Grassucci, Danilo Comminiello, Aurelio Uncini |
ICASSP | 2 |
| 2020 | Differentiable Branching In Deep Networks for Fast InferenceabstractIn this paper, we consider the design of deep neural networks augmented with multiple auxiliary classifiers departing from the main (backbone) network. These classifiers can be used to perform early-exit from the network at various layers, making them convenient for energy-constrained applications such as IoT, embedded devices, or Fog computing. However, designing an optimized early-exit strategy is a difficult task, generally requiring a large amount of manual fine-tuning. In this paper, we propose a way to jointly optimize this strategy together with the branches, providing an end-to-end trainable algorithm for this emerging class of neural networks. We achieve this by replacing the original output of the branches with a 'soft', differentiable approximation. In addition, we also propose a regularization approach to trade-off the computational efficiency of the early-exit strategy with respect to the overall classification accuracy. We evaluate our proposed design approach on a set of image classification benchmarks, showing significant gains in accuracy and inference time. Simone Scardapane, Danilo Comminiello, Michele Scarpiniti, Enzo Baccarelli, Aurelio Uncini |
ICASSP | 2 |
| 2019 | Quaternion Convolutional Neural Networks for Detection and Localization of 3D Sound EventsabstractLearning from data in the quaternion domain enables us to exploit internal dependencies of 4D signals and treating them as a single entity. One of the models that perfectly suits with quaternion-valued data processing is represented by 3D acoustic signals in their spherical harmonics decomposition. In this paper, we address the problem of localizing and detecting sound events in the spatial sound field by using quaternion-valued data processing. In particular, we consider the spherical harmonic components of the signals captured by a first-order ambisonic microphone and process them by using a quaternion convolutional neural network. Experimental results show that the proposed approach exploits the correlated nature of the ambisonic signals, thus improving accuracy results in 3D sound event detection and localization. Danilo Comminiello, Marco Lella, Simone Scardapane, Aurelio Uncini |
ICASSP | 1 |
| 2019 | Frequency-domain Adaptive Filtering: from Real to Hypercomplex Signal ProcessingabstractFrequency-domain adaptive filters (FDAFs) have been widely used over the years, but they are still matter of research due to their powerful capabilities that differentiate them from the whole family of time-domain adaptive filters. This paper aims at providing an overview on FDAFs through a unifying framework that can be used for the derivation of the most popular algorithms of the FDAF family and enables the processing of a wide variety of signals, from real-valued ones to complex- and hypercomplex-valued signals. In particular, we focus on a recent class of FDAFs in the quaternion domain and we show how to derive it from the described framework. Moreover, we evaluate the application of the derived quaternion FDAF to the processing of 3D audio signals. Experimental results show the effectiveness of the proposed adaptive filter in estimating the inverse of a multidimensional acoustic impulse response. Danilo Comminiello, Michele Scarpiniti, Raffaele Parisi, Aurelio Uncini |
ICASSP | 1 |
| 2019 | Widely Linear Kernels for Complex-valued Kernel Activation FunctionsabstractComplex-valued neural networks (CVNNs) have been shown to be powerful nonlinear approximators when the input data can be properly modeled in the complex domain. One of the major challenges in scaling up CVNNs in practice is the design of complex activation functions. Recently, we proposed a novel framework for learning these activation functions neuron-wise in a data-dependent fashion, based on a cheap one-dimensional kernel expansion and the idea of kernel activation functions (KAFs). In this paper we argue that, despite its flexibility, this framework is still limited in the class of functions that can be modeled in the complex domain. We leverage the idea of widely linear complex kernels to extend the formulation, allowing for a richer expressiveness without an increase in the number of adaptable parameters. We test the resulting model on a set of complex-valued image classification benchmarks. Experimental results show that the resulting CVNNs can achieve higher accuracy while at the same time converging faster. Simone Scardapane, Steven Van Vaerenbergh, Danilo Comminiello, Aurelio Uncini |
ICASSP | 3 |
| 2018 | Sparse functional link adaptive filter using an ℓ1-norm regularizationabstractLinear-in-the-parameters nonlinear adaptive filters often show some sparse behavior due to the fact that not all the coefficients are equally useful for the modeling of any nonlinearity. Recently, proportionate algorithms have been proposed to leverage sparsity behaviors in nonlinear filtering. In this paper, we deal with this problem by introducing a proportionate adaptive algorithm based on an ℓ1-norm penalty of the cost function, which regularizes the solution, to be used for a class of nonlinear filters based on functional links. The proposed algorithm stresses the difference between useful and useless functional links for the purpose of nonlinear modeling. Experimental results clearly show faster convergence performance with respect to the standard (i.e., non-regularized) version of the algorithm. Danilo Comminiello, Michele Scarpiniti, Simone Scardapane, Aurelio Uncini |
ISCAS | 1 |
| 2017 | Introducing complex functional link polynomial filtersabstractThe paper introduces a novel class of complex nonlinear filters, the complex functional link polynomial (CFLiP) filters. These filters present many interesting properties. They are a sub-class of linear-in-the-parameter nonlinear filters. They satisfy all the conditions of Stone-Weirstrass theorem and thus are universal approximators for causal, time-invariant, discrete-time, finite-memory, complex, continuous systems defined on a compact domain. The CFLiP basis functions separate the magnitude and phase of the input signal. Moreover, CFLiP filters include many families of nonlinear filters with orthogonal basis functions. It is shown in the experimental results that they are capable of modeling the nonlinearities of high power amplifiers of telecommunication systems with better accuracy than most of the filters currently used for this purpose. Alberto Carini, Danilo Comminiello |
ICASSP | 2 |
| 2017 | Group sparse regularization for deep neural networks
Simone Scardapane, Danilo Comminiello, Amir Hussain 0001, Aurelio Uncini |
Neurocomputing | 2 |
| 2017 | Combined nonlinear filtering architectures involving sparse functional link adaptive filters
Danilo Comminiello, Michele Scarpiniti, Luis Antonio Azpicueta-Ruiz, Jerónimo Arenas-García, Aurelio Uncini |
Signal Process. | 1 |
| 2017 | Frequency domain quaternion adaptive filters: Algorithms and convergence performance
Francesca Ortolani, Danilo Comminiello, Michele Scarpiniti, Aurelio Uncini |
Signal Process. | 2 |
| 2016 | Design of hybrid nonlinear spline adaptive filters for active noise controlabstractIn this paper, we focus on the problem of removing noise in the acoustic domain. To this end, we introduce a class of hybrid nonlinear spline filters, which are designed as a cascade of an adaptive spline function and a single layer adaptive nonlinear network. The adaptive nonlinear networks employed in this work are the functional link network and the even mirror Fourier nonlinear network. Suitable update rules, which not only update the adaptive weights of the nonlinear networks, but also introduce adaptability in the developed spline function are derived. The proposed nonlinear filters have been successfully applied to nonlinear system identification as well as nonlinear active noise control. The new filters have been shown to outperform other popular nonlinear filters. Vinal Patel, Danilo Comminiello, Michele Scarpiniti, Nithin V. George, Aurelio Uncini |
IJCNN | 2 |
| 2016 | A semi-supervised random vector functional-link network based on the transductive framework
Simone Scardapane, Danilo Comminiello, Michele Scarpiniti, Aurelio Uncini |
Inf. Sci. | 2 |
| 2015 | Functional link expansions for nonlinear modeling of audio and speech signalsabstractNonlinear distortions pose a serious problem for the quality preservation of audio and speech signals. To address this problem, such signals are processed by nonlinear models. Functional link adaptive filter (FLAF) is a linear-in-the-parameter nonlinear model, whose nonlinear transformation of the input is characterized by a basis function expansion, satisfying the universal approximation properties. Since the expansion type affects the nonlinear modeling according to the nature of the input signal, in this paper we investigate the FLAF modeling performance involving the most popular functional expansions when audio and speech signals are processed. A comprehensive analysis is conducted to provide the best suitable solution for the processing of nonlinear signals. Experimental results are assessed also in terms of signal quality and intelligibility. Danilo Comminiello, Simone Scardapane, Michele Scarpiniti, Raffaele Parisi, Aurelio Uncini |
IJCNN | 1 |
| 2015 | An interactive optimization procedure for stereophonic acoustic echo cancellation systemsabstractAcoustic echo cancellers are used in teleconferencing systems in order to reduce undesired echoes due to coupling between microphones and loudspeakers. Stereophonic systems provide more realistic experience than single-channel systems, since listeners have spatial information that helps to identify the speaker position. Assuming this scenario, a suitable choice for the system parameters becomes essential to improve the audio reproduction quality. Error-driven optimization strategies are usually used to obtain an optimal system configuration but there is no relationship with the quality desired by the user. In this paper, an interactive evolutionary algorithm is adopted for a stereophonic acoustic echo cancellation system in order to meet subjective specifications in the optimization stage. In this way, the optimal system configuration is derived according to a user-driven approach in order to satisfy the quality requirements demanded by users availing such stereophonic systems. Experimental results prove the effectiveness of the proposed interactive architecture according to both objective and subjective measures. Laura Romoli, Stefania Cecchi, Francesco Piazza, Danilo Comminiello, Michele Scarpiniti, Aurelio Uncini |
IJCNN | 4 |
| 2015 | Improving nonlinear modeling capabilities of functional link adaptive filters
Danilo Comminiello, Michele Scarpiniti, Simone Scardapane, Raffaele Parisi, Aurelio Uncini |
Neural Networks | 1 |
| 2015 | Nonlinear system identification using IIR Spline Adaptive Filters
Michele Scarpiniti, Danilo Comminiello, Raffaele Parisi, Aurelio Uncini |
Signal Process. | 2 |
| 2015 | Intelligent Acoustic Interfaces With Multisensor Acquisition for Immersive ReproductionabstractImmersive speech communication systems have been gaining increasing attention due to their ability to reproduce enhanced acoustic images, and thus achieving good performance in terms of sound quality and accuracy. In this context , a fundamental role is played by intelligent acoustic interfaces (IAIs), which aim at acquiring and/or reproducing desired acoustic information with enhanced perception. The recent widespread availability of multimedia devices, equipped with different kind of sensors, has broadened the range of data processing methods, thus giving a chance for developing advanced IAIs. In this paper, we propose an immersive communication system composed of two IAIs: the first one exploits microphones and cameras, together with a signal processing system, to reduce unwanted noise and enhance the speech quality of the desired information in the transmitting room; the second one is an advanced reproduction system based on a loudspeaker array and on an effective wave field synthesis technique capable of reproducing the spatial perception of the desired speech source in the receiving room. The whole system has been assessed in simulated and real immersive communication scenarios: objective and subjective evaluations have been shown the effectiveness of the proposed system. Danilo Comminiello, Stefania Cecchi, Michele Scarpiniti, Michele Gasparini, Laura Romoli, Francesco Piazza, Aurelio Uncini |
IEEE Trans. Multim. | 1 |
| 2015 | Online Sequential Extreme Learning Machine With KernelsabstractThe extreme learning machine (ELM) was recently proposed as a unifying framework for different families of learning algorithms. The classical ELM model consists of a linear combination of a fixed number of nonlinear expansions of the input vector. Learning in ELM is hence equivalent to finding the optimal weights that minimize the error on a dataset. The update works in batch mode, either with explicit feature mappings or with implicit mappings defined by kernels. Although an online version has been proposed for the former, no work has been done up to this point for the latter, and whether an efficient learning algorithm for online kernel-based ELM exists remains an open problem. By explicating some connections between nonlinear adaptive filtering and ELM theory, in this brief, we present an algorithm for this task. In particular, we propose a straightforward extension of the well-known kernel recursive least-squares, belonging to the kernel adaptive filtering (KAF) family, to the ELM framework. We call the resulting algorithm the kernel online sequential ELM (KOS-ELM). Moreover, we consider two different criteria used in the KAF field to obtain sparse filters and extend them to our context. We show that KOS-ELM, with their integration, can result in a highly efficient algorithm, both in terms of obtained generalization error and training time. Empirical evaluations demonstrate interesting results on some benchmarking datasets. Simone Scardapane, Danilo Comminiello, Michele Scarpiniti, Aurelio Uncini |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2014 | GP-based kernel evolution for L2-Regularization NetworksabstractIn kernel-based learning methods, a crucial design parameter is given by the choice of the kernel function to be used. Although there is, in theory, an infinite range of potential candidates, a handful of kernels covers the majority of actual applications. Partly, this is due to the difficulty of choosing an optimal kernel function in absence of a-priori information. In this respect, Genetic Programming (GP) techniques have shown interesting capabilities of learning non-trivial kernel functions that outperform commonly used ones. However, experiments have been restricted to the use of Support Vector Machines (SVMs), and have not addressed some problems that are specific to GP implementations, such as diversity maintenance. In these respects, the aim of this paper is twofold. First, we present a customized GP-based kernel search method that we apply using an L2-Regularization Network as the base learning algorithm. Second, we investigate the problem of diversity maintenance in the context of kernel evolution, and test an adaptive criterion for maintaining it in our algorithm. For the former point, experiments show a gain in accuracy for our method against fine-tuned standard kernels. For the latter, we show that diversity is decreasing critically fast during the GP iterations, but this decrease does not seems to affect performance of the algorithm. Simone Scardapane, Danilo Comminiello, Michele Scarpiniti, Aurelio Uncini |
IEEE Congress on Evolutionary Computation | 2 |
| 2014 | Advanced intelligent acoustic interfaces for multichannel audio reproductionabstractNowadays, there is a large interest towards multimedia audio systems as a consequence of the development of advanced digital signal processing techniques. In particular, immersive speech communication system has gaining increasing attention since they allow to reproduce realistic acoustic image, and thus achieving good performance in terms of sound quality and accuracy. In this scenario a fundamental role is played by intelligent acoustic interfaces which aim at acquiring audio information, processing it, and returning the processed information to the audio rendering system. In this paper, an effective intelligent acoustic interface composed of a microphone array and a signal processing system capable to enhance the intelligibility of the desired information of the transmitting room is proposed, combined with an advanced reproduction system based on an efficient application of a wave field synthesis technique capable to reproduce an immersive scenario in the receiving room. The whole system has been assessed within a speech communication application involving a moving desired source in a real scenario: objective and subjective evaluation has been reported in order to show the overall system performance. Danilo Comminiello, Stefania Cecchi, Michele Gasparini, Michele Scarpiniti, Aurelio Uncini, Francesco Piazza |
IJCNN | 1 |
| 2014 | An effective criterion for pruning reservoir's connections in Echo State NetworksabstractEcho State Networks (ESNs) were introduced to simplify the design and training of Recurrent Neural Networks (RNNs), by explicitly subdividing the recurrent part of the network, the reservoir, from the non-recurrent part. A standard practice in this context is the random initialization of the reservoir, subject to few loose constraints. Although this results in a simple-to-solve optimization problem, it is in general suboptimal, and several additional criteria have been devised to improve its design. In this paper we provide an effective algorithm for removing redundant connections inside the reservoir during training. The algorithm is based on the correlation of the states of the nodes, hence it depends only on the input signal, it is efficient to implement, and it is also local. By applying it, we can obtain an optimally sparse reservoir in a robust way. We present the performance of our algorithm on two synthetic datasets, which show its effectiveness in terms of better generalization and lower computational complexity of the resulting ESN. This behavior is also investigated for increasing levels of memory and non-linearity required by the task. Simone Scardapane, Gabriele Nocco, Danilo Comminiello, Michele Scarpiniti, Aurelio Uncini |
IJCNN | 3 |
| 2014 | Hammerstein uniform cubic spline adaptive filters: Learning and convergence properties
Michele Scarpiniti, Danilo Comminiello, Raffaele Parisi, Aurelio Uncini |
Signal Process. | 2 |
| 2014 | Nonlinear Acoustic Echo Cancellation Based on Sparse Functional Link RepresentationsabstractRecently, a new class of nonlinear adaptive filtering architectures has been introduced based on the functional link adaptive filter (FLAF) model. Here we focus specifically on the split FLAF (SFLAF) architecture, which separates the adaptation of linear and nonlinear coefficients using two different adaptive filters in parallel. This property makes the SFLAF a well-suited method for problems like nonlinear acoustic echo cancellation (NAEC), in which the separation of filtering tasks brings some performance improvement. Although flexibility is one of the main features of the SFLAF, some problem may occur when the nonlinearity degree of the input signal is not known a priori. This implies a non-optimal choice of the number of coefficients to be adapted in the nonlinear path of the SFLAF. In order to tackle this problem, we propose a proportionate FLAF (PFLAF), which is based on sparse representations of functional links, thus giving less importance to those coefficients that do not actively contribute to the nonlinear modeling. Experimental results show that the proposed PFLAF achieves performance improvement with respect to the SFLAF in several nonlinear scenarios. Danilo Comminiello, Michele Scarpiniti, Luis Antonio Azpicueta-Ruiz, Jerónimo Arenas-García, Aurelio Uncini |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2013 | Combined adaptive beamforming schemes for nonstationary interfering noise reduction
Danilo Comminiello, Michele Scarpiniti, Raffaele Parisi, Aurelio Uncini |
Signal Process. | 1 |
| 2013 | Nonlinear spline adaptive filtering
Michele Scarpiniti, Danilo Comminiello, Raffaele Parisi, Aurelio Uncini |
Signal Process. | 2 |
| 2013 | Functional Link Adaptive Filters for Nonlinear Acoustic Echo CancellationabstractThis paper introduces a new class of nonlinear adaptive filters, whose structure is based on Hammerstein model. Such filters derive from the functional link adaptive filter (FLAF) model, defined by a nonlinear input expansion, which enhances the representation of the input signal through a projection in a higher dimensional space, and a subsequent adaptive filtering. In particular, two robust FLAF-based architectures are proposed and designed ad hoc to tackle nonlinearities in acoustic echo cancellation (AEC). The simplest architecture is the split FLAF, which separates the adaptation of linear and nonlinear elements using two different adaptive filters in parallel. In this way, the architecture can accomplish distinctly at best the linear and the nonlinear modeling. Moreover, in order to give robustness against different degrees of nonlinearity, a collaborative FLAF is proposed based on the adaptive combination of filters. Such architecture allows to achieve the best performance regardless of the nonlinearity degree in the echo path. Experimental results show the effectiveness of the proposed FLAF-based architectures in nonlinear AEC scenarios, thus resulting an important solution to the modeling of nonlinear acoustic channels. Danilo Comminiello, Michele Scarpiniti, Luis Antonio Azpicueta-Ruiz, Jerónimo Arenas-García, Aurelio Uncini |
IEEE Trans. Speech Audio Process. | 1 |
| 2010 | A novel affine projection algorithm for superdirective microphone array beamformingabstractThis paper describes a new adaptive algorithm and assesses its effectiveness within speech enhancement applications. The proposed variable step size block exact APA (VSS-BEAPA) filtering is based on the affine projection algorithm (APA) and introduces a block processing with a variable step size that allows to consider under-modeling scenarios. The algorithm shows improved convergence performance and computational efficiency and its robustness is proved in a typical context of a hands-free teleconferencing application in a noisy environment. The experiments show that a microphone array system joined with VSS-BEAPA filtering is capable of both decreasing the noise level and enhancing the speech signal quality. Danilo Comminiello, Michele Scarpiniti, Raffaele Parisi, Aurelio Uncini |
ISCAS | 1 |