EDBT 2026 Demo / reviewers in the wild / expert
Stefano Squartini
dblp:62/6686
· DBLP profile ↗
90ranked-venue papers
9as first author
20since 2021 · last 2026
0000-0001-9374-0128ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 62 · 5 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 3 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 3 since 2021Systems, architecture and hardware · 4 · 1 first-author · 1 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Experimental Assessment of an Immersive Low-Latency Audio System Designed for Remote Therapeutic Intervention
E. Bonifazi, Giacomo Cucchieri, Elisa Felici, M. Osimani, Lorena Rossi, Giuseppe Bergamino, Valeria Bruschi, Stefania Cecchi, Michael Fioretti, Leonardo Gabrielli, Ilaria Marcantoni, Susanna Spinsante, Stefano Squartini, Alessandro Terenzi |
ICT4AWE | 13 |
| 2026 | Recent trends in distant conversational speech recognition: A review of CHiME-7 and 8 DASR challenges
Samuele Cornell, Christoph Böddeker, Taejin Park, He Huang 0012, Desh Raj, Matthew Wiesner, Yoshiki Masuyama, Xuankai Chang, Zhongqiu Wang 0001, Stefano Squartini, L. Paola García-Perera, Shinji Watanabe 0001 |
Comput. Speech Lang. | 10 |
| 2025 | Dysarthric Speech Classification: A Comparative Analysis of Decision-Support MethodsabstractThis study compares data-driven and non-data-driven methods for classifying dysarthric speech employing the UA-Speech dataset. Data-driven deep-learning architectures, including MobileNetV3, ResNet50, and ResNet152, are evaluated alongside non-data-driven approaches such as pre-trained Automatic Speech Recognition (ASR) based methods and DSP-based methods. The findings reveal that overlapping patient data between training and testing sets leads to inflated performance metrics. By enforcing patient-separated evaluations, the study highlights reduced model accuracy, emphasizing the need for methodologies that generalize effectively to unseen data. Among the tested approaches, the ASR-based approach demonstrates the highest potential, achieving strong correlation with human intelligibility assessments and offering promising clinical applicability. This analysis underscores the importance of robust evaluation protocols and highlights the potential of pre-trained ASR models for reliable dysarthric speech assessment. Davide Lillini, Carlo Aironi, Lucia Migliorelli, Leonardo Gabrielli, Stefano Squartini |
IJCNN | 5 |
| 2025 | Contextual Pooling for Multiple-Instance Learning-based Non-Intrusive Load MonitoringabstractUnderstanding detailed energy consumption patterns is crucial for optimizing resource allocation in microgrids, which often incorporate renewable energy sources and stationary batteries. In this context, this paper addresses the Non-intrusive Load Monitoring problem of appliance-state classification, providing insights into appliance-specific usage by using only the aggregate active power consumed in a building. Deep Neural Networks represent the state of the art in this field, but they require a significant amount of data for training. Consequently, in recent years, the research community has explored weakly supervised approaches. Although weakly supervised methods have proven effective, they rely on the use of pooling functions to calculate bag-level predictions from instance-level ones. These functions should be intrinsically related to the dynamics of the input signal, and while several alternatives have been proposed in the literature, none have specifically addressed the characteristics of appliance power profiles. This paper fills this gap by introducing a new pooling function, Contextual Pooling, that operates on multiple adjacent instance-level predictions to better capture the dynamics of appliance activations. This approach has been evaluated against six different methods on the UK-DALE and REFIT datasets. The results indicate that, on average, Contextual Pooling improves performance by at least 0.9 percentage points on UK-DALE and by 0.7 percentage points on REFIT compared to benchmark methods. McNemar’s test confirms the statistical significance of these improvements. Giulia Tanoni, Emanuele Principi, Paolo Vitulli, Luigi Mandolini, Stefano Squartini |
IJCNN | 5 |
| 2025 | A Deep Cascade Framework for Non-Intrusive Power Disaggregation in Solar-Powered HouseholdsabstractInverter-Based Resources are commonly installed behind the customer meters. Thus, non-intrusive power monitoring systems must handle power signals of different natures measured at the main meter and estimate power generation to ensure the observability of the power grid. This paper proposes a non-intrusive disaggregation approach that includes photovoltaic power production with load monitoring. The approach is based on an innovative cascade learning framework that exploits the solar power estimate to simplify the load monitoring task, thereby improving the overall disaggregation performance. Compared to five state-of-the-art models, our method achieves the lowest disaggregation error on two different real-world public datasets, with improvements of 20.86% and 8.67% over the runner-up benchmark. The code to reproduce the method is available on GitHub1. Giulia Tanoni, Redemptor Laceda Taloma, Emanuele Principi, Danilo Comminiello, Stefano Squartini |
ISCAS | 5 |
| 2025 | Interpretability and reliability-driven knowledge distillation for non-intrusive load monitoring on the edgeabstractThe deployment of deep neural networks (DNNs) on resource-constrained edge devices necessitates efficient, low-complexity algorithms. Knowledge distillation (KD) addresses this through a student-teacher paradigm, transferring knowledge from complex teacher models to simpler student models. Current KD methods often optimize student performance without adequately addressing the reliability and interpretability of transferred knowledge, thus presenting challenges in maintaining both robustness and decision transparency. This paper introduces an Interpretability and Reliability-driven Knowledge Distillation (IR-KD) framework that enhances teacher model interpretability through perception-aligned gradients while leveraging hidden information from weak labels to optimize knowledge transfer. Our approach ensures compressed models remain computationally efficient while improving interpretability, which is essential for trustworthy edge AI deployment. We demonstrate improved predictive performance and model interpretability in non-intrusive load monitoring (NILM) applications as a case study. Quantitative explainability metrics confirm that perception-aligned gradients provide more faithful explanations, validating our approach’s effectiveness in developing reliable and transparent edge AI systems. Djordje Batic, Giulia Tanoni, Emanuele Principi, Lina Stankovic, Vladimir Stankovic 0001, Stefano Squartini |
Expert Syst. Appl. | 6 |
| 2024 | One Model to Rule Them All ? Towards End-to-End Joint Speaker Diarization and Speech RecognitionabstractThis paper presents a novel framework for joint speaker diarization (SD) and automatic speech recognition (ASR), named SLIDAR (sliding-window diarization-augmented recognition). SLIDAR can process arbitrary length inputs and can handle any number of speakers, effectively solving "who spoke what, when" concurrently. SLIDAR leverages a sliding window approach and consists of an end-to-end diarization-augmented speech transcription (E2E DAST) model which provides, locally, for each window: transcripts, diarization and speaker embeddings. The E2E DAST model is based on an encoder-decoder architecture and leverages recent techniques such as serialized output training and "Whisper-style" prompting. The local outputs are then combined to get the final SD+ASR result by clustering the speaker embeddings to get global speaker identities. Experiments performed on monaural recordings from the AMI corpus confirm the effectiveness of the method in both close-talk and far-field speech scenarios. Samuele Cornell, Jee-Weon Jung, Shinji Watanabe 0001, Stefano Squartini |
ICASSP | 4 |
| 2024 | A Graph-Based Neural Approach to Linear Sum Assignment ProblemsabstractLinear assignment problems are well-known combinatorial optimization problems involving domains such as logistics, robotics and telecommunications. In general, obtaining an optimal solution to such problems is computationally infeasible even in small settings, so heuristic algorithms are often used to find near-optimal solutions. In order to attain the right assignment permutation, this study investigates a general-purpose learning strategy that uses a bipartite graph to describe the problem structure and a message passing Graph Neural Network (GNN) model to learn the correct mapping. Comparing the proposed structure with two existing DNN solutions, simulation results show that the proposed approach significantly improves classification accuracy, proving to be very efficient in terms of processing time and memory requirements, due to its inherent parameter sharing capability. Among the many practical uses that require solving allocation problems in everyday scenarios, we decided to apply the proposed approach to address the scheduling of electric smart meters access within an electricity distribution smart grid infrastructure, since near-real-time energy monitoring is a key element of the green transition that has become increasingly important in recent times. The results obtained show that the proposed graph-based solver, although sub-optimal, exhibits the highest scalability, compared with other state-of-the-art heuristic approaches. To foster the reproducibility of the results, we made the code available at https://github.com/aircarlo/GNN_LSAP. Carlo Aironi, Samuele Cornell, Stefano Squartini |
Int. J. Neural Syst. | 3 |
| 2024 | End-to-end integration of speech separation and voice activity detection for low-latency diarization of telephone conversations
Giovanni Morrone, Samuele Cornell, Luca Serafini, Enrico Zovato, Alessio Brutti, Stefano Squartini |
Speech Commun. | 6 |
| 2024 | Knowledge Distillation for Scalable Nonintrusive Load MonitoringabstractSmart meters allow the grid to interface with individual buildings and extract detailed consumption information using nonintrusive load monitoring (NILM) algorithms applied to the acquired data. Deep neural networks, which represent the state of the art for NILM, are affected by scalability issues since they require high computational and memory resources, and by reduced performance when training and target domains mismatched. This article proposes a knowledge distillation approach for NILM, in particular for multilabel appliance classification, to reduce model complexity and improve generalization on unseen data domains. The approach uses weak supervision to reduce labeling effort, which is useful in practical scenarios. Experiments, conducted on U.K.-DALE and REFIT datasets, demonstrated that a low-complexity network can be obtained for deployment on edge devices while maintaining high performance on unseen data domains. The proposed approach outperformed benchmark methods in unseen target domains achieving a$F_{1}$-score 0.14 higher than a benchmark model 78 times more complex. Giulia Tanoni, Lina Stankovic, Vladimir Stankovic 0001, Stefano Squartini, Emanuele Principi |
IEEE Trans. Ind. Informatics | 4 |
| 2023 | Multi-Channel Speaker Extraction with Adversarial Training: The Wavlab Submission to The Clarity ICASSP 2023 Grand ChallengeabstractIn this work we detail our submission to the Clarity ICASSP 2023 grand challenge, in which participants have to develop a strong target speech enhancement system for hearing-aid (HA) devices in noisy-reverberant environments. Our system builds on our previous submission at the Second Clarity Enhancement Challenge (CEC2): iNeuBe-X, which consists in an iterative neural/conventional beamforming enhancement pipeline, guided by an enrollment utterance from the target speaker. This model, which won by a large margin the CEC2, is an extension of the state-of-the-art TF-GridNet model for multi-channel, streamable target-speaker speech enhancement. Here, this approach is extended and further improved by leveraging generative adversarial training, which we show proves especially useful when the training data is limited. Using only the official 6k training scenes data, our best model achieves 0.80 hearing-aid speech perception index (HASPI) and 0.41 hearing-aid speech quality index (HASQI) scores on the synthetic evaluation set. However, our model generalized poorly on the semi-real evaluation set. This highlights the fact that our community should focus more on real-world evaluation and less on fully synthetic datasets. Samuele Cornell, Zhongqiu Wang 0001, Yoshiki Masuyama, Shinji Watanabe 0001, Manuel Pariente, Nobutaka Ono, Stefano Squartini |
ICASSP | 7 |
| 2023 | An experimental review of speaker diarization methods with application to two-speaker conversational telephone speech recordingsabstractWe performed an experimental review of current diarization systems for the conversational telephone speech (CTS) domain. In detail, we considered a total of eight different algorithms belonging to clustering-based, end-to-end neural diarization (EEND), and speech separation guided diarization (SSGD) paradigms. We studied the inference-time computational requirements and diarization accuracy on four CTS datasets with different characteristics and languages. We found that, among all methods considered, EEND-vector clustering (EEND-VC) offers the best trade-off in terms of computing requirements and performance. More in general, EEND models have been found to be lighter and faster in inference compared to clustering-based methods. However, they also require a large amount of diarization-oriented annotated data. In particular EEND-VC performance in our experiments degraded when the dataset size was reduced, whereas self-attentive EEND (SA-EEND) was less affected. We also found that SA-EEND gives less consistent results among all the datasets compared to EEND-VC, with its performance degrading on long conversations with high speech sparsity. Clustering-based diarization systems, and in particular VBx, instead have more consistent performance compared to SA-EEND but are outperformed by EEND-VC. The gap with respect to this latter is reduced when overlap-aware clustering methods are considered. SSGD is the most computationally demanding method, but it could be convenient if speech recognition has to be performed. Its performance is close to SA-EEND but degrades significantly when the training and inference data characteristics are less matched. Luca Serafini, Samuele Cornell, Giovanni Morrone, Enrico Zovato, Alessio Brutti, Stefano Squartini |
Comput. Speech Lang. | 6 |
| 2022 | Learning Filterbanks for End-to-End Acoustic BeamformingabstractRecent work on monaural source separation has shown that performance can be increased by using fully learned filterbanks with short windows. On the other hand it is widely known that, for conventional beamforming techniques, performance increases with long analysis windows. This applies also to most hybrid neural beamforming methods which rely on a deep neural network (DNN) to estimate the spatial covariance matrices. In this work we try to bridge the gap between these two worlds and explore fully end-to-end hybrid neural beamforming in which, instead of using the Short-Time-Fourier Transform, also the analysis and synthesis filterbanks are learnt jointly with the DNN. In detail, we explore two different types of learned filterbanks: fully learned and analytic. We perform a detailed analysis using the recent Clarity Challenge data and show that by using learnt filterbanks it is possible to surpass oracle-mask based beamforming for short windows. Samuele Cornell, Manuel Pariente, François Grondin, Stefano Squartini |
ICASSP | 4 |
| 2022 | Digital Filters Design for Personal Sound Zones: a Neural ApproachabstractThe creation of Personal Sound Zones (PSZ) is a recent application of digital signal processing that allows differentiating the sound intensity in neighboring regions of space (e.g. a “bright” and a “dark” zone). Given the impulse responses of the environment, digital filters can be designed in order to obtain an attenuation of the signal in the dark zone as a result of the superposition of the filtered IR coming from each loudspeaker. A neural optimization approach was recently shown to enable PSZ by designing digital FIR filters. In this work we propose an improvement of that neural optimization approach using a simpler neural network architecture. Furthermore we extend the method to the design of IIR filters, which is computationally more effective for a real-time implementation. The neural technique is compared with two state-of-the-art methods, analyzing the performance in terms of Acoustic Contrast. Experiments have been performed using a vehicle composed of standard loudspeakers and two speaker arrays, and show that the proposed approach achieves remarkable Acoustic Contrast without sacrificing audio quality. Giovanni Pepe, Leonardo Gabrielli, Stefano Squartini, Carlo Tripodi, Nicolo Strozzi |
IJCNN | 3 |
| 2022 | Low-Latency Speech Separation Guided Diarization for Telephone ConversationsabstractIn this paper, we carry out an analysis on the use of speech separation guided diarization (SSGD) in telephone conversations. SSGD performs diarization by separating the speakers signals and then applying voice activity detection on each estimated speaker signal. In particular, we compare two low-latency speech separation models. Moreover, we show a post-processing algorithm that significantly reduces the false alarm errors of a SSGD pipeline. We perform our experiments on two datasets: Fisher Corpus Part 1 and CALLHOME, evaluating both separation and diarization metrics. Notably, our SSGD DPRNN-based online model achieves 11.1% DER on CALL-HOME, comparable with most state-of-the-art end-to-end neural diarization models despite being trained on an order of magnitude less data and having considerably lower latency, i.e., 0.1 vs. 10 seconds. We also show that the separated signals can be readily fed to a speech recognition back-end with performance close to the oracle source signals. Giovanni Morrone, Samuele Cornell, Desh Raj, Luca Serafini, Enrico Zovato, Alessio Brutti, Stefano Squartini |
SLT | 7 |
| 2022 | Overlapped Speech Detection and speaker counting using distant microphone arrays
Samuele Cornell, Maurizio Omologo, Stefano Squartini, Emmanuel Vincent 0001 |
Comput. Speech Lang. | 3 |
| 2022 | Antiderivative Antialiasing for Arbitrary Waveform GenerationabstractIn the last decades many efficient methods have been proposed to generate waveforms with reduced aliasing, proving this field quite mature. However, the introduction of Antiderivative Antialiasing (AA) methods for the reduction of aliasing in nonlinear discrete-time processing, can shed a new light on bandlimited oscillators and provide a general method for dealing with aliasing in arbitrary waveforms generation. In this work we will first bridge the gap between AA methods and bandlimited waveform generation, and then show the flexibility of the recently introduced AA-IIR method for dealing with both classical synthesizer waveforms and arbitrary wavetables. We will show how AA methods can be used for periodic and non-periodic waveform generation, and provide an innovative AA-IIR method able to compute alias-reduced version of any waveform with arbitrary antialiasing filter order, without the effort of computing ad hoc analytical expressions. Antialiasing performance is compared with other well known methods, showing the effectiveness of the approach. Leonardo Gabrielli, Stefano D'Angelo, Pier Paolo La Pastina, Stefano Squartini |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2022 | Deep Optimization of Parametric IIR Filters for Audio EqualizationabstractThis paper describes a novel Deep Learning method for the design of IIR parametric filters for automatic multipoint audio equalization, that is the task of improving the sound quality of a listening environment at multiple listening points employing multiple loudspeakers. The filters are designed to approximate the inverse of the RIR and achieve almost flat magnitude response. A simple and effective neural architecture, named BiasNet, is proposed to determine the IIR equalizer parameters. This novel architecture is conceived for optimization and, as such, is able to produce optimal IIR equalizer parameters at its output, after training, with no input required. In absence of input, the presence of learnable non-zero bias terms ensures that the network works properly. An output scaling method is used to obtain accurate tuning of the IIR filters center frequency, quality factor and gain. All layers involved in the proposed method are shown to be differentiable, allowing backpropagation to optimize the network weights and achieve, after a number of training iterations, the optimal output according to a given RIR. The parameters are optimized with respect to a loss function based on a spectral distance between the measured and desired magnitude response, and a regularization term is used to keep the same microphone-loudspeaker energy balance after equalization. Two experimental scenarios are employed, a room and a car cabin, with several loudspeakers. The performance of the proposed method improves over the baseline techniques and achieves an almost flat band at a lower computational cost. Giovanni Pepe, Leonardo Gabrielli, Stefano Squartini, Carlo Tripodi, Nicolo Strozzi |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2021 | Learning to Rank Microphones for Distant Speech RecognitionabstractFully exploiting ad-hoc microphone networks for distant speech recognition is still an open issue. Empirical evidence shows that being able to select the best microphone leads to significant improvements in recognition without any additional effort on front-end processing. Current channel selection techniques either rely on signal, decoder or posterior-based features. Signal-based features are inexpensive to compute but do not always correlate with recognition performance. Instead decoder and posterior-based features exhibit better correlation but require substantial computational resources. In this work, we tackle the channel selection problem by proposing MicRank, a learning to rank framework where a neural network is trained to rank the available channels using directly the recognition performance on the training set. The proposed approach is agnostic with respect to the array geometry and type of recognition back-end. We investigate different learning to rank strategies using a synthetic dataset developed on purpose and the CHiME-6 data. Results show that the proposed approach is able to considerably improve over previous selection techniques, reaching comparable and in some instances better performance than oracle signal-based measures. Samuele Cornell, Alessio Brutti, Marco Matassoni, Stefano Squartini |
Interspeech | 4 |
| 2021 | Real-World Anomaly Detection by Using Digital Twin Systems and Weakly Supervised LearningabstractThe continuously growing amount of monitored data in the Industry 4.0 context requires strong and reliable anomaly detection techniques. The advancement of Digital Twin technologies allows for realistic simulations of complex machinery; therefore, it is ideally suited to generate synthetic datasets for the use in anomaly detection approaches when compared to actual measurement data. In this article, we present novel weakly supervised approaches to anomaly detection for industrial settings. The approaches make use of a Digital Twin to generate a training dataset, which simulates the normal operation of the machinery, along with a small set of labeled anomalous measurement from the real machinery. In particular, we introduce a clustering-based approach, called cluster centers (CC), and a neural architecture based on the Siamese Autoencoders (SAE), which are tailored for weakly supervised settings with very few labeled data samples. The performance of the proposed methods is compared against various state-of-the-art anomaly detection algorithms on an application to a real-world dataset from a facility monitoring system, by using a multitude of performance measures. Also, the influence of hyperparameters related to feature extraction and network architecture is investigated. We find that the proposed SAE-based solutions outperform state-of-the-art anomaly detection approaches very robustly for many different hyperparameter settings on all performance measures. Andrea Castellani, Stefano Squartini |
IEEE Trans. Ind. Informatics | 3 |
| 2020 | Emergency Siren Recognition in Urban Scenarios: Synthetic Dataset and Deep Learning Models
Michela Cantarini, Luca Serafini, Leonardo Gabrielli, Emanuele Principi, Stefano Squartini |
ICIC (1) | 5 |
| 2020 | A Novel Adversarial Training Scheme for Deep Neural Network based Speech EnhancementabstractIn this work, we propose a novel representation-learning technique for Deep Learning-based Speech Enhancement algorithms inspired by Domain-Adversarial training. A gradient reversal layer and an additional network are employed, only at training time, to explicitly enforce a representation that is orthogonal to the additive noise in the input signal. We show that such learning scheme, which can be applied easily to most mask-based Deep Neural Network Speech Enhancement approaches, is able to improve the denoising performance when used in conjunction with scale-invariant signal-to-distortion ratio loss and allows to reach state-of-the-art performance with no computational overhead at run-time. In particular, on the commonly used VoiceBank-DEMAND benchmarking dataset, we improve signal-to-distortion ratio and signal-to-noise ratio over the nonadversarial model and CSIG, COVL and CBAK over other, state-of-the art, adversarial training techniques. Samuele Cornell, Emanuele Principi, Stefano Squartini |
IJCNN | 3 |
| 2020 | Who Cried When: Infant Cry Diarization with Dilated Fully-Convolutional Neural NetworksabstractIn this paper, we address the problem of the concurrent detection of multiple infant cries by using microphones located in the cribs of a Neonatal Intensive Care Unit (NICU). We term this task as infant cry diarization in resemblance with the "speaker diarization" task related to the speech signal: instead of determining "who spoke when", here the problem is determining "who cried when". The proposed algorithm consists of a fully-convolutional neural network (Conv-DetNet) that processes simultaneously all the audio signals acquired from the microphone in each crib and detects if the infants cried or not. The neural network takes as input Log-Mel coefficients and it is composed of stacked dilated convolutional blocks with increasing dilation factors. Each block is composed of pointwise and depthwise convolutional layers that replace standard convolutions with a mathematically equivalent but more efficient operation. The architecture has been compared to its single-channel equivalent and to single and multi-channel architectures presented in a previous work, composed of standard convolutional layers and fully-connected layers. The experiments have been conducted on a synthetic dataset that simulates the acoustic environment of the Salesi Hospital NICU located in Ancona (Italy). The results have been evaluated in terms of Area Under Precision-Recall Curve (PRC-AUC) and they showed that the proposed multi-channel Conv-DetNet achieves the highest performance with a PRC-AUC equal to 87.58%, outperforming all the comparative methods. Marco Severini, Emanuele Principi, Samuele Cornell, Leonardo Gabrielli, Stefano Squartini |
IJCNN | 5 |
| 2020 | Detecting and Counting Overlapping Speakers in Distant Speech ScenariosabstractInternational audience Samuele Cornell, Maurizio Omologo, Stefano Squartini, Emmanuel Vincent 0001 |
INTERSPEECH | 3 |
| 2020 | Deep Learning for Individual Listening ZoneabstractA recent trend in car audio systems is the generation of Individual Listening Zones (ILZ), allowing to improve phone call privacy and reduce disturbance to other passengers, without wearing headphones or earpieces. This is generally achieved by using loudspeaker arrays. In this paper, we describe an approach to achieve ILZ exploiting general purpose car loudspeakers and processing the signal through carefully designed Finite Impulse Response (FIR) filters. We propose a deep neural network approach for the design of filters coefficients in order to obtain a so-called bright zone, where the signal is clearly heard, and a dark zone, where the signal is attenuated. Additionally, the frequency response in the bright zone is constrained to be as flat as possible. Numerical experiments were performed taking the impulse responses measured with either one binaural pair or three binaural pairs for each passenger. The results in terms of attenuation and flatness prove the viability of the approach. Giovanni Pepe, Leonardo Gabrielli, Stefano Squartini, Luca Cattani, Carlo Tripodi |
MMSP | 3 |
| 2019 | End-to-end Binaural Sound Localisation from the Raw WaveformabstractA novel end-to-end binaural sound localisation approach is proposed which estimates the azimuth of a sound source directly from the waveform. Instead of employing hand-crafted features commonly employed for binaural sound localisation, such as the interaural time and level difference, our end-to-end system approach uses a convolutional neural network (CNN) to extract specific features from the waveform that are suitable for localisation. Two systems are proposed which differ in the initial frequency analysis stage. The first system is auditory-inspired and makes use of a gammatone filtering layer, while the second system is fully data-driven and exploits a trainable convolutional layer to perform frequency analysis. In both systems, a set of dedicated convolutional kernels are then employed to search for specific localisation cues, which are coupled with a localisation stage using fully connected layers. Localisation experiments using binaural simulation in both anechoic and reverberant environments show that the proposed systems outperform a state-of-the-art deep neural network system. Furthermore, our investigation of the frequency analysis stage in the second system suggests that the CNN is able to exploit different frequency bands for localisation according to the characteristics of the reverberant environment. Paolo Vecchiotti, Ning Ma 0002, Stefano Squartini, Guy J. Brown |
ICASSP | 3 |
| 2019 | Processing Acoustic Data with Siamese Neural Networks for Enhanced Road Roughness ClassificationabstractIn recent years, a lot of effort has been put in vehicle safety systems for manned and unmanned driving. Road conditions are crucial among the factors that influence the choice of the driving style and the safety systems. A few works based the detection of the road condition on acoustic sensors mounted on the vehicle using deep learning techniques. In this work we enhance the state of the art by introducing a Siamese Convolutional Neural Network architecture able to achieve improved results for the classification of the road surface roughness. A new dataset is recorded and the approach is tested, achieving a best overall F1-score of 95.6%, improving by 14% the results of the previous method. Leonardo Gabrielli, Livio Ambrosini, Fabio Vesperini, Valeria Bruschi, Stefano Squartini, Luca Cattani |
IJCNN | 5 |
| 2019 | Detection of activity and position of speakers by using deep neural networks and acoustic data augmentation
Paolo Vecchiotti, Giovanni Pepe, Emanuele Principi, Stefano Squartini |
Expert Syst. Appl. | 4 |
| 2019 | A Multi-Stage Algorithm for Acoustic Physical Model Parameters EstimationabstractOne of the challenges in computational acoustics is the identification of models that can simulate and predict the physical behavior of a system generating an acoustic signal. Whenever such models are used for commercial applications, an additional constraint is the time to market, making automation of the sound design process desirable. In previous works, a computational sound design approach has been proposed for the parameter estimation problem involving timbre matching by deep learning, which was applied to the synthesis of pipe organ tones. In this paper, we refine previous results by introducing the former approach in a multi-stage algorithm that also adds heuristics and a stochastic optimization method operating on perceptually motivated objective cost functions. The optimization method shows to be able to refine the first estimate given by the deep learning approach and substantially improve the objective metrics, with the additional benefit of reducing the sound design process time. Subjective listening tests are also conducted to gather additional insights on the results. Leonardo Gabrielli, Stefano Tomassetti, Stefano Squartini, Carlo Zinato, Stefano Guaiana |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2018 | Collaborative Energy Management in Micro-Grid environments through multi-objective optimizationabstractIn this work, we propose a multi-objective Mixed Integer Linear Programming formulation for addressing the Collaborative Energy Management Problem with the aim of maximizing the net profit of both the Building Manager and all the apartments. The planning horizon is discretized into a finite set of periods, i.e., time intervals. In this way, both the residents' tasks and the storage activity can be scheduled over the time horizon. The decision variables take into account both the energy resources shared among the apartments, i.e., the ones administrated by the Building Manager, and the local energy resources of each apartment. Together with the traditional scheduling constraints, we also impose both time windows and priority conditions. In particular, regarding the former, each task can be scheduled starting from a specific period. While, according to the latter, each task has a list of prior tasks, of the same resident, that have to be scheduled before it. The proposed formulation is evaluated by investigating the management of a block of four apartments. In one scenario, the apartments are considered as independent entities. In the other one, a collaborative management is performed. The performance comparison reveals that the collaborative management can improve the energy sale revenue up to 26%, providing additional profit to be shared among the apartment owners and the Building Manager. Marco Severini, Ornella Pisacane, Marco Fagiani, Stefano Squartini |
IJCNN | 4 |
| 2018 | Exploiting the Reactive Power in Deep Neural Models for Non-Intrusive Load MonitoringabstractNon-intrusive load monitoring (NILM) is defined as the task of retrieving the active power consumption of two or more appliances from information gathered at a single metering point. In this work, the use of the reactive aggregate power as an additional feature to the commonly used active power for deep neural models is proposed. The NILM problem is formulated as a denoising problem, and denoising autoencoder (dAE) neural architectures are used to estimate the appliances individual active power consumption. The proposed approach is evaluated on two public datasets: the Almanac of Minutely Power dataset (AMPds) and the UK Domestic Appliance-Level Electricity (UK-DALE) dataset. In order to better evaluate the generalization capabilities of the algorithm, different testing conditions are considered for the UK-DALE dataset, namely a seen and an unseen scenario. The results show that introducing the reactive power can indeed bring and overall performance increase in all scenarios, ranging from +4.9% to +8.4% of the energy-based F1 score. Michele Valenti, Roberto Bonfigli, Emanuele Principi, Stefano Squartini |
IJCNN | 4 |
| 2018 | Snore Sounds Excitation Localization by Using Scattering Transform and Deep Neural NetworksabstractIn this paper, we propose an algorithm for snoring sounds classification based on Deep Scattering Spectrum (SCAT), Gaussian Mixture Models (GMM) Supervectors and Deep Neural Networks (DNN). The task consists in the identification of the type of snoring among four target classes representing the snore sounds' excitation location, which can be highly useful for a successful medical treatment of the habitual snorer or patient afflicted with Obstructive Sleep Apnea. The SCAT is computed from excerpt the acoustic signals, then a GMM Supervector is calculated by adapting the GMM model of the acoustic space with the Maximum a Posteriori (MAP) algorithm and concatenating the mean values of the Gaussians. Resulting supervectors are used to feed the multiclass DNN classifier. The performance of the algorithm has been assessed on the Munich-Passau Snore Sound Corpus (MPSSC), composed of recordings of Drug-Induced Sleep Endoscopy (DISE) examinations. The results are expressed in terms of Unweighted Average Recall (UAR) and a remarkable improvement with respect to the state-of-the-art performance has been registered, achieving a score up to 67.14% and 67.71% respectively on the devel and test datasets. Fabio Vesperini, Andrea Galli, Leonardo Gabrielli, Emanuele Principi, Stefano Squartini |
IJCNN | 5 |
| 2018 | Localizing speakers in multiple rooms by using Deep Neural NetworksabstractIn the field of human speech capturing systems, a fundamental role is played by the source localization algorithms. In this paper a Speaker Localization algorithm (SLOC) based on Deep Neural Networks (DNN) is evaluated and compared with state-of-the art approaches. The speaker position in the room under analysis is directly determined by the DNN, leading the proposed algorithm to be fully data-driven. Two different neural network architectures are investigated: the Multi Layer Perceptron (MLP) and Convolutional Neural Networks (CNN). GCC-PHAT (Generalized Cross Correlation-PHAse Transform) Patterns, computed from the audio signals captured by the microphone are used as input features for the DNN. In particular, a multi-room case study is dealt with, where the acoustic scene of each room is influenced by sounds emitted in the other rooms. The algorithm is tested by means of the home recorded DIRHA dataset, characterized by multiple wall and ceiling microphone signals for each room. In detail, the focus goes to speaker localization task in two distinct neighboring rooms. As term of comparison, two algorithms proposed in literature for the addressed applicative context are evaluated, the Crosspower Spectrum Phase Speaker Localization (CSP-SLOC) and the Steered Response Power using the Phase Transform speaker localization (SRP-SLOC). Besides providing an extensive analysis of the proposed method, the article shows how DNN-based algorithm significantly outperforms the state-of-the-art approaches evaluated on the DIRHA dataset, providing an average localization error, expressed in terms of Root Mean Square Error (RMSE), equal to 324 mm and 367 mm, respectively, for the Simulated and the Real subsets. Fabio Vesperini, Paolo Vecchiotti, Emanuele Principi, Stefano Squartini, Francesco Piazza |
Comput. Speech Lang. | 4 |
| 2018 | Special Issue on Deep Reinforcement Learning and Adaptive Dynamic ProgrammingabstractThe sixteen papers in this special section focus on deep reinforcement learning and adaptive dynamic programming (deep RL/ADP). Deep RL is able to output control signal directly based on input images, which incorporates both the advantages of the perception of deep learning (DL) and the decision making of RL or adaptive dynamic programming (ADP). This mechanism makes the artificial intelligence much closer to human thinking modes. Deep RL/ADP has achieved remarkable success in terms of theory and applications since it was proposed. Successful applications cover video games, Go, robotics, smart driving, healthcare, and so on. However, it is still an open problem to perform the theoretical analysis on deep RL/ADP, e.g., the convergence, stability, and optimality analyses. The learning efficiency needs to be improved by proposing new algorithms or combined with other methods. More practical demonstrations are encouraged to be presented. Therefore, the aim of this special issue is to call for the most advanced research and state-of-the-art works in the field of deep RL/ADP. Dongbin Zhao, Derong Liu 0001, Frank L. Lewis, José C. Príncipe, Stefano Squartini |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2017 | A new open-source Energy Management framework: Functional description and preliminary resultsabstractIn this paper, a new open-source SW framework for energy management is presented. Its name is rEMpy, which stands for residential Energy Management in python. The framework has a modular structure and it is composed by an optimal scheduler, a user interface, a prediction module and the building thermal model. Unlike most of the EMs in literature, rEMpy is open-source, can be fully customized (in terms of tasks, modules and algorithms) and integrates in real-time a thermal modelling software. In this contribution, an overview of the rEMpy and its constitutive parts is given first, followed by a detailed description of the rEMpy modules and the communication system. The Computational Intelligence algorithms which perform forecasting, thermal modelling and optimal scheduling are also presented. The performance of rEMpy is finally evaluated in two case studies with different heating technologies and the results are reported and discussed. Marco Fagiani, Marco Severini, Stefano Squartini, Lucio Ciabattoni, Francesco Ferracuti, Alessandro Fonti, Gabriele Comodi |
CEC | 3 |
| 2017 | Acoustic novelty detection with adversarial autoencodersabstractNovelty detection is the task of recognising events the differ from a model of normality. This paper proposes an acoustic novelty detector based on neural networks trained with an adversarial training strategy. The proposed approach is composed of a feature extraction stage that calculates Log-Mel spectral features from the input signal. Then, an autoencoder network, trained on a corpus of “normal” acoustic signals, is employed to detect whether a segment contains an abnormal event or not. A novelty is detected if the Euclidean distance between the input and the output of the autoencoder exceeds a certain threshold. The innovative contribution of the proposed approach resides in the training procedure of the autoencoder network: instead of using the conventional training procedure that minimises only the Minimum Mean Squared Error loss function, here we adopt an adversarial strategy, where a discriminator network is trained to distinguish between the output of the autoencoder and data sampled from the training corpus. The autoencoder, then, is trained also by using the binary cross-entropy loss calculated at the output of the discriminator network. The performance of the algorithm has been assessed on a corpus derived from the PASCAL CHiME dataset. The results showed that the proposed approach provides a relative performance improvement equal to 0.26% compared to the standard autoencoder. The significance of the improvement has been evaluated with a one-tailed z-test and resulted significant with p <; 0.001. The presented approach thus showed promising results on this task and it could be extended as a general training strategy for autoencoders if confirmed by additional experiments. Emanuele Principi, Fabio Vesperini, Stefano Squartini, Francesco Piazza |
IJCNN | 3 |
| 2017 | A convolutional neural network approach for acoustic scene classificationabstractThis paper presents a novel application of convolutional neural networks (CNNs) for the task of acoustic scene classification (ASC). We here propose the use of a CNN trained to classify short sequences of audio, represented by their log-mel spectrogram. We also introduce a training method that can be used under particular circumstances in order to make full use of small datasets. The proposed system is tested and evaluated on three different ASC datasets and compared to other state-of-the-art systems which competed in the “Detection and Classification of Acoustic Scenes and Events” (DCASE) challenges held in 20161and 2013. The best accuracy scores obtained by our system on the DCASE 2016 datasets are 79.0% (development) and 86.2% (evaluation), which constitute a 6.4% and 9% improvements with respect to the baseline system. Finally, when tested on the DCASE 2013 evaluation dataset, the proposed system manages to reach a 77.0% accuracy, improving by 1% the challenge winner's score. Michele Valenti, Stefano Squartini, Aleksandr Diment, Giambattista Parascandolo, Tuomas Virtanen |
IJCNN | 2 |
| 2016 | Improving the performance of the AFAMAP algorithm for Non-Intrusive Load MonitoringabstractAmong the many electrical load disaggregation methods, often referred to as Non-Intrusive Load Monitoring techniques, the Additive Factorial Approximate MAP (AFAMAP) algorithm has shown outstanding capabilities and, therefore, it is nowadays regarded as a reference model. In order to achieve more accurate disaggregation results, and to satisfy real life environment requirements, further improvements in the algorithm are needed. In this work, the AFAMAP algorithm has been extended, by means of a differential forward model, thus complementing the existing differential backward model. Furthermore, an aggregated data examination method has been employed, aimed to the detection of inadmissible working state combinations of appliances, as well as the constraints setting based on the reactive power disaggregation feedback. The new approach has been evaluated by means of a subset, spanning over 6 months, of the Almanac of Minutely Power dataset (AMPds). On purpose, a real life environment, accounting 6 appliances, has been modelled and the carried out experiments revealed a improvement up to 18% with respect to the baseline AFAMAP. Roberto Bonfigli, Marco Severini, Stefano Squartini, Marco Fagiani, Francesco Piazza |
CEC | 3 |
| 2016 | Exploiting temporal features and pressure data for automatic leakage detection in smart water gridsabstractIn this paper, the unsupervised approach recently proposed by the authors for automatic leakage detection in smart water grids is extended. First of all, the EPANET tool is adopted in order to simulate more realistic leakages. Also, with respect to the original work, an additional time resolution, of 30 minutes, is included, based on the water dataset of the Almanac of Minutely Power Dataset (AMPds). New experiments are performed, as well, to evaluate the results of the application of both temporal features and pressure data. The pressure data is obtained by means of the EPANEt tool, whereas the leakages are induced at run-time for a more realistic behaviour. Two alternative sets of temporal features are evaluated by combining them with the features extracted from both flow and pressure data. Gaussian Mixture Models (GMMs), Hidden Markov Models (HMMs), and One-Class Support Vector Machine (OC-SVM) are used to characterize the normal data behaviour, under a comparative perspective. A feature selection strategy is adopted in computer simulations and the resulting performance indices are evaluated in terms of Area Under Curve (AUC). The obtained results show that the introduction of the temporal information produces a slight performance improvement for both flow and pressure data, but, most importantly, the combination of flow and pressure features allows a significant improvement of leakage detection for both GMM and HMM at every resolution, up to 88% of AUC. Marco Fagiani, Stefano Squartini, Roberto Bonfigli, Marco Severini, Francesco Piazza |
CEC | 2 |
| 2016 | An experimental study on new features for activity of daily living recognitionabstractIn the last few years, the researchers have spent many efforts in developing advanced systems for activity daily living (ADL) recognition in diverse applicative contexts, as home automation and ambient assisted living. Some of these need to know in real time the actions performed by a user, and this involves a number of additional issues to be taken into account during the recognition. In this paper, we present some improvements of a sliding window based approach to perform ADL recognition in a online fashion, i.e., recognizing activities as and when new sensor events are recorded. We describe seven methods used to extract features from the sequence of sensor events. The first four relate to previous works regarding the system of ADL recognition described, while, the last three represent the original contribution of this work. Support Vector Machine (SVM) has been used as classifier. Several experiments have been carried out by using a public smart home dataset and obtained results show that two of the three novel approaches allow to improve the recognition performance of the conventional methods, up to an increment of 5% with respect to the baseline feature extraction approach. Daniele Ferretti, Emanuele Principi, Stefano Squartini, Luigi Mandolini |
IJCNN | 3 |
| 2016 | Combining evolution strategies and neural network procedures for compression driver designabstractCompression driver design involves the study of complex mathematical models characterized by a great number of variables, implying high computational cost and long design time. Therefore, an optimization procedure is required to enhance the design procedure, especially from the parameters point of view. In this paper, a combined approach based both on evolution strategy procedure and neural network model is presented. Taking into consideration several tests on a real compression driver, the proposed method is capable to enhance the design procedure from the point of view of obtained frequency response and of the computational performance. Michele Gasparini, Fabio Vesperini, Stefania Cecchi, Stefano Squartini, Francesco Piazza, Romolo Toppi |
IJCNN | 4 |
| 2016 | Deep neural networks for Multi-Room Voice Activity Detection: Advancements and comparative evaluationabstractThis paper focuses on Voice Activity Detectors (VAD) for multi-room domestic scenarios based on deep neural network architectures. Interesting advancements are observed with respect to a previous work. A comparative and extensive analysis is lead among four different neural networks (NN). In particular, we exploit Deep Belief Network (DBN), Multi-Layer Perceptron (MLP), Bidirectional Long Short-Term Memory recurrent neural network (BLSTM) and Convolutional Neural Network (CNN). The latter has recently encountered a large success in the computational audio processing field and it has been successfully employed in our task. Two home recorded datasets are used in order to approximate real-life scenarios. They contain audio files from several microphones arranged in various rooms, from whom six features are extracted and used as input for the deep neural classifiers. The output stage has been redesigned compared to the previous author's contribution, in order to take advantage of the networks discriminative ability. Our study is composed by a multi-stage analysis focusing on the selection of the features, the network size and the input microphones. Results are evaluated in terms of Speech Activity Detection error rate (SAD). As result, a best SAD equal to 5.8% and 2.6% is reached respectively in the two considered datasets. In addiction, a significant solidity in terms of microphone positioning is observed in the case of CNN. Fabio Vesperini, Paolo Vecchiotti, Emanuele Principi, Stefano Squartini, Francesco Piazza |
IJCNN | 4 |
| 2016 | Acoustic cues from the floor: A new approach for fall classification
Emanuele Principi, Diego Droghini, Stefano Squartini, Paolo Olivetti, Francesco Piazza |
Expert Syst. Appl. | 3 |
| 2015 | A novel approach for automatic acoustic novelty detection using a denoising autoencoder with bidirectional LSTM neural networksabstractAcoustic novelty detection aims at identifying abnormal/novel acoustic signals which differ from the reference/normal data that the system was trained with. In this paper we present a novel unsupervised approach based on a denoising autoencoder. In our approach auditory spectral features are processed by a denoising autoencoder with bidirectional Long Short-Term Memory recurrent neural networks. We use the reconstruction error between the input and the output of the autoencoder as activation signal to detect novel events. The autoencoder is trained on a public database which contains recordings of typical in-home situations such as talking, watching television, playing and eating. The evaluation was performed on more than 260 different abnormal events. We compare results with state-of-the-art methods and we conclude that our novel approach significantly outperforms existing methods by achieving up to 93.4% F-Measure. Erik Marchi, Fabio Vesperini, Florian Eyben, Stefano Squartini, Björn W. Schuller |
ICASSP | 4 |
| 2015 | A novelty detection approach to identify the occurrence of leakage in smart gas and water gridsabstractIn this paper, a novelty detection algorithm for the identification of leakages in smart water/gas grid contexts is proposed. It is based on two separate stages: the first deals with the creation of the statistical leakage-free model, whereas the second evaluates the eventual occurrence of leakage on the basis of the model likelihood. Up to the authors' knowledge, this approach has never been used in the application scenario of interest. A set of several features are extracted from the Almanac of Minutely Power Dataset, and a suboptimal selection is executed to determinate the best combination. The abnormal event (leakage) is induced by manipulating the consumption in the test set. A total of 10 background models are created, by employing both Gaussian Mixture Models (GMMs) and Hidden Markov Models (HMMs) under a comparative perspective, and each of them is adopted to detect 10 leakages, with random duration, length and starting time. Finally, the performance are evaluated in terms of Area Under Curve (AUC) of the Receiver Operating Characteristic (ROC). Obtained results are more than encouraging: the best average AUCs of 85.60% and 87.97% are achieved with HMM, at 1 minute resolution, for natural gas and water, respectively. Specifically, considering true detection rates (TDRs) of 100%, the natural gas exhibits an overall false detection rate (FDR) of 17.11%, and the water achieves an overall FDR of 13.79%. Marco Fagiani, Stefano Squartini, Marco Severini, Francesco Piazza |
IJCNN | 2 |
| 2015 | A Deep Neural Network approach for Voice Activity Detection in multi-room domestic scenariosabstractThis paper presents a Voice Activity Detector (VAD) for multi-room domestic scenarios. A multi-room VAD (mVAD) simultaneously detects the time boundaries of a speech segment and determines the room where it was generated. The proposed approach is fully data-driven and is based on a Deep Neural Network (DNN) pre-trained as a Deep Belief Network (DBN) and fine-tuned by a standard error back-propagation method. Six different types of feature sets are extracted and combined from multiple microphone signals in order to perform the classification. The proposed DBN-DNN multi-room VAD (simply referred to as DBN-mVAD) is compared to other two NN based mVADs: a Multi-Layer Perceptron (MLP-mVAD) and a Bidirectional Long Short-Term Memory recurrent neural network (BLSTM-mVAD). A large multi-microphone dataset, recorded in a home, is used to assess the performance through a multi-stage analysis strategy comprising multiple feature selection stages alternated by network size and input microphones selections. The proposed approach notably outperforms the alternative algorithms in the first feature selection stage and in the network selection one. In terms of area under precision-recall curve (AUC), the absolute increment respect to the BLST-mVAD is 5.55%, while respect to the MLP-mVAD is 2.65%. Hence, solely the proposed approach undergoes the remaining selection stages. In particular, the DBN-mVAD achieves significant improvements: in terms of AUC and F-measure the absolute increments are equal to 10.41% and 8.56% with respect to the first stage of DBN-mVAD. Giacomo Ferroni, Roberto Bonfigli, Emanuele Principi, Stefano Squartini, Francesco Piazza |
IJCNN | 4 |
| 2015 | Non-linear prediction with LSTM recurrent neural networks for acoustic novelty detectionabstractAcoustic novelty detection aims at identifying abnormal/novel acoustic signals which differ from the reference/normal data that the system was trained with. In this paper we present a novel approach based on non-linear predictive denoising autoencoders. In our approach, auditory spectral features of the next short-term frame are predicted from the previous frames by means of Long-Short Term Memory (LSTM) recurrent denoising autoencoders. We show that this yields an effective generative model for audio. The reconstruction error between the input and the output of the autoencoder is used as activation signal to detect novel events. The autoencoder is trained on a public database which contains recordings of typical in-home situations such as talking, watching television, playing and eating. The evaluation was performed on more than 260 different abnormal events. We compare results with state-of-the-art methods and we conclude that our novel approach significantly outperforms existing methods by achieving up to 94.4% F-Measure. Erik Marchi, Fabio Vesperini, Felix Weninger, Florian Eyben, Stefano Squartini, Björn W. Schuller |
IJCNN | 5 |
| 2015 | Energy management with the support of dynamic pricing strategies in real micro-grid scenariosabstractAlthough smart grids are regarded as the technology to overcome the limits of nowadays power distribution grids, the transition will require much time. Dynamic pricing, a straightforward implementation of demand response, may provide the means to manipulate the grid load thus extending the life expectancy of current technology. However, to integrate a dynamic pricing scheme in the crowded pool of technologies, available at demand side, a proper energy manager with the support of a pricing profile forecaster is mandatory. Although energy management and price forecasting are recurrent topics, in literature they have been addressed separately. On the other hand, in this work, the aim is to investigate how well an energy manager is able to perform in presence of data uncertainty originating from the forecasting process. On purpose, an energy and resource manager has been revised and extended in the current manuscript. Finally, it has been complemented with a price forecasting technique, based on the Extreme Learning Machine paradigm. The proposed forecaster has proven to be better performing and more robust, with respect to the most common forecasting approaches. The energy manager, as well, has proven that the energy efficiency of the residential environment can be improved significantly. Nonetheless, to achieve the theoretical optimum, forecasting techniques tailored for that purpose may be required. Marco Severini, Stefano Squartini, Marco Fagiani, Francesco Piazza |
IJCNN | 2 |
| 2015 | Energy-Aware task scheduler for self-powered sensor nodes: From model to firmware
Marco Severini, Stefano Squartini, Francesco Piazza, Massimo Conti |
Ad Hoc Networks | 2 |
| 2015 | An integrated system for voice command recognition and emergency detection based on audio signals
Emanuele Principi, Stefano Squartini, Roberto Bonfigli, Giacomo Ferroni, Francesco Piazza |
Expert Syst. Appl. | 2 |
| 2015 | A review of datasets and load forecasting techniques for smart natural gas and water grids: Analysis and experiments
Marco Fagiani, Stefano Squartini, Leonardo Gabrielli, Susanna Spinsante, Francesco Piazza |
Neurocomputing | 2 |
| 2015 | Acoustic template-matching for automatic emergency state detection: An ELM based algorithm
Emanuele Principi, Stefano Squartini, Erik Cambria, Francesco Piazza |
Neurocomputing | 2 |
| 2015 | Computational Energy Management in Smart Grids
Stefano Squartini, Derong Liu 0001, Francesco Piazza, Dongbin Zhao, Haibo He |
Neurocomputing | 1 |
| 2015 | Signer independent isolated Italian sign recognition based on hidden Markov models
Marco Fagiani, Emanuele Principi, Stefano Squartini, Francesco Piazza |
Pattern Anal. Appl. | 3 |
| 2014 | A nonlinear second-order digital oscillator for Virtual Acoustic FeedbackabstractThe guitar feedback effect, or howling, is well known to the general public and identified with many rock music genres and it is the only case of acoustic feedback employed for musical purposes. Virtual Acoustic Feedback (VAF), is regarded as the extension of this phenomenon to any instrument or sound source by means of virtual acoustics and is meant to enrich the sound palette of a musician. The study of the acoustic feedback as a musical tool and computational techniques for its emulation have been scarcely addressed in literature. In this paper a nonlinear feedback oscillator is proposed and its properties derived. The oscillator does not necessarily need to be connected to a virtual instrument, thus enables to process any kind of pitched real-time input. Leonardo Gabrielli, M. Giobbi, Stefano Squartini, Vesa Välimäki |
ICASSP | 3 |
| 2014 | Multi-resolution linear prediction based features for audio onset detection with bidirectional LSTM neural networksabstractA plethora of different onset detection methods have been proposed in the recent years. However, few attempts have been made with respect to widely-applicable approaches in order to achieve superior performances over different types of music and with considerable temporal precision. In this paper, we present a multi-resolution approach based on discrete wavelet transform and linear prediction filtering that improves time resolution and performance of onset detection in different musical scenarios. In our approach, wavelet coefficients and forward prediction errors are combined with auditory spectral features and then processed by a bidirectional Long Short-Term Memory recurrent neural network, which acts as reduction function. The network is trained with a large database of onset data covering various genres and onset types. We compare results with state-of-the-art methods on a dataset that includes Bello, Glover and ISMIR 2004 Ballroom sets, and we conclude that our approach significantly outperforms existing methods in terms of F-Measure. For pitched non percussive music an absolute improvement of 7.5% is reported. Erik Marchi, Giacomo Ferroni, Florian Eyben, Leonardo Gabrielli, Stefano Squartini, Björn W. Schuller |
ICASSP | 5 |
| 2014 | Computational Intelligence in Smart water and gas grids: An up-to-date overviewabstractComputational Intelligence plays a relevant role in several Smart Grid applications, and there is a florid literature in this regard. However, most of the efforts have been oriented to the electrical energy field, for which many contributions have appeared so far, also facilitated by the availability of suitable databases to use for system training and testing. Different is the case for the water and gas scenarios: this work is thus oriented to present the state-of-the-art techniques for these grids, from 2009 to date. In particular, the focus is on load forecasting and leakage detection applications, that are the most addressed in the literature and present the biggest interest from a commercial point of view as well: the main characteristics and registered performance for all the reviewed approaches are reported. Along this direction, an extensive search of used databases has been performed and thus made available to the research community. Marco Fagiani, Stefano Squartini, Leonardo Gabrielli, Mirco Pizzichini, Susanna Spinsante |
IJCNN | 2 |
| 2014 | Audio onset detection: A wavelet packet based approach with recurrent neural networksabstractThis paper concerns the exploitation of multi-resolution time-frequency features via Wavelet Packet Transform to improve audio onset detection. In our approach, Wavelet Packet Energy Coefficients (WPEC) and Auditory Spectral Features (ASF) are processed by Bidirectional Long Short-Term Memory (BLSTM) recurrent neural network that yields the onsets location. The combination of the two feature sets, together with the BLSTM based detector, form an advanced energy-based approach that takes advantage from the multi-resolution analysis given by the wavelet decomposition of the audio input signal. The neural network is trained with a large database of onset data covering various genres and onset types. Due to its data-driven nature, our approach does not require the onset detection method and its parameters to be tuned to a particular type of music. We show a comparison with other types and sizes of recurrent neural networks and we compare results with state-of-the-art methods on the whole onset dataset. We conclude that our approach significantly increase performance in terms of F-measure without any music genres or onset type constraints. Erik Marchi, Giacomo Ferroni, Florian Eyben, Stefano Squartini, Björn W. Schuller |
IJCNN | 4 |
| 2014 | Power Normalized Cepstral Coefficients based supervectors and i-vectors for small vocabulary speech recognitionabstractTemplate-matching and discriminative techniques, like support vector machines (SVMs), have been widely used for automatic speech recognition. Both methods require that varying length sequences are mapped to vectors of fixed lengths: in template-matching, the problem is solved by means of dynamic time warping (DTW), while in SVM with dynamic kernels. The supervector and i-vector paradigms seem to represent a valid solution to such a problem when SVM are employed for classification. In this work, Gaussian mean supervectors (GMS), Gaussian posterior probability supervectors (GPPS) and i-vectors are evaluated as features both for template-matching and for SVM-based speech recognition in a comparative fashion. All these features are based on Power Normalized Cepstral Coefficients (PNCCs) directly extracted from speech utterances. The different methods are assessed in small vocabulary speech recognition tasks using two distinct corpora, and they have been compared to DTW, dynamic time alignment kernel (DTAK), outerproduct of trajectory matrix, and PocketSphinx as further recognition techniques to be evaluated. Experimental results showed the appropriateness of the supervector and i-vector based solutions with respect to the other state-of-the art techniques here addressed. Emanuele Principi, Stefano Squartini, Francesco Piazza |
IJCNN | 2 |
| 2014 | Computational framework based on task and resource scheduling for micro grid designabstractWithin micro grid scenarios, optimal energy management represents an important paradigm to improve the grid efficiency while lowering its burden. While usually real time energy management is considered, an offline approach can be also adopted to maximize the grid efficiency in certain contexts. Indeed, by evaluating the energy management performance according to the user needs, it is possible to asses which technologies allow the overall system to operate at its best, given the expected load level. From this perspective, a computational framework based on the "Mixed-Integer Linear Programming" paradigm has been proposed in this paper as a tool to simulate the micro grid behaviour in terms of energy consumption and in dependence on the technology of choice. By modelling the energy production and storage means, the pool of electricity tasks, and the thermal behaviour of the building, suitable energy management policies for the micro grid scenario under study can be developed and tested in different operating conditions and time horizons. Moreover, the forecasting paradigm has been integrated into the framework to deal with data uncertainty, and a Neural Network approach has been employed on purpose. Performed computer simulations, related to a six-apartments building scenario, have proven that the suggested framework can fruitfully be adopted to assess the effectiveness of different technical solutions in terms of overall energy cost, thus supporting the decisional process occurring during the micro grid design. Marco Severini, Stefano Squartini, Francesco Piazza |
IJCNN | 2 |
| 2013 | Real-life voice activity detection with LSTM Recurrent Neural Networks and an application to Hollywood moviesabstractA novel, data-driven approach to voice activity detection is presented. The approach is based on Long Short-Term Memory Recurrent Neural Networks trained on standard RASTA-PLP frontend features. To approximate real-life scenarios, large amounts of noisy speech instances are mixed by using both read and spontaneous speech from the TIMIT and Buckeye corpora, and adding real long term recordings of diverse noise types. The approach is evaluated on unseen synthetically mixed test data as well as a real-life test set consisting of four full-length Hollywood movies. A frame-wise Equal Error Rate (EER) of 33.2% is obtained for the four movies and an EER of 9.6% is obtained for the synthetic test data at a peak SNR of 0 dB, clearly outperforming three state-of-the-art reference algorithms under the same conditions. Florian Eyben, Felix Weninger, Stefano Squartini, Björn W. Schuller |
ICASSP | 3 |
| 2013 | A distributed system for recognizing home automation commands and distress calls in the Italian languageabstractThis paper describes a system for recognizing distress calls and home automation voice commands in a smart-home. Distress calls are recognized with the purpose of assisting people in their own homes: when they are detected, a phone call is automatically established with a contact in a address book and the person can request for assistance. The voice call is established through a voice over ip stack, with hands-free communication guaranteed by an acoustic echo canceller. The acoustic environment is constantly monitored by several low-consuming devices distributed throughout the home. In each device, a voice activity detector detects speech segments, and a speech recognition engine recognizes commands and distress calls. Robustness to environmental disturbances has been increased by employing Power Normalized Cepstral Coefficients and by using an adaptive algorithm for interference cancellation. An Italian speech corpus of home automation commands and distress calls has been developed for evaluation purposes. The corpus has been recorded in a real room using multiple microphones, and each sentence has been uttered both in normal and shouted speaking styles. The system performance has been assessed in terms of commands/distress recognition accuracy in order to prove the effectiveness of the approach. Emanuele Principi, Stefano Squartini, Francesco Piazza, Danilo Fuselli, Maurizio Bonifazi |
INTERSPEECH | 2 |
| 2013 | Evaluation of the Wireless M-Bus standard for future smart water gridsabstractThe most recent Wireless Sensor Networks technologies can provide viable solutions to perform automatic monitoring of the water grid, and smart metering of water consumptions. However, sensor nodes located along water pipes cannot access power grid facilities, to get the necessary energy imposed by their working conditions. In this sense, it is of basic importance to design the network architecture in such a way as to require the minimum possible power. This paper investigates the suitability of the Wireless Metering Bus protocol for possible adoption in future smart water grids, by evaluating its transmission performance, through simulations and experimental tests executed by means of prototype sensor nodes. Susanna Spinsante, Mirco Pizzichini, Matteo Mencarelli, Stefano Squartini, Ennio Gambi |
IWCMC | 4 |
| 2013 | Online sequential extreme learning machine in nonstationary environments
Yibin Ye, Stefano Squartini, Francesco Piazza |
Neurocomputing | 2 |
| 2013 | Energy-aware lazy scheduling algorithm for energy-harvesting sensor nodes
Marco Severini, Stefano Squartini, Francesco Piazza |
Neural Comput. Appl. | 2 |
| 2013 | The neural paradigm for complex systems: new algorithms and applications
Stefano Squartini, Jinhu Lü 0001, Qinglai Wei |
Neural Comput. Appl. | 1 |
| 2013 | Hybrid soft computing algorithmic framework for smart home energy management
Marco Severini, Stefano Squartini, Francesco Piazza |
Soft Comput. | 2 |
| 2013 | Optimal Home Energy Management Under Dynamic Electrical and Thermal ConstraintsabstractThe optimization of energy consumption, with consequent costs reduction, is one of the main challenges in present and future smart grids. Of course, this has to occur keeping the living comfort for the end-user unchanged. In this work, an approach based on the mixed-integer linear programming paradigm, which is able to provide an optimal solution in terms of tasks power consumption and management of renewable resources, is developed. The proposed algorithm yields an optimal task scheduling under dynamic electrical constraints, while simultaneously ensuring the thermal comfort according to the user needs. On purpose, a suitable thermal model based on heat-pump usage has been considered in the framework. Some computer simulations using real data have been performed, and obtained results confirm the efficiency and robustness of the algorithm, also in terms of achievable cost savings. Francesco De Angelis 0002, Matteo Boaro, Danilo Fuselli, Stefano Squartini, Francesco Piazza, Qinglai Wei |
IEEE Trans. Ind. Informatics | 4 |
| 2012 | Optimal Task and Energy Scheduling in Dynamic Residential Scenarios
Francesco De Angelis 0002, Matteo Boaro, Danilo Fuselli, Stefano Squartini, Francesco Piazza, Qinglai Wei, Ding Wang 0001 |
ISNN (1) | 4 |
| 2012 | Optimal Battery Management with ADHDP in Smart Home Environments
Danilo Fuselli, Francesco De Angelis 0002, Matteo Boaro, Derong Liu 0001, Qinglai Wei, Stefano Squartini, Francesco Piazza |
ISNN (2) | 6 |
| 2012 | Dominance Detection in a Reverberated Acoustic Scenario
Emanuele Principi, Rudy Rotili, Martin Wöllmer, Stefano Squartini, Björn W. Schuller |
ISNN (1) | 4 |
| 2012 | An Energy Aware Approach for Task Scheduling in Energy-Harvesting Sensor Nodes
Marco Severini, Stefano Squartini, Francesco Piazza |
ISNN (2) | 2 |
| 2012 | Environmental robust speech and speaker recognition through multi-channel histogram equalization
Stefano Squartini, Emanuele Principi, Rudy Rotili, Francesco Piazza |
Neurocomputing | 1 |
| 2011 | Real-Time Speech Recognition in a Multi-talker Reverberated Acoustic Scenario
Rudy Rotili, Emanuele Principi, Stefano Squartini, Björn W. Schuller |
ICIC (2) | 3 |
| 2011 | On-Line Extreme Learning Machine for Training Time-Varying Neural Networks
Yibin Ye, Stefano Squartini, Francesco Piazza |
ICIC (3) | 2 |
| 2011 | Real-Time Joint Blind Speech Separation and Dereverberation in Presence of Overlapping Speakers
Rudy Rotili, Emanuele Principi, Stefano Squartini, Francesco Piazza |
ISNN (2) | 3 |
| 2011 | Robust Multi-stream Keyword and Non-linguistic Vocalization Detection for Computationally Intelligent Virtual Agents
Martin Wöllmer, Erik Marchi, Stefano Squartini, Björn W. Schuller |
ISNN (2) | 3 |
| 2011 | ELM-Based Time-Variant Neural Networks with Incremental Number of Output Basis Functions
Yibin Ye, Stefano Squartini, Francesco Piazza |
ISNN (1) | 2 |
| 2010 | Joint Multichannel Blind Speech Separation and Dereverberation: A Real-Time Algorithmic Implementation
Rudy Rotili, Claudio De Simone, Alessandro Perelli, Simone Cifani, Stefano Squartini |
ICIC (3) | 5 |
| 2010 | Incremental-Based Extreme Learning Machine Algorithms for Time-Variant Neural Networks
Yibin Ye, Stefano Squartini, Francesco Piazza |
ICIC (1) | 2 |
| 2010 | Robust speech recognition using feature-domain multi-channel bayesian estimatorsabstractThis paper proposes innovative multi-channel bayesian estimators in the feature-domain for robust speech recognition. Both minimum-mean-squared-error (MMSE) and maximum-a-posteriori (MAP) criteria have been explored: the related algorithms extend the multi-channel frequency-domain counterparts and generalize the single-channel feature-domain MMSE solution, recently appeared in the literature. Computer simulations conducted on a modified AURORA2 database show the efficacy of the frequency-domain multi-channel estimators when used as a pre-processing stage of a speech recognition engine, and that the proposed multi-channel MAP approach outperforms single-channel estimators by at least 3% on average. Emanuele Principi, Rudy Rotili, Simone Cifani, Lorenzo Marinelli, Stefano Squartini, Francesco Piazza |
ISCAS | 5 |
| 2008 | Stability analysis of natural gradient learning rules in overdetermined ICA
Stefano Squartini, Andrea Arcangeli, Francesco Piazza |
Signal Process. | 1 |
| 2007 | Discrete Stockwell Transform and Reduced Redundancy Versions from Frame Theory ViewpointabstractThe present work gives an interpretation of the discrete Stockwell transform (DST) from the theory of frame (TOF) perspective, showing first of all that the complete set of expansion functions of the DST is a frame. Starting from this interpretation, two versions of the DST are derived that allow to get a reduced redundancy representation of the original time series, when compared to the complete DST: the first method is based on a dyadic tiling of the time-frequency plane and leads to a perfect reconstruction representation using O(N) coefficients, whilst the second method is inspired by the matching pursuit technique and only provides an approximate representation of the original time series, but achieves better performances in terms of compactness, useful property in many practical applications. Alessandro Bastari, Stefano Squartini, Francesco Piazza |
ISCAS | 2 |
| 2007 | Gaussianization Based Approach for Post-Nonlinear Underdetermined BSS with Delays
Alessandro Bastari, Stefano Squartini, Stefania Cecchi, Francesco Piazza |
ISNN (3) | 2 |
| 2007 | Echo State Networks for Real-Time Audio Applications
Stefano Squartini, Stefania Cecchi, Michele Rossini, Francesco Piazza |
ISNN (3) | 1 |
| 2007 | Stability Analysis of Natural Gradient Learning Rules in Complete ICA: A Unifying PerspectiveabstractThis letter deals with the independent component analysis (ICA) problem in the complete case. As appeared recently in the literature, different Riemannian metrics can be defined within the parameter space (i.e., the general linear group), allowing to derive correspondingly various ICA learning rules based on the relative natural gradients (NGs). This letter proposes a general framework to analyze the stability of such learning rules, including the already published study focusing on the Amari's NG approach as a special case thereof. In particular, it is shown that the stability conditions known in the literature still hold in all cases addressed Stefano Squartini, Andrea Arcangeli, Francesco Piazza |
IEEE Signal Process. Lett. | 1 |
| 2006 | New Riemannian metrics for speeding-up the convergence of over- and underdetermined ICAabstractIn this paper some alternative Riemannian metrics are defined on the parameter space of non-square matrices, corresponding to various translations defined therein. Such metrics allow the authors to derive novel learning rules for two ICA based algorithms for over-determined blind source separation (BSS), which tries to separate less sources from more sensors. Computer simulations show a significant improvement of the convergence speed when second-order translations are employed in contrast to their first-order counterparts, extending known results for complete BSS Stefano Squartini, Francesco Piazza, Fabian J. Theis |
ISCAS | 1 |
| 2004 | An approach employing signal sparse representation in wavelet domain for underdetermined blind source separationabstractSeveral contributions in literature have recently proposed techniques based on assumption of source sparsity in some representation domain to give a solution to the problem of blind source separation in the underdetermined case. This work investigates how to employ wavelet based sparse representation of signals in an already existing algorithm for the problem under study, in order to improve separability of sources, in comparison to application of short time Fourier transform. Different wavelet transforms are considered. Moreover, this approach allows to perform a suitable de-noising operation after the separation algorithm, by thresholding the wavelet coefficients corresponding to extracted sources. This occurs at a very low computational cost, resulting in a further improvement of source recovering when noise is present at mixture level. Experimental results confirm the effectiveness of what implemented. Eraldo Pomponi, Stefano Squartini, Francesco Piazza |
IJCNN | 2 |
| 2003 | A recurrent multiscale architecture for long-term memory prediction taskabstractIn the past few years, researchers have been extensively studying the application of recurrent neural networks (RNNs) to solving tasks where detection of long term dependencies is required. This paper proposes an original architecture termed the Recurrent Multiscale Network, RMN, to deal with these kinds of problems. Its most relevant properties are concerned with maintaining conventional RNNs' capability of information storing whilst simultaneously attempting to reduce their typical drawback occurring when they are trained by gradient descent algorithms, namely the vanishing gradient effect. This is achieved through RMN which preprocesses the original signal separating information at different temporal scales through an adequate DSP tool, and handling each information level with an autonomous recurrent architecture; the final goal is achieved by a nonlinear reconstruction section. This network has shown a markedly improved generalization performance over conventional RNNs, in its application to time series prediction tasks where long range dependencies are involved. Stefano Squartini, Amir Hussain 0001, Francesco Piazza |
ICASSP (2) | 1 |
| 2003 | Attempting to reduce the vanishing gradient effect through a novel recurrent multiscale architectureabstractThis paper proposes a possible solution to the vanishing gradient problem in recurrent neural networks, occurring when such networks are applied to solving tasks where detection of long term dependencies is required. The main idea consists of pre-processing the signal (a time series typically) through a discrete wavelet decomposition, in order to separate the short term information from the long term ones, and treating each scale by different recurrent neural networks. The partial results concerning all the sequences at diverse time/frequency resolutions are combined through an adaptive nonlinear structure in order to achieve the final goal. This new preprocessing based approach is distinct from the other one reported in literature to-date, as it tends to mitigate the effects of the problem under study avoiding relevant changing in network's architecture and learning techniques. The overall system (called recurrent multiscale network, RMN) is described and its performances tested through typical tasks namely the latching problem and time series prediction. Stefano Squartini, Amir Hussain 0001, Francesco Piazza |
IJCNN | 1 |