Daniele Salvati

dblp:119/8669 · DBLP profile ↗
← Back
22ranked-venue papers
18as first author
10since 2021 · last 2026
0000-0002-8042-0333ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 13 first-author · 6 since 2021Artificial intelligence and machine learning · 11 · 8 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Iterative high-order spherical harmonic diagonal-unloading beamforming for localizing direct sound and dominant reflections in reverberant rooms
abstract
Reverberation and early reflections create spurious peaks in acoustic power maps and bias direction-of-arrival (DOA) estimates. At the same time, the direct path and dominant reflections carry useful information about the acoustic environment. This paper investigates joint localization of direct sound and multiple dominant reflected components in single-source reverberant environments using spherical microphone arrays (SMAs). It introduces an iterative spherical-harmonic (SH) diagonal-unloading (DU) beamforming framework for multipath analysis. Microphone signals are encoded in the SH domain and fused into a broadband covariance matrix by frequency smoothing with per-bin trace normalization. Multi-peak localization alternates DU map evaluation and iterative covariance filtering, with detected directions suppressed through a two-sided null-steering update and per-iteration trace renormalization. Performance is assessed through controlled simulations spanning reverberation time, signal-to-noise ratio (SNR), signal-to-diffuse-noise ratio (SDNR), and gain mismatch, together with experiments based on multichannel room impulse responses recorded with a spherical microphone array. Results in terms of root-mean-square error (RMSE) and accuracy rate (AR) show that increasing SH order improves localization of the direct path and strongest reflected components, whereas later iterations become more challenging as weaker components are extracted. The proposed method outperforms the baselines while retaining a computationally attractive structure for SH processing.
Daniele Salvati
Signal Process.1
2026 Entropy-Based Geometry Design for SRP-PHAT Acoustic Source Localization
abstract
A novel entropy-based method is proposed for microphone array geometry design in scalable steered response power phase transform (SRP-PHAT) acoustic source localization (ASL). The approach relies on a spatial sensitivity map that models the density of time-difference-of-arrival (TDOA) samples projected onto the search grid through volumetric accumulation of generalized cross-correlation (GCC) functions. By maximizing the entropy of the normalized sensitivity map, the proposed formulation promotes a spatially uniform distribution of TDOA information over the region of interest. The resulting geometry optimization, solved via a derivative-free pattern search, consistently leads to perimeter (edge-based) microphone configurations. Simulation and real-world experiments demonstrate improved ASL performance compared to conventional array geometries.
Daniele Salvati
IEEE Signal Process. Lett.1
2025 Raw Audio Deep Learning Filter Banks for Acoustic Scene Classification
abstract
An acoustic scene classification (ASC) framework based on filter banks and raw audio signals is proposed. The time-domain waveform captured by a microphone is initially processed by a filter bank to produce raw audio subbands, which are then processed by a deep neural network (DNN) to estimate the scene class. The network model consists of parallel convolutional neural network (CNN) branches and utilizes late fusion for output class prediction. Each CNN pipeline is designed to handle both the time-domain waveform and the raw audio subbands. The filter bank enhances the frequency components that are crucial for scene recognition, thereby improving classification accuracy. Experiments conducted using the DCASE 2019 dataset show a significant performance improvement compared to models that use raw audio or log-mel spectrogram inputs.
Daniele Salvati
ICASSP1
2023 A late fusion deep neural network for robust speaker identification using raw waveforms and gammatone cepstral coefficients
abstract
Speaker identification aims at determining the speaker identity by analyzing his voice characteristics, and relies typically on statistical models or machine learning techniques. Frequency-domain features are by far the most used choice to encode the audio input in sound recognition. Recently, some studies have also analyzed the use of time-domain raw waveform (RW) with deep neural network (DNN) architectures. In this paper, we hypothesize that both time-domain and frequency-domain features can be used to increase the robustness of speaker identification task in adverse noisy and reverberation conditions, and we present a method based on a late fusion DNN using RWs and gammatone cepstral coefficients (GTCCs). We analyze the characteristics of RW and spectrum-based short-time features, reporting advantages and limitations, and we show that the joint use can increase the identification accuracy. The proposed late fusion DNN model consists of two independent DNN branches made primarily by convolutional neural networks (CNN) and fully connected neural networks (NN) layers. The two DNN branches have as input short-time RW audio fragments and GTCCs, respectively. The late fusion is computed on the predicted scores of the DNN branches. Since the method is based on short segments, it has the advantage of being independent from the size of the input audio signal, and the identification task can be computed by summing the predicted scores over several short-time frames. Analysis of speaker identification performance computed with simulations show that the late fusion DNN model improves the accuracy rate in adverse noise and reverberation conditions in comparison to the RW, the GTCC, and the mel-frequency cepstral coefficients (MFCCs) features. Experiments with real-world speech datasets confirm the efficiency of the proposed method, especially with small-size audio samples.
Daniele Salvati, Carlo Drioli, Gian Luca Foresti
Expert Syst. Appl.1
2023 Efficient parallel kernel based on Cholesky decomposition to accelerate multichannel nonnegative matrix factorization
Antonio Jesús Muñoz-Montoro, Julio J. Carabias-Orti, Daniele Salvati, Raquel Cortina
J. Supercomput.3
2022 Efficient Detection and Localization of Acoustic Sources with a low complexity CNN network and the Diagonal Unloading Beamforming
abstract
Detection of acoustic events and direction of arrival estimation of acoustic sources are nowadays central topics in the field of acoustic array signal processing, providing both theoretical and practical relevant perspectives. Reconnaissance and surveillance against intrusions, search and rescue in hostile environments, speaker detection and localization are examples of real applications in which an accurate and efficient analysis of the acoustic scene is required. In this context, we have been investigating an efficient CNN neural network-based method capable of learning from the multi-channel signal an intrinsic function between a specific choice of audio-related features and both the nature of the acoustic event and its spatial location. In this work, we investigate an extended CNN network and compare its performance with the accuracy of a reduced complexity CNN network. The main novelty introduced in this research with respect to the state-of-the-art is that we propose a Diagonal Unloading (DU) Beamforming-based method that produces acoustic maps of azimuth and elevation angles to generate the feature representation of the acoustic signal. A comparative study with the Log-Mel Spectrogram feature representation is also conduced along with a method with fusion of the two feature representations that has been experimented in this work. The dataset for both training and validation of the CNN network belongs to the DCASE challenge. The experiments demonstrated the benefits introduced by the DU-Acoustic Map feature representation that provides additional information about the position of acoustic sources, in terms of angles of azimuth and elevation, through the acoustic maps. The accuracy and efficiency of the proposed deep learning-based method are confirmed by the results.
Andrea Toma, Daniele Salvati, Carlo Drioli, Gian Luca Foresti
IJCNN2
2022 Acoustic Source Localization Using a Geometrically Sampled Grid SRP-PHAT Algorithm With Max-Pooling Operation
abstract
The steered response power phase transform (SRP-PHAT) is a well-known algorithm for acoustic source localization using microphone arrays. It consists in the computation of the generalized cross-correlation (GCC) between each microphone pair, and in the coherent summation of the GCC values in the grid search space. Several improvements based on the volumetric grid have been proposed in order to achieve spatial resolution scalability and to reduce the computational cost by using a coarser grid. In general, the problem of the volumetric based methods is that the noise and the reverberation are projected into the search space since all GCC information is used to build the acoustic map. It is hence proposed a volumetric grid SRP-PHAT algorithm based on the geometrically sampled grid (GSG) that incorporates a max-pooling (MP) operation in the volume accumulation of the GCC values in order to improve the localization performance. The MP is the solution of a minimization-maximization problem that aims at minimizing the deleterious effect of noise and reverberation and at maximizing the accuracy of the GCC values related to the target sound source. Simulations and real-world experiments demonstrate the efficiency of the proposed SRP-GSG-MP algorithm in adverse conditions.
Daniele Salvati, Carlo Drioli, Gian Luca Foresti
IEEE Signal Process. Lett.1
2021 Time Delay Estimation for Speaker Localization Using CNN-Based Parametrized GCC-PHAT Features
Daniele Salvati, Carlo Drioli, Gian Luca Foresti
Interspeech1
2021 CNN-Based Processing of Acoustic and Radio Frequency Signals for Speaker Localization from MAVs
abstract
A novel speaker localization algorithm from micro aerial vehicles (MAVs) is investigated. It introduces a joint direction of arrival (DOA) and distance prediction method based on processing and fusion of the multi-channel speech data with radio frequency (RF) measurements of the received signal strength. Possible applications include unmanned aerial vehicles (UAVs)based reconnaissance and surveillance against intrusions and search and rescue in hostile environments. A 3-stages convolutional neural network (CNN) with a fusion layer is proposed to perform this task with the objective of augmenting the source localization from multi-channel speech signals. Two parallel CNNs process the speech and RF data, and the regression network produces predictions of the angle and distance from the source after the fusion layer. To show the performance and effectiveness of this RF-assisted method, the experimental scenario and datasets are presented and experiments are then discussed along with the results that have been obtained.
Andrea Toma, Daniele Salvati, Carlo Drioli, Gian Luca Foresti
Interspeech2
2021 Acoustic Target Tracking Through a Cluster of Mobile Agents
abstract
This paper discusses the problem of tracking a moving target by means of a cluster of mobile agents that is able to sense the acoustic emissions of the target, with the aim of improving the target localization and tracking performance with respect to conventional fixed-array acoustic localization. We handle the acoustic part of the problem by modeling the cluster as a sensor network, and we propose a centralized control strategy for the agents that exploits the spatial sensitivity pattern of the sensor network to estimate the best possible cluster configuration with respect to the expected target position. In order to take into account the position estimation delay due to the frame-based nature of the processing, the possible positions of the acoustic target in a given future time interval are represented in terms of a compatible set, that is, the set of all possible future positions of the target, given its dynamics and its present state. A frame-by-frame cluster reconfiguration algorithm is presented, which adapts the position of each sensing agent with the goal of pursuing the maximum overlap between the region of high acoustic sensitivity of the entire cluster and the compatible set of the sound-emitting target. The tracking scheme iterates, at each observation frame, the computation of the target compatible set, the reconfiguration of the cluster, and the target acoustic localization. The reconfiguration step makes use of an opportune cost function proportional to the difference of the compatibility set and the acoustic sensitivity spatial pattern determined by the mobile agent positions. Simulations under different geometric configurations and positioning constraints demonstrate the ability of the proposed approach to effectively localize and track a moving target based on its acoustic emission. The Doppler effect related to moving sources and sensors is taken into account, and its impact on performance is analyzed. We compare the localization results with conventional static-array localization and positioning of acoustic sensors through genetic algorithm optimization, and results demonstrate the sensible improvements in terms of localization and tracking performance. Although the method is discussed here with respect to acoustic target tracking, it can be effectively adapted to video-based localization and tracking, or to multimodal information settings (e.g., audio and video).
Carlo Drioli, Giulia Giordano, Daniele Salvati, Franco Blanchini, Gian Luca Foresti
IEEE Trans. Cybern.3
2020 Two-Microphone End-to-End Speaker Joint Identification and Localization Via Convolutional Neural Networks
abstract
We present an end-to-end scheme based on convolutional neural networks (CNNs) for speaker joint identification and localization. We investigate the possibility to estimate both the direction of arrival (DOA) and the identity of the speaker in far-field noisy and reverberant conditions using a two-channel microphone array. The proposed CNN network is designed to map the raw waveform of the two channels into the speaker identity and into the DOA of its speech signal. We analyze the identification and localization performance with simulated experiments in noisy and reverberation conditions.
Daniele Salvati, Carlo Drioli, Gian Luca Foresti
IJCNN1
2020 Diagonal Unloading Beamforming in the Spherical Harmonic Domain for Acoustic Source Localization in Reverberant Environments
abstract
Spherical microphone arrays allow the sound field analysis in three dimensions with the advantage of having the same resolution in all directions. By considering the frequency-independent character of the steering vectors in the spherical harmonic (SH) domain, we propose a very low-complexity SH diagonal unloading (DU) beamforming with a novel frequency smoothing power transform (FSPT) of the covariance matrices. We consider the direction of arrival (DOA) estimation problem of acoustic sources in reverberant conditions. The DU beamforming provides high resolution directional response since it exploits the subspace orthogonality property of the covariance matrix by the removal or the attenuation of the signal subspaces, obtained through the subtraction of an opportune diagonal matrix from the covariance matrix. The FSPT aims at smoothing the narrowband covariance matrices of the entire set of frequency domain components, and it pursues this goal by minimizing the narrowband error contributions due to reverberation in the broadband frequency smoothing covariance matrix. We analyze the DOA estimation performance using speech signals with simulations and real acoustic data in reverberant conditions. The results show that the proposed SH-DU-FSPT has a DOA estimation performance comparable to that of high resolution state-of-the-art methods with a significant reduction of the computational cost, since the steering directional responses are computed on the broadband frequency smoothing covariance matrix.
Daniele Salvati, Carlo Drioli, Gian Luca Foresti
IEEE ACM Trans. Audio Speech Lang. Process.1
2019 End-to-End Speaker Identification in Noisy and Reverberant Environments Using Raw Waveform Convolutional Neural Networks
Daniele Salvati, Carlo Drioli, Gian Luca Foresti
INTERSPEECH1
2019 Power Method for Robust Diagonal Unloading Localization Beamforming
abstract
We propose a robust version of the diagonal unloading (DU) beamforming for the acoustic source localization problem in high noise conditions. The DU beamformer exploits the subspace orthogonality property by the removal or the attenuation of the signal subspaces, obtained through the subtraction of an opportune diagonal matrix from the covariance matrix. As a result, it provides high-resolution directional response with low computational complexity. We show that a robust DU beamformer can be implemented by subtracting the largest eigenvalue of the estimated covariance matrix from the diagonal elements, and that this implementation is valid in general (i.e., for both the single-source and the multiple-source case). We propose the use of the power method for the estimation of the largest eigenvalue in the DU procedure. We show with numerical simulations that the proposed method improves the localization performance in high noise conditions without substantial increment of the computational cost. Applications for this method include a number of scenarios involving multirotor aerial systems due to its robustness to the noise and its low computational complexity.
Daniele Salvati, Carlo Drioli, Gian Luca Foresti
IEEE Signal Process. Lett.1
2018 Sensitivity-based region selection in the steered response power algorithm
Daniele Salvati, Carlo Drioli, Gian Luca Foresti
Signal Process.1
2018 A Low-Complexity Robust Beamforming Using Diagonal Unloading for Acoustic Source Localization
abstract
In acoustic array processing, beamforming is a class of algorithms commonly used to estimate the position of a radiating sound source. This paper presents a diagonal unloading (DU) transformation method for the conventional response power beamforming to achieve robust localization with low computational complexity. The transformation is obtained by subtracting an opportune diagonal matrix from the covariance matrix of the array output vector. Specifically, the DU beamformer aims at subtracting the signal subspace from the noisy signal space. It is, hence, a data-dependent covariance matrix conditioning method. We show how to calculate precisely the unloading parameters, and we present a comparison of the proposed DU beamforming, the robust minimum variance distortionless response (MVDR) filter, and the multiple signal classification (MUSIC) method, in terms of their respective eigenanalyses. Theoretical analysis and experiments conducted on both simulated and real acoustic data demonstrate that the DU beamformer localization performance is comparable to that of robust MVDR and MUSIC. Since its computational cost is equivalent to that of a conventional beamformer, the proposed DU beamformer method can, thus, be very attractive due to its effectiveness and computational efficiency.
Daniele Salvati, Carlo Drioli, Gian Luca Foresti
IEEE ACM Trans. Audio Speech Lang. Process.1
2016 A weighted MVDR beamformer based on SVM learning for sound source localization
Daniele Salvati, Carlo Drioli, Gian Luca Foresti
Pattern Recognit. Lett.1
2016 Sound Source and Microphone Localization From Acoustic Impulse Responses
abstract
This letter proposes a new method for source and microphone localization in reverberant environments using a randomly arranged sensor array, under the hypothesis that the position of one reference sensor and the geometry of the environment are known, and the other microphone positions are unknown. A minimum mean square error (MMSE) estimator that exploits early reflections is proposed. The MMSE estimator is solved by a grid search method that combines the information on early reflections estimated using a multichannel blind system identification and the time difference of arrivals between the reflections and the direct-path calculated with the image-source model. Simulations under different reverberant scenarios demonstrate the ability of the proposed approaches in localizing source and microphone.
Daniele Salvati, Carlo Drioli, Gian Luca Foresti
IEEE Signal Process. Lett.1
2015 Frequency map selection using a RBFN-based classifier in the MVDR beamformer for speaker localization in reverberant rooms
Daniele Salvati, Carlo Drioli, Gian Luca Foresti
INTERSPEECH1
2014 Incoherent Frequency Fusion for Broadband Steered Response Power Algorithms in Noisy Environments
abstract
The steered response power (SRP) algorithms have been shown to be among the most effective and robust ones in noisy environments for direction of arrival (DOA) estimation. In broadband signal applications, the SRP methods typically perform their computations in the frequency-domain by applying a fast Fourier transform (FFT) on a signal portion, calculating the response power on each frequency bin, and subsequently fusing these estimates to obtain the final result. We introduce a frequency response incoherent fusion method based on a normalized arithmetic mean (NAM). Experiments are presented that rely on the SRP algorithms for the localization of motor vehicles in a noisy outdoor environment, focusing our discussion on performance differences with respect to different signal-to-noise ratios (SNR), and on spatial resolution issues for closely spaced sources. We demonstrate that the proposed fusion method provides higher resolution for the delay-and-sum SRP, and improved performances for minimum variance distortionless response (MVDR) and multiple signal classification (MUSIC).
Daniele Salvati, Carlo Drioli, Gian Luca Foresti
IEEE Signal Process. Lett.1
2013 Adaptive Time Delay Estimation Using Filter Length Constraints for Source Localization in Reverberant Acoustic Environments
abstract
Adaptive time delay estimation based on blind system identification (BSI) focuses on the impulse responses between a source and a microphone to estimate the time difference of arrival (TDOA) in reverberant environments. In this letter, we consider the adaptive eigenvalue decomposition (AED) BSI method based on the normalized multichannel frequency-domain least mean square (NMCFLMS) algorithm. We show that the use of filter length constraints (FLC) based on the maximum TDOA between microphones improves the performance of the NMCFLMS filter for the localization of different sound types in highly reverberant environments. The experimental results demonstrate the improvement of the proposed method for reverberation times$({\rm RT}_{60})$of up to 2 s. Applications for this method include teleconferencing systems, musical interfaces, videogames, and monitoring systems.
Daniele Salvati, Sergio Canazza
IEEE Signal Process. Lett.1
2011 Multiple acoustic sources localization using incident Signal Power comparison
abstract
We present a novel approach to locate multiple acoustic sources in far-field environments, in order to solve an interesting problem in different application domain, such as: audio surveillance systems and soundscape analysis frameworks. This approach aims at finding a solution to the ambiguities in Direction Of Arrivals (DOAs) combination caused by simultaneous multiple sources. The algorithm is based on two steps: the separation of the sources by means of beamforming techniques and the comparison of the Incident Signal Power (ISP) spectrum by means of a spectral distance measure. We implemented a prototype, composed by two linear arrays, that has been successfully tested in a real noisy environment.
Daniele Salvati, Antonio Rodà, Sergio Canazza, Gian Luca Foresti
AVSS1